Skip to content
Jatin Sikri
All work

BALANCED+ Inc. · 2025–2026

A multi-night production incident, and the recommendation that came out of it

Production went down several nights running. The pressure in that situation is to restore service and move on, which leaves the cause in place. I drove it to an actual diagnosis and wrote the RCA.

Role
Incident lead and author of the RCA
Focus
Incident management, Root-cause analysis, SQL Server, Hyper-V

The problem

A production Hyper-V and SQL Server environment was failing across consecutive nights. Each night it came back up, and each night nothing had been established about why it went down, so the next failure was already scheduled.

What I did

I led the resolution and drove it to an actual diagnosis: an antivirus filter-driver interaction and NVMe storage faults, rather than the application layer where the symptoms appeared.

I produced a formal root-cause analysis and an infrastructure-replacement recommendation, and added proactive SQL Server alerting through Database Mail so the next occurrence is caught early rather than reported by the client.

The business outcome was a decision the client could actually make. A recurring outage became a documented cause and a specific recommendation to weigh, instead of another night of restored service and no answer.

Want to talk about work like this?

I am open to technical account management, service delivery, and client-facing technology leadership roles in Toronto.