Skip to content
Jatin Sikri
All work

BALANCED+ Inc. · 2025–2026

A multi-night production incident, and the recommendation that came out of it

I led resolution of a multi-night Hyper-V and SQL Server production incident, diagnosed the root cause to the storage and antivirus layer, and produced the formal RCA and infrastructure-replacement recommendation.

Role
Incident lead and author of the RCA
Focus
Incident management, Root-cause analysis, SQL Server, Hyper-V

The problem

A production Hyper-V and SQL Server environment was failing across consecutive nights. The pressure in that situation is to restore service and move on, which leaves the cause in place and guarantees a repeat.

What I did

I led the resolution and drove it to an actual diagnosis: an antivirus filter-driver interaction and NVMe storage faults, rather than the application layer where the symptoms appeared.

I produced a formal root-cause analysis and an infrastructure-replacement recommendation, and added proactive SQL Server alerting through Database Mail so the next occurrence is caught early rather than reported by the client.

Want to talk about work like this?

I am open to technical account management and professional services roles in Toronto.