Every production incident hides the same expensive moment: the gap between a page firing and knowing what's wrong. Sherlock closes it — investigating like a senior engineer and posting an evidence-backed root-cause verdict in 2–5 minutes. Built, running, and hardened toward beta today.
Skillz runs a complex, distributed estate — hybrid, spread across multiple production accounts and clusters. When an alert fires, no single screen tells the story.
Today, a paged engineer logs into one dashboard after another just to figure out where the issue even is. Then it's hopping across accounts and clusters to work out what could have caused it, and a hunt through boards and runbooks to reach the actual root cause — all by hand, under pressure, while players are impacted and nothing has been decided. That's why diagnosis alone takes 10–45+ minutes.
And the outcome swings on who's paged: a seasoned engineer and a new engineer on-call run very different investigations. It's almost entirely work an AI SRE can do in parallel, in seconds — the same way, every time.
From the on-call engineer's seat, Sherlock is four steps — from page to root cause in about 2–5 minutes:
A production monitor detects an issue and pages on-call in Slack. The clock starts.
The on-call engineer says @sherlock triage this — or, for opted-in services, it auto-triages the moment the alert fires.
Metrics, logs, Kubernetes, recent deploys and code changes, and the team's own runbooks — gathered in parallel.
Root cause with confidence, who's impacted, and ranked fixes — every claim linked to its evidence. The engineer decides and acts.
Read-only by architecture: Sherlock finds the why; On-Call makes the call.
Sherlock runs today, exercised against real alerts. This isn't a plan or a slide — it's built.
An agentic AI investigation loop with hard guardrails — bounded turns, time and spend — ending in a structured, evidence-cited verdict.
Uses the team's own runbooks as the primary playbook, then extends the investigation to pin down root cause.
Correlates alerts with recent deploys and commits to spot the offending change — because the most common cause is a recent one.
Live & historical triages, incident cockpit, topology, connector health, audit — plus a leadership ROI dashboard.
Every alert class is graded for accuracy before on-call is asked to trust it — the launch gate for beta.
Detection and resolution time tracked per service across every triage — the hard numbers behind hours saved and coverage.
The worst case of any failure is a wasted investigation, never a changed system.
Sherlock can look but cannot touch — write actions do not exist in the system.
Sherlock supplies evidence and ranked recommendations; the decision is always a human's.
Production clusters accept no inbound connections — they answer read-only requests over a locked, signed channel.
Every question, every piece of evidence, every conclusion is recorded end-to-end.
Every investigation runs inside hard limits on turns, wall-clock time and model tokens.
A production beta with a cohort of real on-call teams — hardened, secured, gated on a measured accuracy bar. Tracked via verdict-acceptance rate and MTTD/MTTR.
All priority alert classes, 10+ services via self-serve onboarding, high availability.
Conversational triage in Slack, Jira handoff and owner outreach, and guarded, reversible auto-remediation.
Built an agentic AI for SRE — an autonomous AIOps agent that reasons with an LLM to investigate incidents on its own, calling live observability tools and retrieving runbooks and incident memory via RAG. An engineer building for engineers.
Built in-house — by an engineer, for engineers.