Project Sherlock
STATUS · LIVE · BETA READY
AI SRE Agent · Built in-house at Skillz

Incident minutes are customer minutes.

Every production incident hides the same expensive moment: the gap between a page firing and knowing what's wrong. Sherlock closes it — investigating like a senior engineer and posting an evidence-backed root-cause verdict in 2–5 minutes. Built, running, and hardened toward beta today.

The investigation window — the largest, most variable slice of every incident
time from page → root cause
Today — manual investigation10–45+ min
With Sherlock~2–5 min
015 min30 min45 min
What takes an engineer 10–45 minutes of manual digging happens automatically, every alert, at any hour — often before they're fully online.
10–45+ min
of hands-on investigation per incident — before a fix can even start
~30 / wk
production alerts paging on-call engineers
3:00 AM
when it happens most — cold starts, context-switching, pressure
01 — THE PRODUCT

Sherlock: an AI teammate on every incident

Skillz runs a complex, distributed estate — hybrid, spread across multiple production accounts and clusters. When an alert fires, no single screen tells the story.

WHY PURPOSE-BUILT

One product to stitch a fragmented estate

Today, a paged engineer logs into one dashboard after another just to figure out where the issue even is. Then it's hopping across accounts and clusters to work out what could have caused it, and a hunt through boards and runbooks to reach the actual root cause — all by hand, under pressure, while players are impacted and nothing has been decided. That's why diagnosis alone takes 10–45+ minutes.

And the outcome swings on who's paged: a seasoned engineer and a new engineer on-call run very different investigations. It's almost entirely work an AI SRE can do in parallel, in seconds — the same way, every time.

From the on-call engineer's seat, Sherlock is four steps — from page to root cause in about 2–5 minutes:

STEP 1

Alert fires

A production monitor detects an issue and pages on-call in Slack. The clock starts.

STEP 2

Sherlock engages

The on-call engineer says @sherlock triage this — or, for opted-in services, it auto-triages the moment the alert fires.

STEP 3

It investigates

Metrics, logs, Kubernetes, recent deploys and code changes, and the team's own runbooks — gathered in parallel.

STEP 4

Verdict lands

Root cause with confidence, who's impacted, and ranked fixes — every claim linked to its evidence. The engineer decides and acts.

Read-only by architecture: Sherlock finds the why; On-Call makes the call.

SEE IT IN ACTION

Watch Sherlock triage a live alert

OPEN IN DRIVE →
02 — WHAT'S RUNNING TODAY

Not a prototype — a working product

Sherlock runs today, exercised against real alerts. This isn't a plan or a slide — it's built.

BUILT

Autonomous triage engine

An agentic AI investigation loop with hard guardrails — bounded turns, time and spend — ending in a structured, evidence-cited verdict.

BUILT

Runbook-first investigation

Uses the team's own runbooks as the primary playbook, then extends the investigation to pin down root cause.

BUILT

Code-aware root cause

Correlates alerts with recent deploys and commits to spot the offending change — because the most common cause is a recent one.

BUILT

Orchestration console

Live & historical triages, incident cockpit, topology, connector health, audit — plus a leadership ROI dashboard.

BUILT

Eval & shadow harness

Every alert class is graded for accuracy before on-call is asked to trust it — the launch gate for beta.

BUILT

MTTD / MTTR analytics

Detection and resolution time tracked per service across every triage — the hard numbers behind hours saved and coverage.

BUILT

Six evidence sources — read-only across our whole operational surface

CONNECTED DatadogELKKubernetesAWSGitHubConfluence
03 — BUSINESS IMPACT

Where it moves the business

CONSISTENCY
Every alert, the same rigor. A 3am cold start gets the same senior-level investigation as your best engineer at their desk — no key-person dependency.
PRODUCTIVITY
~5–20+ engineer-hours / week returned from investigative toil to feature work — tracked as "hours saved."
MTTD / MTTR
Both drop. The slowest, most human-dependent link is removed from the critical path — measured per service.
CUSTOMERS
Minutes, not hours. Players feel incidents for less time — protecting player experience and revenue.
ON-CALL LIFE
Fewer 3am cold-starts — on-call opens to a credible hypothesis and its evidence, not a blank dashboard.
KNOWLEDGE
Compounds, doesn't walk out. Every triage is recorded and replayable — expertise stays in the building.

Trust & guardrails — engineered in, not policy

The worst case of any failure is a wasted investigation, never a changed system.

Read-only, provably

Sherlock can look but cannot touch — write actions do not exist in the system.

Humans make every call

Sherlock supplies evidence and ranked recommendations; the decision is always a human's.

Zero exposure to production

Production clusters accept no inbound connections — they answer read-only requests over a locked, signed channel.

Fully audited & replayable

Every question, every piece of evidence, every conclusion is recorded end-to-end.

Bounded spend

Every investigation runs inside hard limits on turns, wall-clock time and model tokens.

04 — THE PATH FORWARD

From beta-ready to a teammate that acts

PLANNED · SCALE

Across the estate

All priority alert classes, 10+ services via self-serve onboarding, high availability.

AHEAD · VISION

Acts, not just advises

Conversational triage in Slack, Jira handoff and owner outreach, and guarded, reversible auto-remediation.

Abhishek Kumar

Abhishek Kumar

Lead Cloud Engineer

Built an agentic AI for SRE — an autonomous AIOps agent that reasons with an LLM to investigate incidents on its own, calling live observability tools and retrieving runbooks and incident memory via RAG. An engineer building for engineers.

Sherlock finds the why.
On-Call makes the call.

Built in-house — by an engineer, for engineers.