Skip to content

Production on fire? We pick up the pager.

Same-day embed. We triage, stabilize, and ship the fix. You get a written post-mortem, the runbook that should have existed, and a 90-day hardening plan so it doesn’t happen again.

Rate $180 – $250 USD / hour · fixed-fee and retainer available

What you get

Triage
we get on the call within 15 minutes, find the active failure, and stop the bleeding.
Stabilization
kill switches, traffic shedding, rollback, region failover, whatever it takes.
Root cause
written analysis within 48 hours of stabilization.
Post-mortem
blameless, written for the team and the board, with action items and owners.
Runbook
the runbook that should have existed, so the next on-call knows what to do.
90-day hardening
a phased plan to fix the systemic gaps that allowed the incident.

How we engage

  1. 01

    Emergency engagement (first 8 hours)

    Triage + stabilization. You pay for hours worked, not a retainer. No minimum.

  2. 02

    Stabilization (8–40 hours)

    Root cause + immediate fixes + a written post-mortem.

  3. 03

    Hardening retainer (3 months, post-incident)

    The 90-day plan, in writing, executed. Prevents the next incident.

Stack we work in

On-call

PagerDuty, Opsgenie, FireHydrant, incident.io

Observability

Datadog, Honeycomb, Grafana, Sentry, Lightstep

Forensics

CloudTrail, VPC flow logs, packet captures, k8s audit, custom logs

Communication

Zoom, Slack war rooms, Statuspage, dedicated incident channels

Reference architectures

Anonymized patterns from real engagements. Client names omitted; details available under NDA.

Series A SaaS — 6-hour outage on Black Friday

Same-day embed. Found a connection pool exhaustion from a new background job. Killed the job, restarted, restored service. Root cause: missing max-connection limit. Fix shipped in 24h, post-mortem in 48h, runbook in a week.

Healthtech — HIPAA breach scoping

Embedded within 2 hours of detection. Scoped the breach, contained the access, coordinated with the privacy officer and outside counsel. Cleared the breach notification timeline with 6 hours to spare.

Logistics — 3-region failover fail

Diagnosed a regional failover that didn’t actually fail over. Implemented a tested, automated failover with quarterly drills. The next regional incident recovered in 47 seconds.

Questions we get asked

How fast can you respond?

Within 15 minutes for the initial triage call. We’re in your time zone (or close to it) and we don’t sleep on active incidents.

Do you do security incidents?

Yes — we coordinate with your security firm and outside counsel. We do the technical containment, the scoping, and the post-mortem. We do not provide legal advice.

Can you take over on-call permanently?

Optional — 24/7 managed on-call with a 15-minute response SLA. Most clients use us during the 90-day hardening period, then take the pager back themselves.

What about the post-mortem?

Blameless, written for the team, written for the board. We turn it into action items with owners and dates, then we own following up.

Do you do table-top exercises?

Yes — quarterly or annual. We run the scenario, your team responds, we write up the gaps. Most teams find 3–5 process gaps that are cheap to fix and prevent a real incident.

Let’s scope it properly.

A 30-minute call. No deck, no pitch — we read your repo or your architecture diagram and tell you what’s realistic.