Break it before your customers do.

A game day run once proves the system survived that afternoon. Firedrill runs the same failures every week against a namespace or a compose stack, scores what came back, and keeps a dated record an auditor can be handed.

shape: monthly subscription, per estate grows from: chaos-toolbox, backup-verify, dark-canary

firedrill · plan checkout-prod weekly, Sunday 03:00 UTC

run 2026-09-21 · 12 failure modes · blast radius within limits

01survivedpod terminationrecovered in 2.1s
02survivedmemory pressure 90%degraded cleanly, no errors
03faileddisk fill 95%writes blocked, no alert fired
04survivednetwork latency +500mstimeouts held, retries bounded
05failedDNS withheldcascade: auth at 4s, gateway 502 at 8s
06failedcertificate expiryTLS handshake errors at the edge
07surviveddependency unavailablecircuit breaker opened
08degradedclock skew +5mtoken validation errors for 40s
...

survived 7 degraded 2 failed 3 previous run: 6 · 2 · 4

record: evidence/checkout-prod/2026-09-21

Example output. The last line is the point: a dated record of what was tested and what came back.

Reading 01
survived=7/12
failure modes the estate came through on the last run
Reading 02
time_to_impact=4s
from DNS withheld to the first 502 at the gateway
Reading 03
last_drill=2026-09-21
the date on the newest record, which is what an auditor asks for

Readings from the last run.

The problem

Resilience is tested once and asserted forever

Most teams have run a game day. Almost none have run the same one twice. A pod was killed in a staging namespace before the launch, the service recovered, and the slide has said the platform is resilient ever since. The dependency graph has changed, a sidecar was added, a timeout was tuned, and the claim has not been re-examined.

The tools that inject failure exist and are free: chaos-toolbox, Chaos Mesh, LitmusChaos, the fault-injection services the cloud providers sell. What they produce is an experiment log for the engineer who ran them. What a director, an auditor or an insurer asks for is different: which failure modes the estate survives today, which it does not, and when that was last checked.

DORA asks financial entities and their suppliers for resilience testing with evidence, and NIS2 asks for continuity and incident-handling measures that hold under stress. Both turn an occasional exercise into a recurring obligation, and a recurring obligation needs a record.

Failure scenarios

Twelve ways your system can break

Each scenario runs in isolation against the real estate, not a mock and not a simulation, and each one carries a declared blast radius and a stop condition.

Pod / Container KillCritical

Terminates a pod or container at random, then measures recovery time and how requests were routed meanwhile.

Memory Pressure 90%High

Consumes memory until the container sits at 90% utilisation. Tests OOM handling and graceful degradation.

Disk Fill 95%High

Fills the filesystem to 95% capacity. Reveals missing disk alerts and write-path failures.

Network Latency +500msMedium

Injects 500ms of latency on inter-service traffic. Tests timeout configuration and retry logic.

DNS Resolution FailureCritical

Blocks DNS lookups for targeted services. Exposes missing caches and cascade risks.

Certificate ExpiryCritical

Replaces a service certificate with an expired one. Tests TLS error handling and alerting.

Dependency UnavailableHigh

Takes a downstream dependency fully offline. Validates circuit breakers and fallback paths.

Network PartitionCritical

Isolates a service from the rest of the cluster. Tests split-brain behaviour and leader election.

CPU SaturationMedium

Pins CPU to 100% on a target container. Checks autoscaling triggers and request queuing.

Credential RevocationHigh

Invalidates API keys and service-account tokens mid-request. Tests token refresh and auth fallbacks.

Clock Skew +5minMedium

Shifts the system clock forward by five minutes. Reveals time-dependent logic and JWT validation gaps.

Config CorruptionHigh

Injects invalid values into environment variables and config maps. Tests startup validation and safe defaults.

How it works

The same failures, on a schedule, scored the same way each time

01

Point

firedrill init reads a Kubernetes namespace or a docker-compose file and lists the services, their health checks and their dependencies. Anything it cannot infer, the operator records once.

02

Break

Each failure mode runs in isolation, on the cadence set: a pod killed, a disk filled, latency added to one hop, DNS withheld, a certificate expired, a credential revoked, a dependency made unavailable, a clock skewed. Health checks and error rates are watched throughout, and the run stops the moment a blast radius exceeds the limit declared for it.

03

Score

Each mode is marked survived, degraded or failed, with the time to first impact and the time to recovery. The score compares run to run because the failures and the thresholds only change when the plan does.

04

Record

Every run leaves a dated record: what was injected, what was observed, what broke, and the remediation the engine suggests. The record is the deliverable, and it lives in the same dashboard as Lastresort's restore log, so resilience and recovery evidence sit side by side.

Detailed reports

Every failure tells you how to fix it

A failed row is not the deliverable. Each mode that breaks produces a report: what was injected, what it took down and when, what that cost, and the remediation to apply.

05 · DNS resolution failure

failed
What was injected

DNS resolution for payment-service.internal blocked for 30 seconds, with the rest of the namespace left alone.

What happened

A cascade. The auth service failed its health checks at T+4s, having no cached resolution to fall back to. The API gateway returned 502 to every request at T+8s. Both stayed down until the injection window closed at T+30s.

Impact

One internal name stopped resolving and checkout stopped serving eight seconds later. No alert fired on the auth health checks, so the outage would have been found by a customer rather than by a pager.

Remediation
  1. Add a DNS caching sidecar, such as dnsmasq or CoreDNS with the cache plugin.
  2. Retry DNS lookups with bounded exponential backoff instead of failing the health check on the first miss.
  3. Put a circuit breaker between the auth service and payment-service, so one unreachable dependency degrades a path instead of the gateway.
  4. Alert on consecutive auth health-check failures, which nothing currently watches.
Time to first impact

T+0s T+4s T+8s T+30s

T+4s auth down · T+8s full outage · T+30s injection ends

Example output. The row in the scorecard is the same drill: the report is what the row expands into.

Who it is for

Teams whose resilience claim is older than their architecture

Case 01

The platform team with a resilience slide

A game day was run before the last launch and nothing since. The estate has changed and the claim has not been re-examined.

Case 02

The supplier being asked for evidence

A financial-services customer, an auditor or an insurer asks how resilience is tested and when. A dated log of drills answers the question a slide from last year cannot.

What it costs

Priced per estate, on a monthly subscription

How it would be priced

One price covers a namespace or a compose stack, every failure mode, on any cadence. Charging per drill would price the thing worth doing more often. The request form asks for the shape of the estate so the tiers fit real ones.

Request access

Say what you would drill first

Requests decide which failure modes and which platforms are supported first.

Free for the whole beta. The first twenty-five teams keep 50% off for twelve months after launch, locked in at signup.

Your address is used for this request and nothing else. One email when there is something to show.

Reasonable objections

The questions this gets asked first

Chaos Mesh and LitmusChaos are free.
They are, and they inject failure well. They leave the operator to decide what to run, when, what counts as survival, and where to keep the results so an auditor can read them. Firedrill is that layer: a fixed set of failures on a schedule, a score that compares across runs, and a dated record, with chaos-toolbox doing the injection underneath.
A drill against production is a risk in itself.
Every failure mode carries a declared blast radius and a stop condition, and the run halts the moment either is crossed. The first runs go against a staging namespace; production is opted into per failure mode, and the record says which.
Is a scorecard really resilience testing?
It is the repeatable part of it. The record claims exactly what was tested and what came back on that date, and nothing wider, which is what makes it defensible when somebody pushes back.