Index

Recovery Engineering

Recover the First Useful Service, Not Every Server

Six components sit on the recovery checklist: a database, queue, cache, carrier adapter, identity service, and API. Restore them in that order and the first four can be perfectly healthy while no customer can retrieve a shipment.

The OPS-05 fixture gives every recovery action a deterministic duration. Inventory order reaches an accepted customer journey at 540 virtual milliseconds. Restoring identity, shipment data, and the API first reaches a bounded read-only journey at 240. Both schedules finish all six components at 540.

Those values are synthetic scheduling inputs, not a recovery benchmark. The useful observation is why the schedules differ: component order follows nouns in an inventory; journey order follows the dependencies of behavior.

Recovery should first restore the smallest verified path to one useful customer outcome, then widen from that working slice. “Useful” must include authorization, data integrity, and honest degradation. It cannot mean the first process which returns 200.

A component list does not contain a customer journey

The fixture models one private shipment-status request:

customer -> identity -> tracking API -> shipment database
                              |              |
                              |              +-> last-known status
                              +-> cache (optional acceleration)
                              +-> queue -> carrier adapter (fresh updates)

Identity, the API, and the shipment database form the hard path. Without identity, the service cannot establish that user-100 owns shipment-100. Without the database, the API has no accepted shipment or event history. Without the API, healthy foundations still expose no customer behavior.

The cache changes the cost of a read, not its correctness. A miss can fall through to the database, so the fixture classifies cache recovery as optional for this journey.

The queue and carrier adapter are more interesting. They are required for a fresh external update, but the database already contains a last-known status. The fixture allows that status to be read while update ingestion is down only under an explicit degraded contract:

{
  "shipment": "shipment-100",
  "status": "in_transit",
  "updated_at": "2026-10-01T09:10:00Z",
  "freshness": "degraded",
  "refresh_enabled": false
}

That response is useful for a narrower purpose. It does not pretend that tracking is current, and it does not offer a refresh action which cannot work. Once the queue and carrier return, a separate update probe must ingest event three before the fixture calls the journey fully capable.

Drawing the graph is not the hard part. Classifying each edge is. An optional dependency must have a tested fallback. A degraded dependency must have a bounded user-visible contract. Everything else remains hard, however convenient it would be to omit during recovery.

Stop the first clock after behavior, not readiness

OPS-05 compares four sequential schedules on a virtual clock:

ScheduleFirst useful journeyFull capabilityTotal actions
Journey first240480540
Readiness first400540540
Inventory order540540540
Invalid narrow orderNeverNever220

Journey-first recovery spends 80 units on identity, 120 on the database, and 40 on the API. It then runs the authenticated shipment probe. Queue and carrier recovery continue afterward; full update capability arrives at 480. Cache warming finishes last without changing correctness.

Inventory order starts with the database, queue, cache, and carrier. Each node becomes ready on schedule. The customer path remains broken until identity and the API finish at 540.

Readiness-first order prioritizes short actions: API, cache, then identity. It looks active and accumulates green components quickly. The customer probe still waits for accepted data at 400, and fresh updates wait for the carrier until 540.

The invalid order is nominally fastest. It starts the API, database, and cache in 220 units, skipping identity. The fixture refuses to record a first-useful time. A private shipment read without authorization is not degraded service; it is a different and unsafe behavior.

This is why the recovery clock needs an acceptance event. Process health answers whether a component can run its probe. Application acceptance answers whether the chosen dependency slice now produces the outcome used to justify its priority.

AWS’s reliability guidance makes a similar distinction: workload availability is based on delivering business value, not merely component health . The fixture turns that boundary into an executable checkpoint for one small system.

Degradation is a contract, not an excuse

The queue and carrier can be postponed only because the fixture states what remains true without them. The known user is authenticated. Ownership is checked. The shipment and ordered event history pass integrity checks. The last update time is visible. Freshness is labelled degraded. Refresh is disabled.

Remove any of those conditions and the fallback fails:

  • returning private state anonymously is rejected;
  • a known user without ownership is rejected;
  • two database rows without shipment-100 are rejected;
  • stale data without the freshness marker is rejected; and
  • claiming full capability without queue and carrier fails the update probe.

AWS describes graceful degradation as turning applicable hard dependencies into soft dependencies . That qualifier matters. Identity is not applicable here because anonymous access changes the security contract. Data integrity is not applicable because a response about the wrong shipment does no useful work. Fresh ingestion is applicable for a bounded last-known read because the user can see exactly what was lost.

Google’s SRE guidance recommends serving degraded results while doing as much useful work as possible , and warns that rarely exercised degraded paths carry their own risk. A fallback which exists only in a diagram should still be treated as a hard dependency during recovery. The contract becomes credible when the reduced path is exercised regularly and observable as a distinct mode.

Degraded operation also needs an exit. In this fixture, queue and carrier readiness does not automatically declare success. A new carrier event must travel through ingestion and appear as sequence three. That probe closes the gap between components returning and capability returning.

Optional dependencies can become accidental blockers

Marking the cache as hard moves journey-first acceptance from 240 to 300 virtual units. Nothing about authorization, shipment state, event order, or freshness changes. Recovery gets slower solely because the graph confused an optimization with a correctness dependency.

The opposite mistake is just as easy. Marking the queue and carrier optional for full recovery makes the schedule look finished at 240. The fresh-update probe then fails. Dependency labels belong to a specific behavior: the cache is optional for a correct read; the queue and carrier are optional for a declared last-known read; all three may be required for a latency or freshness objective not modelled here.

That specificity keeps the graph reviewable. “Database is critical” says little. “The authenticated last-known tracking read requires identity, shipment data, and the API; current tracking additionally requires queue and carrier ingestion” can be challenged, tested, and revised.

NIST’s business-impact guidance starts from mission-essential functions and asks which assets enable those objectives . The fixture applies the same direction at application scale: outcome first, supporting components second. It does not imply that this shipment journey would outrank payments, safety controls, or contractual data delivery in another business.

Sometimes the infrastructure really does come first

Journey-first is not a universal sorting algorithm. A shared identity plane, network control plane, or primary database may unlock many high-value journeys at once. Restoring that foundation first can be the fastest route to business value, even if its name appears low in an infrastructure diagram.

Risk can also outrank customer visibility. Containment, financial integrity, audit capture, legal retention, or a safety interlock may need to recover before any customer traffic. A narrow read path is invalid if it can write against inconsistent state, bypass authority, conceal material staleness, or make later reconciliation unsafe.

Parallel recovery changes the schedule too. Independent teams may restore the hard journey path and the remaining foundations at the same time. The useful graph still matters because it identifies the acceptance checkpoint and the components which can actually block it; the sequential numbers no longer predict elapsed time.

The recommendation reverses when the narrow path is unsafe, when a shared foundation dominates several more important journeys, or when platform dependencies impose an unavoidable order. Those are reasons to choose a different path through the graph, not reasons to return to an unranked list.

A recovery plan can keep the inventory. It should add something the inventory cannot express:

first outcome
hard dependency path
allowed degraded behavior
integrity and authorization probes
first-useful acceptance
full-capability acceptance
conditions which invalidate the order

Recovering every server may eventually be necessary. It is not a useful first checkpoint. Pick the customer or business behavior which matters first, prove the smallest safe path to it, and let that acceptance result—not the number of green boxes—decide when recovery has begun to work.