fin1te/dsv1.0.0
  1. ds
  2. Components
  3. Incident replay

Incident replay

Failing states and the timeline rail.

Fig. 01Checkout outage5 components · 4 links
6 stepsWalk through with the arrows
14:0214:34
Impact 28m
rolloutCustomersCheckout API6 podsPgBouncerpool: 200PostgreSQLprimaryArgo CDdeploys

A walkthrough whose steps carry a time and status changes becomes a replay. Components go amber or red and recover, data stops on links touching a failed component, and the rail under the step bar places each event at its real time with the impact window in red. Press play.

Written in topo

step 14:06 "PgBouncer's queue starts to grow" bouncer warn=bouncer,api
step 14:09 "Postgres hits max_connections" db down=db,bouncer
step 14:26 "Rollback lands, Postgres recovers" argo db ok=db,bouncer warn=api

States

ClassOnWhat it does
.inc-down.nodeRed outline and wash, a slow red glow, red status dot.
.inc-warn.nodeAmber outline and wash.
.inc-recovered.nodeOne green flash when a component comes back.
.inc-set.nodeHides the authored status dot while a replay dot is shown.
.inc-broken.edgeRed dashed line, flow hidden: nothing moves.
.inc-slow.edgeAmber line, flow slowed to 3.3s.

Control and fallback links are left alone: scrapes, deploys and pages keep working during an outage, which is usually how the fix arrives.

Rail

<div class="irail">
  <div class="rail">
    <div class="impact" style="left:12%;width:72%"></div>
    <button class="tick sev-warn cur" style="left:12%" aria-label="Step 2"></button>
    <button class="tick sev-down" style="left:21%" aria-label="Step 3"></button>
    …
  </div>
  <div class="ends"><span>14:02</span><span>14:34</span></div>
  <span class="sum bad">Impact 28m</span>
</div>
ClassOnWhat it does
.iraildivRail, end labels and summary.
.impact / .opendivThe red band from first failure to recovery; .open fades out when it never recovered.
.tick.sev-*buttonOne per step, coloured by the worst change it makes: down, warn, ok, none.
.cur.tickThe current step.
.sum.badspanImpact length in red.