The operational incident register: mis-picks, mis-ships, parcels lost in transit, gate mismatches, count escalations, dead devices, integration outages. Severity, owner, an SLA clock that runs, a root cause from a closed list, recurrence detection, and the module that owns the fix. For the hash-chained record of every data event see the audit log; this register is the human narrative that explains it.
INC-2026-0729-002 is a SEV-2 past its resolution SLA by 3h 40m. Gate G1 at HQ_319 has been reading without a matching ops2 dispatch since 09:12, 31 units affected, AED 24.6K at risk. Three drift cases are already claiming it as their explanation. A gate that reads and a dispatch log that does not is the exact ambiguity the authority matrix exists to settle: RFID wins on physical existence, so those units left the building.
2 recurrence groups fired this week.packaging_fault at HQ_319 is on its 4th occurrence in 61 days and auto-escalated from SEV-3 to SEV-2 on open. A problem that keeps happening is a bigger problem than one that happened once, and the register now says so without waiting for somebody to notice the pattern.Recurrence groups →
v2's incident register could not close a single drift case. Every one of its fifteen entries was an engineering incident: a Postgres pool exhaustion, an HMAC rotation gap, a bad CSD deploy, a Shopify regional outage, a Cloudflare header, a Tigris throttle storm. Exactly one was operational. Meanwhile the things that actually explain inventory drift are a mis-pick, a mis-ship, a parcel lost in transit, a gate read with no dispatch behind it, and a carton opened and never decremented. None of those had a home. i3's kind is a closed list of twelve spanning both worlds, and owning_module names which of the fifteen i3 modules ships the structural fix, so an incident is a routable work item and not only a story.
Variant SKU rename MSK15-B to MSK15-BLK orphaned 22 encoded tags
Ahmad Y.
03 SKU
24-07 11:20
ack 4h
resolved 8d
22
0
sku_rename
×2
waived
1 claimed
INC-2026-0722-005
SEV-2
damage_cluster
Packaging fault on PO-2026-109: 6 crushed cartons of 22
Khaled A.
06 GRN
22-07 09:15
ack 11m
resolved 17h
18
8,640
packaging_fault
×4
written
-
INC-2026-0721-002
SEV-1
security
Write-off posted with no approval id during a shadow-mode gap
Ahmad Y.
09 Damages
21-07 14:02
ack 4m
resolved 1h55
2
4,180
gate_bypass
-
written
-
INC-2026-0719-007
SEV-3
mis_pick
Composite BX-MSR31 picked as a parent, components never decremented
Sara M.
03 SKU
19-07 12:50
ack 1h26
resolved 2d
9
8,100
composite_handling
-
n/a
1 claimed
INC-2026-0718-003
SEV-3
gate_mismatch
G1 antenna 2 double-reading outbound transfers after a reconfig
Khaled A.
12 RFID
18-07 09:12
ack 22m
resolved 6h
12
4,880
device_desync
×2
n/a
2 claimed
INC-2026-0716-009
SEV-4
external_dependency
Shopify eu-west webhook relay queued 1,400 events for 2h
Ahmad Y.
14 Events
16-07 08:33
ack 18m
resolved 2h4
0
0
upstream_outage
-
waived
-
INC-2026-0714-006
SEV-2
process_breach
Partial replenishment moved stock with no movement row written
Ahmad Y.
04 Locations
14-07 10:08
ack 9m
resolved 21h
148
51,300
swallowed_error
-
written
4 claimed
INC-2026-0711-004
SEV-3
count_variance
HQ_325 blind count variance -9, escalated by dual-approval rule
Sara M.
10 Counts
11-07 07:00
ack 1h46
resolved 3d
9
3,780
count_method
×3
n/a
1 claimed
18 of 36 shown · sorted by detection descending · keys J/K move · Enter open · A acknowledge · E escalate · R resolve · C set root cause13 of 18 are operational rather than technical. In v2's register that ratio was 1 in 15.
Incident rate · 12 weeks
Stacked by severity, with a 4-week moving average and the SLA-breach count as a second series. The rate matters more than the count: w26 has more incidents than w22 and fewer breaches, which means detection improved rather than the operation degrading.
Bars descending by count, cumulative percentage overlaid, 80% reference line marked. Four causes of thirteen account for 79% of incidents, which is the conversation worth having.
The cumulative line crosses 80% at the sixth cause. Those six own four modules between them: 06 GRN (packaging, supplier data), 12 RFID (device desync), 04 Locations (open carton, bin instability), 10 Counts (count method), 08 Reverse (carrier process). That is the fix backlog, in priority order, derived rather than argued.
Top 4 causesNext 4Long tailCumulative %
Recurrence groups
Two or more incidents sharing a root cause and an affected entity inside 90 days. The second and later occurrences escalate one severity level automatically, because a problem that keeps happening is a bigger problem than one that happened once.
2 active · 6 incidents
packaging_fault at HQ_319 · 4 occurrences in 61 days
auto-escalated SEV-3 to SEV-2
Four suppliers, one carrier lane, one packaging spec, and the crushed-carton count rising by one each time. The group's owning module is 06 and its open action is a packing-spec revision. Read one at a time these are four small problems; read as a group they are one worsening one.
device_desync at gate G1 · 2 occurrences in 11 days
auto-escalated SEV-3 to SEV-2
INC-2026-0718-003#1antenna 2 double-reading after a reconfigSEV-3
INC-2026-0729-002#2reading with no matching dispatch · 11d laterSEV-2
The first was closed as an antenna retune. The second is the same gate 11 days later with a different symptom, which is the signature of a fix that treated the symptom. Owning module 12, and the open action is a heartbeat check with a quorum rather than a single probe.
Groups need 90 days of history to mean anything, which is why this panel renders gate-blocked rather than empty during the first quarter of operation. An empty recurrence panel and an uncomputable one look identical and mean opposite things.
Incident to drift handoff
When an incident explains a drift case, module 13 closes that case citing the incident. The claim is made here, accepted there, and it is reversible: if the incident is voided every drift case it closed reopens automatically.
INC-…0729-002DRF-0142MSR31-07 at HQ_319 · drift window 09:12 to 14:40 sits inside the incident window-128,988accepted
INC-…0729-002DRF-0143LBA01 at HQ_319 · same gate, same window-918,450accepted
INC-…0729-002DRF-0147FYN01 at HQ_319 · claimed 22m ago, module 13 has not accepted yet-104,200claimed
INC-…0726-004DRF-0131Carton MC-4402 contents at HQ_319 · opened, never decremented+247,920accepted
INC-…0714-006DRF-0118Picking-face replenishment at HQ_319 · 148 units across 11 SKUs-14851,300accepted
INC-…0728-002DRF-0138JAE17 at MCC · module 13 rejected: the count method explains the variance but not its sign+62,520rejected
INC-…0703-011DRF-0092TNS01 at YAS · incident later voided as a duplicate, so this case reopened on the next cron cycle-31,410revoked
One drift case, at most one incident. A partial unique index on (drift_case_id) WHERE state IN ('claimed','accepted') enforces it, so two incidents cannot both take credit for the same variance and make it look explained twice.
Severity and SLA
The clock is computed from severity at open time and frozen on the row. Escalating later sets a new clock through an event; it never silently rewrites the original, because the original is what the first responder was actually held to.
SEV-1Stock truth is wrong or unavailable in a way customers or Finance can see. All hands.ack 15m · fix 4h
SEV-2A material quantity of stock is unaccounted for, or a sibling contract is silently failing.ack 1h · fix 24h
SEV-3A process broke, a workaround exists, and the ledger can still be reconciled by hand.ack 4h · fix 5d
SEV-4Data quality or cosmetic. No stock effect, no customer effect, worth recording so it can recur into something.ack 1d · fix 30d
Post-mortems are mandatory at SEV-1 and SEV-2. A CHECK constraint forbids post_mortem_status = 'not_required' at those severities, and a waiver must name a person and a reason. "We were busy" is a reason; a blank field is not.
Fix ownership
Which of the fifteen modules ships the structural fix for each open root cause. Read this as a backlog, not a chart.
Module
Open causes
Inc
AED
04 Locations
iwms_open_carton, bin_instability, bin_adjacency
29
62,760
12 RFID
device_desync, hardware_fault
21
29,480
06 GRN
packaging_fault, supplier_data
29
18,240
10 Counts
count_method
11
6,300
08 Reverse
carrier_process, pack_verification
9
11,290
03 SKU
sku_rename, composite_handling
5
8,100
14 Events
sibling_outage
4
0
09 Damages
gate_bypass
1
4,180
external
upstream_outage
1
0
unknown is a valid root cause with a null owning module. It is a decision that nobody could name the cause, recorded as such, and it is not the same thing as an empty field. Three incidents closed as unknown in the last 90 days.
Post-mortem template
What went well
Detection time, blast-radius containment, the rollback decision, comms clarity. Name what to keep, not what to praise.
What went poorly
Missed signals, tooling gaps, runbook inaccuracies, dependency surprises. Blameless about people, unsparing about systems.
Action items
Concrete, owned, dated. No "investigate further" unless the investigation itself is scoped, owned and dated.
Structural fix and owning module
i3 addition. Which of the fifteen modules changes, and whether the change is a constraint or a code path. A constraint cannot be forgotten; a code path can.
written 12pending 3waived 8
Open action items
Add a quorum probe to the gate heartbeat, 2 of 3 Khaled · due 31-07 · INC-…0729-002
Re-level MCC shelf-A3 and re-inspect the adjacent bays Store MCC · due 30-07 · INC-…0729-001
Revise the packing spec with all four affected suppliers Ahmad · due 05-08 · packaging_fault group
Make carton_events.event_type a foreign key so an unregistered type cannot be swallowed Module 04 · due Phase 5 · INC-…0714-006
Suppress EPC-decay findings during a device_failure window Module 09 · done 28-07 · INC-…0727-009
Incident register
loading
SLA clocks are computed server side and arrive with the rows, because a clock that starts rendering before its deadline is known would count up from the wrong zero.
No incidents in this window. Nice.
Genuinely empty, not filtered. Worth a glance at the drift register before believing it: 14 drift cases are open and 3 of them have no explanation, and an unexplained drift case with no incident behind it usually means an incident that nobody opened.
The incident service returned 500 after 6 seconds. Nothing was changed and no SLA clock was affected: deadlines are stored on the row, so an outage here does not quietly extend anyone's deadline.
Error id inc-7c04a1 · 29-07-2026 14:45:33 GST · read failed, no write attempted
Your grant on app i3 is viewer. Narratives, timelines, root causes and drift claims are all readable, because an incident register that only responders can read stops being an organisational memory. Open, acknowledge, escalate, resolve and claim are absent.
Nightly hash-chain verify. Opening an incident still works, because refusing to record an incident during an outage is how an outage stops being recorded. Drift claims and status transitions are held and drain in order at 21:25.
The register itself is live and complete. The recurrence panel, the Pareto and the moving average are marked gate-blocked rather than rendered empty, because 61 days of data would produce a chart that looks like an answer and is not one. They unlock on 27-09, and the panel says so instead of showing a plausible shape.