- What happened
- Failover had been verified against test logic but never against a genuine, unplanned supplier failure. Separately, per-supplier provisioning had never been checked one supplier at a time, in isolation, to rule out one supplier's failure masking another's.
- Root cause
- No incident had yet occurred, under the tightened two-signal failover rule, in which a real (not simulated) machine died mid-job. And prior multi-supplier tests ran suppliers concurrently, which could hide a single supplier's problem behind a successful failover to a different one.
- What was fixed
- None needed — this was a verification exercise, not a bug fix. A real production job (f3c8c3c9) was launched on Kilawatt infrastructure; the underlying machine failed on its own. The platform's two-signal rule confirmed the outage at 23:22:01 UTC, dispatched a capped-loss replacement to Kilawatt infrastructure immediately, and the job completed successfully rather than failing. Separately, each of the four suppliers (Kilawatt infrastructure, Kilawatt infrastructure, Kilawatt infrastructure, Kilawatt infrastructure) was dispatched and observed in isolation, one at a time, with no other supplier running concurrently.
- Current status
- Verified live and documented in a standalone report ("MCP Agent Layer — Live Multi-Supplier Dispatch & Capped-Loss Failover Verification", published Sep 27, 2026): Kilawatt infrastructure, Kilawatt infrastructure and Kilawatt infrastructureeach reached a healthy state in 7-8 minutes when tested alone; Kilawatt infrastructure did not reach healthy within the test window when isolated, and its capped-loss failover during a live outage is the proof case above. This closes the previously open "live natural-failure confirmation" item on the failover fix and gives the first isolated (non-concurrent) per-supplier comparison.