Platform status

Kilawatt Cloud system status

No active faults — not enough customer traffic to rate every component

Derived from real payment, provisioning, execution and health records over the last 24 hours. Payments, Provisioning and API include our own live verification runs through the production API; open health incidents count customer machines only.

Checked Wed, 30 Sep 2026 05:39:15 GMT

Components

Payments (x402)

4 of 4 payments succeeded in the last 24 hours. Too few to rate reliably.

Basis: Real payments through the production API in the last 24 hours. Includes 4 of our own live verification runs. A rating needs at least 20 attempts.

Not enough data yet

Provisioning

10 of 17 provisioning requests succeeded in the last 24 hours. Too few to rate reliably.

Basis: Real provisioning requests through the production API in the last 24 hours. Includes 17 of our own live verification runs. A rating needs at least 20 attempts.

Not enough data yet

API

10 of 17 API executions succeeded in the last 24 hours. Too few to rate reliably.

Basis: Real API executions through the production API in the last 24 hours. Includes 17 of our own live verification runs. A rating needs at least 20 attempts.

Not enough data yet

GPU Health Monitoring

No open health incidents. Last health sample recorded 238 minutes ago.

Basis: Open incident windows on customer machines and the most recent health sample.

Operational

Incident history

Newest first. These are real incidents from this platform, including ones we caused ourselves. Each entry states what happened, why, what changed, and whether the work is genuinely finished.

2026-09-27Resolved

Live failover proof — real supplier outage handled without job loss, all four suppliers verified in isolation

What happened
Failover had been verified against test logic but never against a genuine, unplanned supplier failure. Separately, per-supplier provisioning had never been checked one supplier at a time, in isolation, to rule out one supplier's failure masking another's.
Root cause
No incident had yet occurred, under the tightened two-signal failover rule, in which a real (not simulated) machine died mid-job. And prior multi-supplier tests ran suppliers concurrently, which could hide a single supplier's problem behind a successful failover to a different one.
What was fixed
None needed — this was a verification exercise, not a bug fix. A real production job (f3c8c3c9) was launched on Kilawatt infrastructure; the underlying machine failed on its own. The platform's two-signal rule confirmed the outage at 23:22:01 UTC, dispatched a capped-loss replacement to Kilawatt infrastructure immediately, and the job completed successfully rather than failing. Separately, each of the four suppliers (Kilawatt infrastructure, Kilawatt infrastructure, Kilawatt infrastructure, Kilawatt infrastructure) was dispatched and observed in isolation, one at a time, with no other supplier running concurrently.
Current status
Verified live and documented in a standalone report ("MCP Agent Layer — Live Multi-Supplier Dispatch & Capped-Loss Failover Verification", published Sep 27, 2026): Kilawatt infrastructure, Kilawatt infrastructure and Kilawatt infrastructureeach reached a healthy state in 7-8 minutes when tested alone; Kilawatt infrastructure did not reach healthy within the test window when isolated, and its capped-loss failover during a live outage is the proof case above. This closes the previously open "live natural-failure confirmation" item on the failover fix and gives the first isolated (non-concurrent) per-supplier comparison.
2026-09-27Resolved

Resolved — Customers could trigger downtime credits without a real fault

What happened
A customer's launch response included their machine's health-report token. With it, a customer could report GPU faults that never happened, and the system would mark the machine as down and pay automatic downtime (SLA) credits.
Root cause
The health-report endpoint trusted the token alone, and the token was given to the customer instead of only to the machine. Credits for faults seen only inside the machine didn't need any confirmation from the compute supplier.
What was fixed
The token is no longer returned to customers and uses its own dedicated secret. Stopped machines can't send reports. Faults reported only from inside a machine now earn credit only when the supplier's own monitoring also saw the machine unhealthy at that time, and the report endpoint is rate-limited.
Current status
Resolved and verified live on production on Sep 27, 2026. A fake fault report sent with a real machine token was accepted but paid no credit, while a fault confirmed by the supplier still paid correctly. We found no sign it was used by any customer; the only credits ever paid through it were from our own tests, and these were reversed.
2026-09-27Resolved

No limit on repeated launches that never became usable

What happened
After the never-ready full refund was introduced, an account could repeatedly launch jobs that never became usable and be refunded each time, while the platform still paid its compute suppliers for each provisioning attempt.
Root cause
Launch checks covered balance, spend caps, request rate and concurrency, but nothing tracked how many of an account's recent launches ended without ever becoming ready, and there was no platform-wide cost breaker.
What was fixed
Each account (including pay-per-request wallets) is now limited to 3 never-ready launches per rolling 24 hours, after which new launches are refused before any funds are held. New launches also pause platform-wide if supplier cost on never-ready machines exceeds a set hourly threshold.
Current status
Resolved and verified live on production. The test launches all became ready, so the account was given test never-ready history; a real launch from that account was then refused before any funds were held, while a second account was still allowed.
2026-09-27Resolved

Resolved — x402 credit withdrawal

What happened
When a job paid via on-chain pay-per-request (x402) never became ready, the automatic full refund was credited to an internal account balance tied to the paying wallet, with no way to withdraw it back on-chain.
Root cause
x402 revenue settles to fiat, so there was no company-held USDC wallet or payout path for the internal balance a refund lands in.
What was fixed
Internal credit on an x402 wallet, including never-ready refunds, can now be returned as real USDC on Base to the wallet that paid. Payouts are approved by an operator, not automatic: they are sanctions-screened, capped at $50 each, sent from a company-held wallet, and logged.
Current status
The first live payout (2.81 USDC) is on-chain in Base transaction 0x900dc485…c937. Self-service withdrawals are not offered yet.
2026-09-27Resolved

Billing double-credit possible between the never-ready refund and the downtime-credit system

What happened
A job that received the new never-ready full refund could also separately receive a downtime/SLA credit for the same period, paying the customer twice for one failure.
Root cause
The two credit systems did not check each other's state: the downtime-credit path had no knowledge of whether the never-ready refund had already been paid for the same job.
What was fixed
The downtime-credit system now skips crediting any job that already received the never-ready refund.
Current status
Resolved. Caught and fixed during live testing of the never-ready refund the same day, and re-verified live: the refund posts exactly once and no second credit follows.
2026-09-27Resolved

Jobs that never became usable were not automatically refunded

What happened
If a job ended — timed out, failed, or was stopped — without ever reaching a confirmed-healthy state, the customer was still charged in full with no refund.
Root cause
No automated check existed linking a job's outcome to refund eligibility; ending a job simply kept whatever had been charged.
What was fixed
An automatic full refund now triggers whenever a job ends without ever confirming readiness, across every end-of-life path (completion, stop, failure, reaper).
Current status
Resolved and verified live on production with a real refund posted to a real test account's wallet.
2026-09-27Resolved

Failover could trigger on a single transient signal

What happened
A single failed status check from a provider — for example a transient rate-limit or server error — could have triggered failover of a genuinely healthy running job.
Root cause
The failover threshold did not require sustained or repeated negative evidence; one bad response was enough.
What was fixed
Failover now requires two consecutive negative signals with a minimum time gap, or a sustained failure window. A single blip no longer triggers it.
Current status
Resolved and now confirmed under a real failure: on Sep 27 a live Kilawatt infrastructure machine died unexpectedly mid-job (job f3c8c3c9); the two-signal rule caught it, failover replaced the job on Kilawatt infrastructure, and the job completed successfully rather than failing. See the capped-loss failover verification entry below for the full timeline.
2026-09-27Resolved

Jobs reported as running before any health confirmation

What happened
Jobs were reported as running and billed the moment a GPU provider accepted the request, before any health check confirmed the machine was actually usable.
Root cause
No readiness gate existed between provider acceptance and marking a job live; provider acceptance was treated as proof of a working machine.
What was fixed
Jobs now start in a provisioning state and only transition to healthy/running once a genuine readiness signal is confirmed — a provider status check or an in-machine health report.
Current status
Resolved and verified live on production.
2026-09-27Resolved

Signed-in customers could directly modify their own machine records

What happened
Signed-in customers had direct database write access to their own machine records — including billing rate and status fields — instead of those changes being restricted to server-side logic only.
Root cause
The table's access rules granted broad write permissions to any signed-in user for their own rows. An earlier change had limited what customers could read on those records but never touched the write permissions.
What was fixed
Customer write access to machine records was removed entirely; only server-side logic can create or change them now. The one console flow that previously wrote with the customer's own access was moved server-side, with ownership pinned to the signed-in customer.
Current status
Resolved and verified live against production with a real signed-in customer account: attempts to change the billing rate, change the status, delete a record, and insert a new one were all refused, and the customer's normal machine list still loads.
2026-09-27Resolved

Signed-in customers could read supplier-identifying fields directly from the database

What happened
The public API masks which underlying GPU supplier handles a job, but signed-in customers could bypass that masking by reading supplier-identifying fields — supplier name, supplier machine ID, and raw supplier error text — directly from the database for their own jobs and machines.
Root cause
Masking was applied in the API layer, but the underlying database rows were also directly readable by the owning customer, and those rows carried the unmasked fields.
What was fixed
Column-level access restrictions now block customers from reading supplier-identifying fields on their own records, while normal customer-visible fields (status, GPU, cost, credits) remain available. Admin and server-side access is unchanged.
Current status
Resolved and independently re-verified live against production with a real signed-in customer account: every attempt to read the restricted fields was refused, and the API's masked responses were re-confirmed to contain no supplier names or supplier machine IDs.
2026-09-22Resolved

Sanctions screening: stale-list fail-open window closed

What happened
After the sanctions refresh was moved to an hourly background job, a silently stalled refresh could keep serving the cached list for up to seven days and still allow a payment through. Cold start was safe (no list blocks), but a list that had gone stale between refreshes was trusted for far too long.
Root cause
MAX_STALENESS_MS was set to 7 days. That ceiling was meant as a backstop, but because the hourly job could fail repeatedly without surfacing, a week-old list could still read as usable and let a payment pass. The payment path fails closed on a missing or unusable list, but the definition of 'unusable' was too permissive.
What was fixed
Found by a technically precise public comment on the x402 postmortem that asked specifically whether a cold start or a stalled refresh fails open or closed. The ceiling was tightened from 7 days to 48 hours — a healthy hourly cache is never more than an hour old, so 48 hours tolerates a full day of missed refreshes and then blocks. A real OFAC-listed address was re-confirmed still blocked, and a clean address still passes.
Current status
Resolved and verified live. Before the fix, backdating the cache to 3 days old returned 'allowed'; after the fix the same stale cache returns 'blocked'. The sanctioned address stays blocked and the clean address still passes.
2026-09-22Resolved

API documentation accuracy: wrong host, inflated GPU limit, missing A100 price

What happened
Three documentation defects found in one review. The docs told developers to send requests to api.kilawattcloud.dev, a host that never existed and does not resolve. The GPU 'count' parameter was documented as accepting 1–512. And the homepage named the A100 as available while the rate card carried no A100 price at all.
Root cause
The public documentation and rate card were written ahead of the API and were never reconciled against what the gateway actually enforces. No customer request was affected by a system fault — the failure was that anyone following the docs literally would have hit an unresolvable host or a rejected request.
What was fixed
The docs, API explorer, quick-start and console snippets now point at the real host, www.kilawattcloud.dev/api/public/v1/, with the four endpoints that actually exist. The documented GPU count was corrected to the limits the gateway really enforces: 1–8 for H100, H200, A100 and L40S, 1–4 for B200, and up to 16 for workstation-class cards. The A100 was confirmed genuinely quotable across all four providers by a live quote and added to the rate card as a live-quoted card with a conservative floor price, since real cost ranges from about $0.54 to $15.92 per GPU-hour depending on provider.
Current status
Resolved, verified live and published the same day. A real API call to the corrected endpoint returned a priced, pre-authorized response; a request for 512 GPUs was rejected by the live gateway, as were 64 and 9, while 8 succeeded; and a live A100 quote returned real cost from all four providers.
2026-09-22Resolved

SLA-linked health monitoring and automatic credits

What happened
Until this change, a GPU that degraded or died mid-job was only caught by the customer. Hardware was verified at creation time and never again for the rest of the run.
Root cause
No continuous in-flight health checking existed, and any credit for lost time would have required the customer to notice and ask for it.
What was fixed
Every running instance is now probed on a short interval. Degraded and down windows are recorded with exact start and end times, and the proportional cost is credited back to the wallet automatically on recovery — no claim, no support ticket. The credit and the exact window appear on the job receipt.
Current status
Resolved (with caveat) — SLA-linked health monitoring and automatic credits. The in-machine health reporting is proven live end-to-end. On Sep 27 a simulated fault run went through the real system on a production machine that had been confirmed ready by its own health reports: a GPU fault report (Xid + ECC) was sent from outside the machine to the real reporting endpoint, using the machine's own credentials. The fault was simulated, and no real hardware fault occurred. The real system opened incidents, marked the machine down, logged the report, closed the incidents on the next clean report, and posted automatic SLA credits. The test found a bug where a milder incident could hide a more serious one; it was fixed the same day. It also found that overlapping incidents could each post their own credit; that was fixed the same day too. Not yet proven: catching a real hardware fault as it happens, which can't be forced without damaging rented equipment.
2026-09-21Resolved

Jobs requested as A40/A10 were fulfilled with an RTX 4090

What happened
Jobs submitted for an A40 or A10 were provisioned on an RTX 4090 instead, without the request failing and without the substitution being disclosed.
Root cause
Ours, not the provider's. Our own GPU mapping carried a fallback allow-list that included the RTX 4090 and was missing proper A40 and A10 entries, so an unmatched request quietly fell through to whatever the fallback listed.
What was fixed
The fallback was removed. Each GPU now maps only to true equivalents, an unavailable card is refused outright instead of substituted, and the hardware actually provisioned is verified against the request before the charge is taken. A mismatch is torn down and failed over rather than billed. Receipts state the GPU that was actually provisioned.
Current status
Resolved and verified with real paid test jobs: an A40 request provisioned as an A40, and an A10 request that could not be honoured was refused with the funds returned to wallet credit instead of being silently swapped.
2026-09-15Resolved

x402 payment authorization windows expiring (120–165+ second settlements)

What happened
Some agent payments took 120 to 165 seconds or longer to settle and exceeded the payment authorization window, so the job was never created. No customer was charged for a failed settlement.
Root cause
A stale OFAC sanctions-list cache. When the cache aged out, the payment path itself triggered a live, untimed re-download of the list from the US Treasury and blocked on it. Slow or failing Treasury responses stalled the payment, not the payment system.
What was fixed
The sanctions refresh was moved out of the payment path entirely and onto an hourly database-side job with timeouts and backoff. Screening still fails closed against a locally stored list. The authorization window was widened as a safety margin rather than as the fix.
Current status
Resolved. A live re-test after the change settled in 18.8 seconds.
2026-09-15Resolved

Billing race condition: payments could settle without a matching job

What happened
Under simultaneous load, six concurrent execution requests could settle six payments while only three jobs were actually created — money taken with no compute behind it.
Root cause
The balance check and the reservation were separate steps, so concurrent requests each read the same balance before any of them wrote. Settlement also happened before the limit and profitability checks that could still reject the job.
What was fixed
Reservations are now atomic in the database: check and hold happen in one operation, so concurrent requests cannot both win the same funds. Rate, concurrency, spend and margin checks all run before any payment is settled, and a reservation that does not become a job is released back to the wallet.
Current status
Resolved and verified with a live six-request burst: exactly three jobs were created, three requests were rejected, and the rejected ones left the balance untouched.
2026-09-15Resolved

x402 payments settling before safety checks ran

What happened
On the x402 execution route, a payment could be settled before the rate-limit, concurrency and spend-cap checks ran at all — so a request that was going to be rejected on limits could still take the money first and refund it after.
Root cause
The route settled the payment early in the request flow and only afterwards evaluated whether the job was allowed to run under the caller's limits.
What was fixed
The full pre-authorization dry-run (balance, rate limit, concurrency, spend cap) was moved ahead of payment settlement, so nothing is charged for a request the checks would reject. Verified against a real settled transaction.
Current status
Resolved. Later the same day the reservation itself was also made atomic under concurrency — see the separate entry above; that was the deeper fix for the same class of problem, found and fixed at a different time.
2026-09-15Resolved

Live multi-agent run: 18 real jobs held through a full Kilawatt infrastructure outage

What happened
An actual registered agent wallet ran 18 real completed jobs through the public x402 endpoint, settling real USDC across 22 payments, with zero jobs dropped. During the run, Kilawatt infrastructure failed 100% of the attempts made to it that day (11 of 11).
Root cause
Nothing to fix on our side — this was a live, unplanned provider outage, not a simulation. It was the first real-world test of the failover system under load.
What was fixed
No code change was needed. The routing layer automatically failed every failed Kilawatt infrastructure attempt over to Kilawatt infrastructure — a provider that had never previously appeared as a live failover destination — and to Kilawatt infrastructure, and every job still completed and settled correctly.
Current status
Resolved and proven. Our records for that evening show the job burst and the Kilawatt infrastructure-to-Kilawatt infrastructure/Kilawatt infrastructure failovers with real settled payments; none of the 18 jobs was lost.
2026-09-03Resolved

Original billing and provisioning safety overhaul: seven gaps closed together

What happened
A launch-safety review found seven real gaps at once: launch paths only checked the first hour of cost instead of the full job, there were no per-key rate or concurrency limits, customers had no spend cap of their own, running instances were never metered or reaped, Kilawatt infrastructure provisioning was hitting the wrong endpoint and returning 404s, provider failover did not actually fall through from Kilawatt infrastructure to Kilawatt infrastructure on failure, and gateway-launched instances were not recorded or auto-terminated at their paid-through time.
Root cause
The platform had been built feature-first: jobs launched and providers were called, but the safety rails around money and lifecycle had never been systematically checked.
What was fixed
All seven were fixed in one pass: full-cost pre-authorization on every launch path, default per-key limits of 3 concurrent jobs and 30 requests/minute (locked against self-elevation), an optional customer-set monthly spend cap, a 15-minute metering and auto-reap scheduler, the corrected Kilawatt infrastructure endpoint, a real Kilawatt infrastructure-to-Kilawatt infrastructure failover path, and recording plus paid-through termination for gateway-launched instances.
Current status
Resolved and verified live with real instance creation and destruction on both Kilawatt infrastructure and Kilawatt infrastructure.

Includes our own live verification runs through the production API.

Component states are computed from our control-plane records at page load; they are not a synthetic external probe of every provider. Questions: hello@kilawattcloud.dev