Projects / paused

Inference gateway

A nine-day project whose GPU-leasing capability was deleted five days in, and whose safety doctrine outlived it.

Status
paused
Languages
C#
Current release
—
Latest
—

Rate limit Bounds the rate

Worst-case accrual Bounds the total

A per-hour price cap does not bound spend over time. The design commits worst-case accrual at creation and persists it immediately.

How it works

A gateway with a bounded-concurrency path, an inference proxy, and (formerly) a GPU fleet autoscaler with a lease store. The autoscaler was the only writer of leases and the only reader of load reports.

An inference gateway extracted from another project, with a GPU fleet autoscaler. The fleet capability was deleted five days later — 70 files and 6,801 lines — because three of its seams were unreachable by construction. What survived is a written doctrine about the cost of renting GPUs that is more clearly argued than the code it was built for.

  • 70 files and 6,801 lines deleted in one commit, five days after extraction.
  • Three seams removed because they were unreachable by construction, not because they were wrong.
  • 'MaxFleetPricePerHour is a rate, and a rate bounds nothing over time. Timeouts are not money limits.'

Releases