Apache-2.0 · one binary · v0.2.0

Kill it mid-batch. Resume finishes the job.

forge fans a giant JSONL of prompts across N OpenAI-compatible engines — vLLM, SGLang, llama.cpp, Ollama, or a hosted API — on your own spot GPUs. The GPU work stays in the engines. forge is only the orchestration shell: work queue, sharding, retries, spot-interruption handling, backpressure, result aggregation.

§01 — the claim, tested

A hard kill costs eight re-runs. A signal costs 601 ms.

One 300-item batch at concurrency 8, ended two ways: kill -9, the way a spot instance dies with no warning, and SIGTERM, the way it dies with about thirty seconds of one. Both finish 300 of 300. What differs is the bill. No output file to join back against the input on keys, no bespoke checkpoint recipe: the lease queue is the resume state, so it answers the question directly.

kill -9 · no warning

at the kill

8 live leases

108 done, 184 never started. The in-flight items are still leased to a process that no longer exists.

after the lease TTL

8 orphaned

Those leases lapse and the items become interrupted: true — the fingerprint of a machine that died.

reclaimable

192

184 pending plus 8 orphaned. Exactly the set resume re-dispatches, join-free.

after resume

300/300

retried=8. The engine served 308 requests for 300 items: the eight interrupted ones ran twice and were stored once.

$ forge audit --checkpoint ckpt.db pending=184 live_leases=8 orphaned_leases=0 done=108 dead=0 total=300 reclaimable=184 (184 pending + 0 orphaned) live workers still holding leases: forge-17244=8 # …once the lease TTL lapses, the dead worker's 8 items join the reclaim set $ forge resume --checkpoint ckpt.db --workers http://127.0.0.1:8099 --concurrency 8 $ forge verify --input in.jsonl --results out.jsonl exit 0 # 300 unique outputs, 0 missing, 0 extra, 0 duplicates

SIGTERM · draining

at the signal

601 ms to drain

It stops leasing new work, lets what is already in flight finish and store its results, then exits. 91 results written.

live leases

0

Nothing was still in flight, so nothing holds a lease. Exit code 2: items never leased are still pending, which is what an orchestrator branches on.

orphaned

0

There is no TTL to wait out. The checkpoint is resumable the moment the process is gone, rather than a lease-length later.

after resume

300/300

retried=0. Nothing had to be reclaimed, and no GPU time was paid for twice.

# SIGTERM: stop leasing, let the in-flight finish, store their results, exit $ forge run --input in.jsonl --workers http://127.0.0.1:8099 --concurrency 8 exit 2 # Incomplete — drained in 601 ms, 91 results written $ forge audit --checkpoint ckpt.db live_leases=0 orphaned_leases=0 # nothing to wait out, nothing to reclaim $ forge resume --checkpoint ckpt.db --workers http://127.0.0.1:8099 --concurrency 8 $ forge verify --input in.jsonl --results out.jsonl exit 0 # 300 unique outputs, 0 missing, 0 extra, 0 duplicates

Why that is safe rather than lucky. A lease is fenced by (custom_id, attempt), so a result posted for a stale attempt by a resurrected worker is rejected. An item cannot commit twice. An item is only done once its result is durably stored, so a crash between store and acknowledge re-runs an idempotent store rather than losing the work. resume re-dispatches exactly pending + orphaned; done rows are never re-leased.

Being hard-killed was always correct. It was never free. Every item in flight is GPU time already spent that has to be spent again, and none of it can be reclaimed until the lease TTL elapses. A spot preemption gives about thirty seconds of warning, which is far more than the 601 ms this drain took, so on the interruption you actually get that loss becomes nothing.

A second signal does not escalate. A drain finishes only what is already in flight, so an impatient second Ctrl-C would throw away precisely the work the first one was protecting. The escape hatch is SIGKILL, and it stays safe for the reason the top of this section shows: a hard kill costs GPU minutes, never correctness.

§02 — heterogeneous by default

A spot fleet is never made of identical boxes

Dispatch used to be wave-based: lease a batch, dispatch it, wait for all of it, lease the next. That puts a barrier at every batch boundary — so one slow worker gates the whole fleet.

400-item batch, concurrency 8 — sliding window

1 fast box 12.99/s
2 fast boxes 25.76/s
2 fast + 1 weak 25.59/s
the fleet’s ceiling
if nothing were wasted
~30/s

Under wave dispatch the third row falls below the first: adding a third machine made the fleet slower than one box on its own. That is the head-of-line block, measured, and it is forge's own regression rather than a competitor's. Dispatch is now a sliding window with per-worker free-slot accounting: every completion immediately frees a slot topped up with a fresh lease, and a saturated worker is skipped rather than queued behind.

fleetwave dispatchsliding windowgain
1 fast box10.86 items/s12.99 items/s+20%
2 fast boxes21.08 items/s25.76 items/s+22%
2 fast + 1 weak9.15 items/s25.59 items/s2.8×

In the mixed run the weak box chews its own 49 items at its own pace while the fast boxes stay saturated — zero 429s, engine-side peak queue depth of 4. Throughput is the sum of what each engine can do rather than a multiple of the slowest. A naive sequential client against the same single box manages 1.58 items/s.

§03 — when you get the numbers wrong

Declare 64 against a cap of 8 and it still finishes

Misdeclare a worker's concurrency eight times over the engine's real limit. A naive concurrent client permanently loses work; forge converges on the actual cap.

client — 200 items, declared concurrency 64completedpermanently lostengine 429s
naive 64-thread, fixed 1 s retry, 10 attempts158/200421,107
forge — AIMD + header cooldown200/200072

AIMD halves the in-flight window on each 429 burst, and a shared per-worker cooldown honours Retry-After so the whole fleet backs off in lockstep rather than each thread discovering the limit alone. That is 94% less rejected traffic, and nothing dead-lettered.

The coordinator is nearly free

Saturating three remote worker nodes over real VPC HTTP at 26.89 items/s, forge itself peaked at 9.7 MB RSS and 0.9% of one core (2% peak). It adds scheduling, durability and accounting without meaningfully taxing the box it runs on.

Against a sequential client

16.1× on that remote fleet, and about 90% of the fleet's theoretical 30/s ceiling — with no tuning. Every run: all items done, zero dead-letters.

What these numbers are, and are not They are measured against examples/engine_sim.py — a mock that models a hard concurrency cap, queueing beyond it, 429 with Retry-After, and per-request jitter — not against real GPUs. That is deliberate: the harness is the variable under test, not GPU throughput, which lives in vLLM and SGLang and is out of forge's scope. These figures tell you what the coordinator does with a fleet. They tell you nothing about how fast your model runs.

§04 — the arbitrage, priced by you

It will also tell you the batch wasn’t worth a fleet

Offline batch inference is roughly five to ten times cheaper per token than online serving, and spot GPUs cut another 60–90% off on-demand. forge cost turns the real usage forge captured into the invoiceable difference — and it never guesses a price.

# a 50M-item summarization run — ~35B input / ~7.5B output tokens — # the rent you actually paid: rate x wall-clock hours x GPUs, # priced against an online rate you also supply $ forge cost --results results.jsonl \ --gpu-usd-per-hour 5.00 --gpu-hours 50 --gpus 8 \ --online-per-mtok-input 0.50 --online-per-mtok-output 1.50 forge: $2000.0000 ($0.0471/Mtok · 21250000 tokens/$) online: $28750.0000 → saved $26750.0000 (93.0%)

Both sides of that sum are yours. forge does not shop for GPUs and does not know what spot costs today — it multiplies the rate you give it by the hours the run took, and prices the same tokens against the online rate you also give it. Pass --gpu-cost instead if you already have one figure for the run. What forge contributes is the token side: those totals are summed from the real usage the engines reported, not estimated from your input file.

It is honest in the other direction too: on a job too small to amortise the GPU bill the saved figure goes negative, which tells you the batch was not worth a dedicated fleet. Cached-prompt tokens price at the full input rate unless you pass a discounted rate for them — conservative by default, so the saving is never overstated.

Against the alternative

Cloud Batch APIs hide the operational pain but give a flat ~50% discount, with no choice of model weights, data locality or GPU economics. forge is the other trade: you keep all three, and you own the procurement decision — forge runs the shell, it does not buy your GPUs.

Keep the prefix cache hot

--prefix-bucket reorders the input by shared (model, system-prompt) so the engine's automatic prefix cache actually stays warm and those cached_tokens materialise instead of being theoretical.

The only numbers on this page that are not measured The GPU rate, the hours and the online prices above are an illustration — plausible figures put through the real arithmetic, not a run anybody performed. Everything else here comes from a harness in the repository. Substitute your own four numbers and the answer changes completely, which is the entire point of the command.

§05 — the boundary

One flat fan-out. No graph, no provisioner.

forge is a single homogeneous fan-out of independent items across interchangeable workers. The boundary is enforced by the data model rather than by discipline — there is nowhere in the core to express what a workflow engine exists to express.

Not a workflow engine

No task graph, no depends_on or fan-in, no cron or triggers, no YAML topology. Arbitrary task graphs with inter-task dependencies are dagron's job. forge can hand a flat plan to dagron; the arrow points outward only.

Not a provisioner

“Get me 400 spot GPUs across 12 regions” is out of scope. forge consumes interruption signals and re-queues at item granularity. It never launches, tears down or autoscales an instance.

Not an inference engine

The GPU heavy-lifting stays in vLLM, SGLang, llama.cpp or Ollama, where it belongs. forge never batches on the wire either — the engine's continuous batching packs concurrent requests.

Getting it

A multi-arch image, or build the static binary from source. The v0.2.0 GitHub release carries no attached binaries yet, and images are not yet signed.

$ docker pull mancube/forge:v0.2.0 $ cargo build --release -p forge-cli

Results can land in object storage (S3, GCS, Azure) rather than a local file, and forge can serve the OpenAI Batch REST API surface so existing clients point at it unchanged.

§06 — but isn’t this just…

Six things forge gets mistaken for

The slot forge sits in is narrow, and it is next to four much larger ones. Every answer below is about what kind of thing forge is, not about who is faster — no number here compares forge to anything but a naive client, because that is the only comparison this repository actually ran.

Isn’t this a cluster framework?

The standard self-hosted answer to batch inference is a distributed runtime: a cluster, actors, placement groups, an operator to install. It is mature and genuinely powerful, and forge does not beat it at inference — the GPU work lives in the engine either way.

What forge trades away is the cluster. One static binary, no control plane, pointed at OpenAI-compatible endpoints wherever they are: a laptop, a bare VM, a spot fleet with no Kubernetes on it. That is a cold-start and operational-weight argument, not a throughput one.

The honest half: resume and engine-failure recovery are actively landing in those stacks. “They cannot resume” is a shrinking difference and we do not lean on it. “No cluster required” is the durable one.

Isn’t this a provisioner?

No, and the two compose rather than compete. Getting you 400 spot GPUs across a dozen regions is a launcher's job. A launcher is not a per-item work queue: it hands you machines, not a durable record of which of your 50 million prompts has been answered.

forge consumes interruption signals and re-queues at item granularity. It never launches, tears down or autoscales an instance — §05 states that as a boundary rather than a roadmap gap.

Isn’t this a workflow engine?

There is no task graph, no depends_on or fan-in, no cron and no YAML topology. forge is one flat fan-out of independent items across interchangeable workers, and the boundary is enforced by the data model: there is nowhere in the core to express what a workflow engine exists to express.

Arbitrary graphs are dagron's job. forge can hand a flat plan to a workflow engine; the arrow points outward only.

Why not just use a hosted batch API?

Often you should. They are zero-infrastructure and the ergonomics are hard to beat, and they set the contract forge deliberately adopts: a JSONL of independent requests each keyed by a unique custom_id, output keyed by the same id with order not guaranteed. Migration off one is a file copy for exactly that reason.

What you give up is the choice of model weights, data locality and the GPU economics underneath — and the discount floors at a flat rate somebody else sets. forge is the other trade: you keep all three and you own the procurement decision. §04 prices it with your numbers, including the case where the answer is that the batch was not worth a fleet.

Why not rent the GPUs and write the script myself?

You can, and for a one-off you probably should. Compute platforms and serving frameworks sell you the machines or the engine; the queue, the sharding, the checkpoint and the aggregation are still yours to build. That shell is forge's entire scope.

The page measures the naive version rather than asserting it is bad. A sequential client against one box manages 1.58 items/s. A 64-thread one that misdeclares the engine's concurrency finishes 158 of 200 and loses 42 permanently, with 1,107 rejections — §03 has the run. Fencing, lease expiry and backpressure are the parts that are not a weekend.

Does forge make my model run faster?

No. It never batches on the wire — the engine's own continuous batching packs concurrent requests, and the GPU heavy-lifting stays in vLLM, SGLang, llama.cpp or Ollama where it belongs.

Every figure on this page is about what the coordinator does with a fleet: how much of its ceiling it reaches, what it loses when a box dies, what it does when you get the numbers wrong. They tell you nothing about how fast your model runs, and §03 says so in its own callout.

§07 — the long run

A long job should not look the same as a wedged one

A batch that takes six hours is a different operational object from one that takes six seconds. You need to know it is alive, know what it cost you when something goes wrong, and be able to take a backup without stopping it. That was the work of v0.2.0.

You can tell it apart from a hang

A progress heartbeat every 30 seconds carries done/total, how many are in flight, items per second and an ETA. Before it, a multi-hour run and a wedged one logged exactly the same thing: nothing.

Dead letters have a return leg

forge requeue moves quarantined items back to pending with a fresh retry budget, on the same job: same checkpoint, same results file, no second run to merge. Select with --ids, --reason or --all; there is no default, because reversing a terminal state should not guess at scope.

The prune is the feature

The store rebuilds its emitted-id set from the results file and the dead-letter sidecar, so an item requeued while its dead record still stood would be re-run on a real GPU and have its fresh result silently dropped. requeue removes those records first, atomically, and refuses outright while a coordinator still holds a live lease, naming the holders.

A backup you can take while it runs

The queue snapshots through SQLite's VACUUM INTO: a self-contained file holding exactly the state committed when it started, without blocking the writer. It is an API rather than tar because the queue is WAL-mode: committed state lives across three files, and an archiver reading them at three different moments gets a torn copy that opens cleanly and is simply short.

Something to quote at support

Version, git SHA and build target on startup and on /v1/health. An inbound x-request-id is echoed on every response, not only errors, because “quote the request id” is useless advice if a successful-but-wrong response has none. --log-format json when a shipper is reading.

stdout is the answer

Diagnostics go to stderr. They used to share stdout with the payload, so a single rejected-input warning could break forge status --json for whoever hit it. There is a test.

The checkpoint carries a schema version now, so a file written by a newer forge is refused by name with the remedy rather than half-read. A checkpoint from before the field existed is stamped rather than migrated, because nothing about its tables changed, so a job in progress is not broken by the introduction of a version field. The protection is forward.

Apache-2.0

Every number here is reproducible from the repository

The kill, the drain, the fleet comparison and the overload run all ship as scripts and mocks you can run in three terminals. The regression that caused the head-of-line block is pinned by a test. A verification run on 2026-09-10 re-ran every claim above in one sitting and landed within about 5% of the figures shown, which is what sharing a box with a soak test costs.