forge fans a giant JSONL of prompts across N OpenAI-compatible engines
— vLLM, SGLang, llama.cpp, Ollama, or a hosted API — on your own
spot GPUs. The GPU work stays in the engines. forge is only the orchestration
shell: work queue, sharding, retries, spot-interruption handling,
backpressure, result aggregation.
A hard kill costs eight re-runs. A signal costs 601 ms.
One 300-item batch at concurrency 8, ended two ways: kill -9, the
way a spot instance dies with no warning, and SIGTERM, the way it
dies with about thirty seconds of one. Both finish 300 of 300. What differs is
the bill. No output file to join back against the input on keys, no bespoke
checkpoint recipe: the lease queue is the resume state, so it answers
the question directly.
kill -9 · no warning
at the kill
8 live leases
108 done, 184 never started. The in-flight items are still leased to a process that no longer exists.
after the lease TTL
8 orphaned
Those leases lapse and the items become interrupted: true — the fingerprint of a machine that died.
reclaimable
192
184 pending plus 8 orphaned. Exactly the set resume re-dispatches, join-free.
after resume
300/300
retried=8. The engine served 308 requests for 300 items: the eight interrupted ones ran twice and were stored once.
It stops leasing new work, lets what is already in flight finish and store its results, then exits. 91 results written.
live leases
0
Nothing was still in flight, so nothing holds a lease. Exit code 2: items never leased are still pending, which is what an orchestrator branches on.
orphaned
0
There is no TTL to wait out. The checkpoint is resumable the moment the process is gone, rather than a lease-length later.
after resume
300/300
retried=0. Nothing had to be reclaimed, and no GPU time was paid for twice.
# SIGTERM: stop leasing, let the in-flight finish, store their results, exit$ forge run --input in.jsonl --workers http://127.0.0.1:8099 --concurrency 8
exit 2# Incomplete — drained in 601 ms, 91 results written$ forge audit --checkpoint ckpt.db
live_leases=0 orphaned_leases=0 # nothing to wait out, nothing to reclaim$ forge resume --checkpoint ckpt.db --workers http://127.0.0.1:8099 --concurrency 8
$ forge verify --input in.jsonl --results out.jsonl
exit 0# 300 unique outputs, 0 missing, 0 extra, 0 duplicates
Why that is safe rather than lucky. A lease is fenced by
(custom_id, attempt), so a result posted for a stale attempt
by a resurrected worker is rejected. An item cannot commit twice.
An item is only done once its result is durably stored, so a
crash between store and acknowledge re-runs an idempotent store rather
than losing the work. resume re-dispatches exactly
pending + orphaned; done rows are never
re-leased.
Being hard-killed was always correct. It was never free. Every item
in flight is GPU time already spent that has to be spent again, and none of
it can be reclaimed until the lease TTL elapses. A spot preemption gives
about thirty seconds of warning, which is far more than the 601 ms this
drain took, so on the interruption you actually get that loss
becomes nothing.
A second signal does not escalate. A drain finishes only what is
already in flight, so an impatient second Ctrl-C would throw
away precisely the work the first one was protecting. The escape hatch is
SIGKILL, and it stays safe for the reason the top of this
section shows: a hard kill costs GPU minutes, never correctness.
§02 — heterogeneous by default
A spot fleet is never made of identical boxes
Dispatch used to be wave-based: lease a batch, dispatch it, wait for all of
it, lease the next. That puts a barrier at every batch boundary — so
one slow worker gates the whole fleet.
400-item batch, concurrency 8 — sliding window
1 fast box12.99/s
2 fast boxes25.76/s
2 fast + 1 weak25.59/s
the fleet’s ceiling if nothing were wasted~30/s
Under wave dispatch the third row falls below the
first:adding a third machine made the fleet slower than one box
on its own. That is the head-of-line block, measured, and it
is forge's own regression rather than a competitor's. Dispatch is now a
sliding window with per-worker free-slot accounting: every completion
immediately frees a slot topped up with a fresh lease, and a saturated worker
is skipped rather than queued behind.
fleet
wave dispatch
sliding window
gain
1 fast box
10.86 items/s
12.99 items/s
+20%
2 fast boxes
21.08 items/s
25.76 items/s
+22%
2 fast + 1 weak
9.15 items/s
25.59 items/s
2.8×
In the mixed run the weak box chews its own 49 items at its own pace while
the fast boxes stay saturated — zero 429s, engine-side peak queue
depth of 4. Throughput is the sum of what each engine can do rather
than a multiple of the slowest. A naive sequential client against the same
single box manages 1.58 items/s.
§03 — when you get the numbers wrong
Declare 64 against a cap of 8 and it still finishes
Misdeclare a worker's concurrency eight times over the engine's real limit.
A naive concurrent client permanently loses work; forge converges on the
actual cap.
client — 200 items, declared concurrency 64
completed
permanently lost
engine 429s
naive 64-thread, fixed 1 s retry, 10 attempts
158/200
42
1,107
forge — AIMD + header cooldown
200/200
0
72
AIMD halves the in-flight window on each 429 burst, and a shared per-worker
cooldown honours Retry-After so the whole fleet backs off in
lockstep rather than each thread discovering the limit alone. That is
94% less rejected traffic, and nothing dead-lettered.
The coordinator is nearly free
Saturating three remote worker nodes over real VPC HTTP at 26.89 items/s, forge itself peaked at 9.7 MB RSS and 0.9% of one core (2% peak). It adds scheduling, durability and accounting without meaningfully taxing the box it runs on.
Against a sequential client
16.1× on that remote fleet, and about 90% of the fleet's theoretical 30/s ceiling — with no tuning. Every run: all items done, zero dead-letters.
What these numbers are, and are not
They are measured against examples/engine_sim.py — a mock
that models a hard concurrency cap, queueing beyond it, 429 with
Retry-After, and per-request jitter — not against real
GPUs. That is deliberate: the harness is the variable under test, not GPU
throughput, which lives in vLLM and SGLang and is out of forge's scope.
These figures tell you what the coordinator does with a fleet. They tell you
nothing about how fast your model runs.
§04 — the arbitrage, priced by you
It will also tell you the batch wasn’t worth a fleet
Offline batch inference is roughly five to ten times cheaper per token than
online serving, and spot GPUs cut another 60–90% off on-demand.
forge cost turns the realusage forge
captured into the invoiceable difference — and it never guesses a
price.
# a 50M-item summarization run — ~35B input / ~7.5B output tokens —# the rent you actually paid: rate x wall-clock hours x GPUs,# priced against an online rate you also supply$ forge cost --results results.jsonl \
--gpu-usd-per-hour 5.00 --gpu-hours 50 --gpus 8 \
--online-per-mtok-input 0.50 --online-per-mtok-output 1.50
forge: $2000.0000 ($0.0471/Mtok · 21250000 tokens/$)
online: $28750.0000 → saved $26750.0000 (93.0%)
Both sides of that sum are yours. forge does not shop for GPUs and
does not know what spot costs today — it multiplies the rate you give
it by the hours the run took, and prices the same tokens against the online
rate you also give it. Pass --gpu-cost instead if you already
have one figure for the run. What forge contributes is the token
side: those totals are summed from the real usage the
engines reported, not estimated from your input file.
It is honest in the other direction too: on a job too small to amortise the
GPU bill the saved figure goes negative, which tells you the batch
was not worth a dedicated fleet. Cached-prompt tokens price at the full
input rate unless you pass a discounted rate for them — conservative
by default, so the saving is never overstated.
Against the alternative
Cloud Batch APIs hide the operational pain but give a flat ~50% discount, with no choice of model weights, data locality or GPU economics. forge is the other trade: you keep all three, and you own the procurement decision — forge runs the shell, it does not buy your GPUs.
Keep the prefix cache hot
--prefix-bucket reorders the input by shared (model, system-prompt) so the engine's automatic prefix cache actually stays warm and those cached_tokens materialise instead of being theoretical.
The only numbers on this page that are not measured
The GPU rate, the hours and the online prices above are an
illustration — plausible figures put through the real
arithmetic, not a run anybody performed. Everything else here comes from a
harness in the repository. Substitute your own four numbers and the answer
changes completely, which is the entire point of the command.
§05 — the boundary
One flat fan-out. No graph, no provisioner.
forge is a single homogeneous fan-out of independent items across
interchangeable workers. The boundary is enforced by the data model rather
than by discipline — there is nowhere in the core to express what a
workflow engine exists to express.
Not a workflow engine
No task graph, no depends_on or fan-in, no cron or triggers, no YAML topology. Arbitrary task graphs with inter-task dependencies are dagron's job. forge can hand a flat plan to dagron; the arrow points outward only.
Not a provisioner
“Get me 400 spot GPUs across 12 regions” is out of scope. forge consumes interruption signals and re-queues at item granularity. It never launches, tears down or autoscales an instance.
Not an inference engine
The GPU heavy-lifting stays in vLLM, SGLang, llama.cpp or Ollama, where it belongs. forge never batches on the wire either — the engine's continuous batching packs concurrent requests.
Getting it
A multi-arch image, or build the static binary from source. The v0.2.0 GitHub release carries no attached binaries yet, and images are not yet signed.
Results can land in object storage (S3, GCS, Azure) rather than a local
file, and forge can serve the OpenAI Batch REST API surface so existing
clients point at it unchanged.
§06 — but isn’t this just…
Six things forge gets mistaken for
The slot forge sits in is narrow, and it is next to four much larger ones.
Every answer below is about what kind of thing forge is, not about
who is faster — no number here compares forge to anything but a naive
client, because that is the only comparison this repository actually ran.
Isn’t this a cluster framework?
The standard self-hosted answer to batch inference is a distributed
runtime: a cluster, actors, placement groups, an operator to install.
It is mature and genuinely powerful, and forge does not beat it at
inference — the GPU work lives in the engine either way.
What forge trades away is the cluster. One static binary, no control
plane, pointed at OpenAI-compatible endpoints wherever they are: a
laptop, a bare VM, a spot fleet with no Kubernetes on it. That is a
cold-start and operational-weight argument, not a throughput one.
The honest half: resume and engine-failure recovery are
actively landing in those stacks. “They cannot resume” is a
shrinking difference and we do not lean on it. “No cluster
required” is the durable one.
Isn’t this a provisioner?
No, and the two compose rather than compete. Getting you 400 spot
GPUs across a dozen regions is a launcher's job. A launcher is not a
per-item work queue: it hands you machines, not a durable record of
which of your 50 million prompts has been answered.
forge consumes interruption signals and re-queues at item
granularity. It never launches, tears down or autoscales an instance
— §05 states that as a boundary rather
than a roadmap gap.
Isn’t this a workflow engine?
There is no task graph, no depends_on or fan-in, no cron
and no YAML topology. forge is one flat fan-out of independent items
across interchangeable workers, and the boundary is enforced by the data
model: there is nowhere in the core to express what a workflow engine
exists to express.
Arbitrary graphs are dagron's job.
forge can hand a flat plan to a workflow engine; the arrow
points outward only.
Why not just use a hosted batch API?
Often you should. They are zero-infrastructure and the ergonomics are
hard to beat, and they set the contract forge deliberately adopts: a
JSONL of independent requests each keyed by a unique
custom_id, output keyed by the same id with order not
guaranteed. Migration off one is a file copy for exactly that reason.
What you give up is the choice of model weights, data locality and
the GPU economics underneath — and the discount floors at a flat
rate somebody else sets. forge is the other trade: you keep all three
and you own the procurement decision. §04 prices
it with your numbers, including the case where the answer is that the
batch was not worth a fleet.
Why not rent the GPUs and write the script myself?
You can, and for a one-off you probably should. Compute platforms and
serving frameworks sell you the machines or the engine; the queue, the
sharding, the checkpoint and the aggregation are still yours to build.
That shell is forge's entire scope.
The page measures the naive version rather than asserting it is bad.
A sequential client against one box manages 1.58 items/s. A
64-thread one that misdeclares the engine's concurrency finishes
158 of 200 and loses 42 permanently, with 1,107 rejections
— §03 has the run. Fencing, lease
expiry and backpressure are the parts that are not a weekend.
Does forge make my model run faster?
No. It never batches on the wire — the engine's own continuous
batching packs concurrent requests, and the GPU heavy-lifting stays in
vLLM, SGLang, llama.cpp or Ollama where it belongs.
Every figure on this page is about what the coordinator does with a
fleet: how much of its ceiling it reaches, what it loses when a box
dies, what it does when you get the numbers wrong. They tell you nothing
about how fast your model runs, and §03 says
so in its own callout.
§07 — the long run
A long job should not look the same as a wedged one
A batch that takes six hours is a different operational object from one that
takes six seconds. You need to know it is alive, know what it cost you when
something goes wrong, and be able to take a backup without stopping it. That
was the work of v0.2.0.
You can tell it apart from a hang
A progress heartbeat every 30 seconds carries done/total, how many are in flight, items per second and an ETA. Before it, a multi-hour run and a wedged one logged exactly the same thing: nothing.
Dead letters have a return leg
forge requeue moves quarantined items back to pending with a fresh retry budget, on the same job: same checkpoint, same results file, no second run to merge. Select with --ids, --reason or --all; there is no default, because reversing a terminal state should not guess at scope.
The prune is the feature
The store rebuilds its emitted-id set from the results file and the dead-letter sidecar, so an item requeued while its dead record still stood would be re-run on a real GPU and have its fresh result silently dropped. requeue removes those records first, atomically, and refuses outright while a coordinator still holds a live lease, naming the holders.
A backup you can take while it runs
The queue snapshots through SQLite's VACUUM INTO: a self-contained file holding exactly the state committed when it started, without blocking the writer. It is an API rather than tar because the queue is WAL-mode: committed state lives across three files, and an archiver reading them at three different moments gets a torn copy that opens cleanly and is simply short.
Something to quote at support
Version, git SHA and build target on startup and on /v1/health. An inbound x-request-id is echoed on every response, not only errors, because “quote the request id” is useless advice if a successful-but-wrong response has none. --log-format json when a shipper is reading.
stdout is the answer
Diagnostics go to stderr. They used to share stdout with the payload, so a single rejected-input warning could break forge status --json for whoever hit it. There is a test.
The checkpoint carries a schema version now, so a file written by a newer
forge is refused by name with the remedy rather than half-read. A checkpoint
from before the field existed is stamped rather than migrated, because nothing
about its tables changed, so a job in progress is not broken by the
introduction of a version field. The protection is forward.
Apache-2.0
Every number here is reproducible from the repository
The kill, the drain, the fleet comparison and the overload run all ship as
scripts and mocks you can run in three terminals. The regression that caused
the head-of-line block is pinned by a test. A verification run on 2026-09-10
re-ran every claim above in one sitting and landed within about 5% of the
figures shown, which is what sharing a box with a soak test costs.