The ReAct loop, priced

We priced the pattern everyone builds on.It leaks in three places.

ReAct — think, act, observe, repeat — is the most widely used agent pattern there is. Its problems are well known and almost never priced. So we priced them.

All three leaks close. Two need nothing but a different model in a different step. The third is a single change inside one step — and we name it.

Where it goes

Three leaks, three fixes.

A two-day horizon, twelve people running the workflow, every model in every model-bearing position — ranked by true cost per delivered task, which is API spend plus modeled failure exposure.

63%

The loop pays for its own memory, over and over. Nearly two-thirds of all model spend sits in the reasoning step, because the scratchpad is append-only — iteration five re-sends iterations one through four and pays for them again.

The fix: trim or summarize between iterations. No model choice substitutes for it.

2.07×

One model everywhere costs double. Running a single mid-catalog model at both steps costs 2.07× the best configuration on the board. Nothing about the workflow changes — only which model runs where.

The fix: pay for the step that repeats, economize on the step that runs once.

196 of 196

The cheapest tokens finish last. The configuration with the lowest published API price ranked dead last of 196 — 3.13× the winner's true cost, bought by paying a hundredth as much per token.

The fix: rank on cost per delivered task, never on the rate card.

Failure exposure is priced at $25 per silent defect — a declared assumption, not a measurement, and one the conclusion survives at anything above $0.08.

The part nobody expects

You don't trade quality for it.

You would expect the cheaper configuration to be the riskier one. It is not. The best configuration costs less and produces half the silent failures — 8.5 per run against 17.2. Same workflow, same steps, cheaper and safer at once.

Nor does spending more fix it. Put the flagship in the final step and that configuration lands third, not first — 14% more per delivered task than the winner. And most of that gap is not the token price: in this workload the flagship let more answers through wrong than the first-ranked configuration did — 209 silent failures per thousand against 186.

Dated on purpose

Today's models. Today's prices.

This was run against the model catalog of July 2026, and that date is doing real work. The catalog moves every few weeks; new models land, prices fall, and the winner moves with them. An answer computed against last quarter's catalog is not a cheaper answer — it is a wrong one.

Run identity aaafbcebe2e55630 · catalog July 2026 · ensemble of 64 seeds · deterministic replay verified.

Send us one workflow

Yours won't behave like this one.

That is the whole point — and the reason
the answer has to be computed, not looked up.

$15,000 introductory offer · one workflow.