Run your agents
under one unified API.

Run production agents on Onefold with one OpenAI-compatible endpoint for every frontier model, per-call cost and latency tracing, and real traces you can replay as tests.

app.onefoldhq.com/keys

API keys

New key
Balance
$482.19
Spent this month
$133.25
Active keys
4
Projects
3
NameKeyProjectLimitSpentLast used
runway-prodsk-ag-…4f2arunway$200 / mo$84.212m ago
tasks-prodsk-ag-…9c11tasks$100 / mo$41.0914m ago
atlas-devsk-ag-…20e7atlasNone$6.833h ago
localsk-ag-…7b55runway$20 / mo$1.122d ago

runway / traces / tr_8f2a4c

runway.forecast_query

ok

sess_7c2 · trace 3 of 4

Duration
3.84s
Cost
$0.0536
Tokens
14,208
Spans
6
  • plan.intent340ms$0.0021
  • tool.fetch_txns890mstimeout
  • tool.fetch_txns210msretry · ok
  • retrieve.policy180ms8 docs
  • reason.projection1.92s$0.0448
  • compose.summary480ms$0.0067

reason.projection llm

model
claude-sonnet-4-5
ttft
312ms
tokens
8,140 in / 602 out
Seeded from tr_8f2a4c
Run all
claude-opus-5

How long is our runway?

Net burn is $18.4k a month. At the current balance you have 14 months of runway, ending Nov 2027.

1.84s · 312ms ttft · 8,140 → 602 · $0.0448

Sync chats
claude-sonnet-5

How long is our runway?

Runway is roughly 14 months on $18.4k net burn.

0.91s · 198ms ttft · 8,140 → 214 · $0.0076

Sync chats
claude-haiku-4-5

How long is our runway?

14 months.

0.38s · 96ms ttft · 8,140 → 41 · $0.0011

Sync chats

vs claude-opus-5  ·  cost −83%  ·  latency−51%  ·  output −64%

01

Run

  • One baseURL, every frontier model
  • Automatic failover across providers
  • Per-key spend limits and budgets
  • Streaming passthrough, no added latency

02

Trace

  • Full span trees, tool calls included
  • Cost, tokens, and latency per step
  • Sessions across multi-turn runs
  • OpenTelemetry native, any framework

03

Experiment

  • Compare models side by side
  • Seed any run from a real trace
  • Real cost, not estimates
  • Every run traced automatically

Write-ups

A 72:1 input-to-output ratio in a five-step agent run

Sixty thousand tokens in, eight hundred out. Where the rest of it went, step by step.

Coming soon

What 3 coding agents cost to finish the same task

Cursor, Claude Code and Cline, one ticket, one repo.

Coming soon

The cheapest model that still calls the right tool

Coming soon

One model, 4 providers: does the endpoint change the answer?

Same weights, different serving stack. Latency and throughput measured per endpoint.

Coming soon

Is Kimi better at building UI than the top 3 models?

Coming soon

Accuracy per dollar, not accuracy

The ranking nobody publishes, because it needs first-party pricing.

Coming soon