Know the cost of intelligence.

Model your AI cost architecture before you build. Optimize it before you commit. Measure it against plan after you ship.

Build a cost architecture → Get early access // no card, no signup to try
Studio // building a cost architecture 0/5 models chosen
parallel branches · 100% routed
1 chatbot/general pick model
2a rag/support pick model
2b agentic/tools pick model
2c classify/route pick model
3 summarize/batch pick model
● Running estimate $0 / mo corridor — ≈ $— per conversation · blended $—/1M tok
Delivery mix set on step 5
Compute lanes API · self-hosted GPU H100 80GB · CoreWeave · 3-yr
01 // What it prices

Price the architecture, not a prompt.

Real AI apps are phases and parallel branches, and every node can run on a different model and a different delivery mode — pay-per-token, batch, provisioned, cached, or your own GPUs. That's where the money actually moves, and it's what per-token calculators can't see.

Cost Architecture // the map you build
PARALLEL BRANCHES · 100% ROUTED 1 chatbot/general Sonnet 4.6 2a rag/support Llama 3.3 70B 2b agentic/tools Qwen3 235B A22B 2c classify/route Llama 3.1 8B 3 summarize/batch Sonnet 4.6
// each node carries its own workload shape, model and delivery mode — costs roll up to one number
your architecture
Model / versionTokens / moUsed in DeliverySpend / mo
Claude Sonnet 4.6 AWS Bedrock
6.64B$1.44/1M 13 $0
Llama 3.3 70B GPU Self-hosted
1.00B$1.12/1M 2a $0
Qwen3 235B A22B GPU Self-hosted
8.20B$1.29/1M 2b $0
Llama 3.1 8B GPU Self-hosted
250M$0.06/1M 2c $0
Model spend · month 1 16.09B   $0
// same five nodes, same traffic, same token counts — only the model and delivery choices change
REQUESTS / MO TOKENS / MO PHASES SPEND / MO 200k requests chatbot · rag · agentic First-pass 16.09B tok Loop + retry what-if only 5 nodes 2a·2b·2c fan-out LOOP + RETRY · WHAT-IF ON COST SENSITIVITIES, NEVER IN THE QUOTE $23,875 / mo ✓ WITHIN QUOTE · $15,555 – $32,785
// stress-test it: loop depth, retry rate, and a harness regression that silently multiplies calls
// the quote adds tool schemas, hidden costs and production overhead on top of model spend
02 // Why you can trust it

An honest range beats a confident guess.

Every estimate ships with a corridor whose width reflects how much you actually told us. Answer the quick pass in minutes for a wide, honest range; keep going and watch it tighten. Precision comes from supplying information — never from a tool declaring confidence it hasn't earned.

Corridor // narrows as you answer ±5% at full depth
quick + routing + tools full ±36% ±5% $23,875
parametercache_hit_rate · rag/support
default0% · engineering default, unmeasured (FD-09)
still at defaults53 of 55 inputs · each one widens the band
pricingpriced against the live catalogue · 2026-10-05
// every default carries its basis and a provenance dot — confirm one and the band tightens to ±10% on it
03 // What you do with it
quote$23,875 / mo
corridor$15,555 – $32,785
assumptions pinned16
volume200k req/mo · locked
cache hit rate0% · locked
LLM calls / task10 · locked
countersignedbuyer ✓ · seller ✓

Lock the assumptions, not just the quote.

shipping with early access

Every AI cost dispute starts the same way: the quote was right, but the assumptions moved. Infermaven pins what sits underneath a number — volumes, cache rates, tool-call counts — and both sides countersign. When reality drifts, you renegotiate the assumption, not the relationship.

// locked quotes export as an OTEL monitoring schema, and feed the anonymized benchmark pool behind Insights

04 // What Studio unlocks

Swap the model. Keep the architecture. Watch the whole bill move.

Most calculators price a prompt. Studio prices the system you actually built, then lets you swap any node’s model or delivery mode and reprices every phase that depends on it. Three options side by side, each a real estimate with its own corridor. Pick one and it becomes your quote.

Cost options // one architecture, three prices Review & Budget · step 9 of 9 · priced against the live catalogue · 2026-10-05
🔒 Architecture · locked5 nodes · 200k req/mo
Option 1 · your estimateAs built2 API · 3 GPU
Option 2 · StudioCheaper frontier tier2 swaps · 3 inherited
Option 3 · StudioSelf-host the front door2 swaps · 3 inherited
Intent & routing 1chatbot/general · 200k req/mo
Claude Sonnet 4.6AWS Bedrock$2,612stdbase
Claude Haiku 4.5AWS Bedrock$871std−$1,741⇄
Llama 3.3 70BSelf-hosted · H100 80GB$1,046gpu/hr−$1,566⇄
RAG · support copilot 2aparallel branch · 100% routed
Llama 3.3 70BSelf-hosted · H100 80GB$1,124gpu/hrbase
Llama 3.3 70BSelf-hosted · H100 80GB$1,124gpu/hr—
Llama 3.3 70BSelf-hosted · H100 80GB$1,124gpu/hr—
Agentic workflow 2bparallel branch · 100% routed
Qwen3 235B A22BSelf-hosted · H100 80GB$10,613gpu/hrbase
Qwen3 235B A22BSelf-hosted · H100 80GB$10,613gpu/hr—
Qwen3 235B A22BSelf-hosted · H100 80GB$10,613gpu/hr—
Classification / extraction 2cparallel branch · 100% routed
Llama 3.1 8BSelf-hosted · H100 80GB$15gpu/hrbase
Llama 3.1 8BSelf-hosted · H100 80GB$15gpu/hr—
Llama 3.1 8BSelf-hosted · H100 80GB$15gpu/hr—
Summarize · batch 3merge of 2a · 2b · 2c
Claude Sonnet 4.6AWS Bedrock$6,954batchbase
GPT-5 miniMicrosoft Azure$658batch−$6,296⇄
Llama 3.1 8BSelf-hosted · H100 80GB$341gpu/hr−$6,613⇄
Production overhead A-2612% on the subtotal
$2,558
$1,594
$1,577
Total / mo
Σ node cells + overhead, exactly. Every column is a real estimate.
$23,875baselineP10–P90 $15,555–$32,785 · ±36%
$14,874-38%P10–P90 $6,705–$23,997 · ±58%Use this option →
$14,715-38%P10–P90 $6,400–$24,088 · ±60%Use this option →
⇄ swap sheet · 3 · Summarize · batch · what Option 2 chose fromranked by closeness on the attributed benchmarks, then $/mo · scores stay attributed to their publisher · no composite, no recommendation
Claude Sonnet 4.6 current · AWS Bedrock$6,954
GPT-5 mini Microsoft Azure$658 -91%
GPT-5.5 Microsoft Azure$12,115 +74%
GPT-5.4 Microsoft Azure$6,058 -13%
🔒 Free prices Option 1 · $23,875/mo Studio prices the same architecture three ways, keeps every version in My Drive, and ships the finance audit pack behind each number. Unlock the board →
01

The architecture is locked. Only the models move.

Phases, routing shares, volumes and pass rates stay exactly as you built them. So when a total drops 38%, you know it was the swap, not a quiet change to an assumption.

02

Alternatives are ranked, never recommended.

The swap sheet lists the nearest models by attributed benchmark scores and by cost, with the source of every score. There is no blended index and no “best pick”. You read the numbers and decide.

03

Every column is a real estimate.

Each option runs through the same engine as your quote, carries its own corridor, and exports to the audit pack. Select one and the whole review page reprices against it. Nothing on the board is a projection of a projection.

// swaps that cross a tokenizer boundary raise a blocking pre-ship question instead of a silent conversion. one token count cannot describe two tokenizers, and the board says so.

05 // Pricing

Model cost architecture free. Optimize what you build in Studio. Benchmark what you deploy with Insights.

Every tier produces honest numbers — corridors are never paywalled. What you buy is depth, memory, and the market view.

// Quote & TCO

Free

for sellers & buyers on a deal
$0
  • Everything through Performance & Delivery, with the rest applied as badged defaults
  • Model Explorer — live prices, every delivery mode
  • Assumption locks — pinned, versioned, exportable
  • Buyer/seller countersignature coming
Get early access →
// The practitioner's workbench

Studio

for engineers & FDEs
$79 / seat / mo
// early-access pricing — locked 12 months
  • everything in Free, and:
  • Full depth: sensitivities, tools, hidden costs, architecture
  • Quality benchmarks in Model Explorer — scores from Artificial Analysis
  • Saved quotes & portfolio comparison
  • OTEL schema export mapped to your phases, models and delivery modes
Get early access →
// The market view

Insights

for FinOps & engineering leaders
$500 / org / mo
  • everything in Studio, and:
  • Benchmarks from real locked quotes, recency-weighted
  • Quartiles by workload shape, industry & region
  • Sample sizes always shown — no quartile without its n
Join the beta →

Price your AI before you build it.

Early access opens in cohorts. Join the list and we'll run your first TCO estimate with you — free, hands-on, no commitment.

// joining the list means we’ll email you about your cohort. nothing else, unless you tick the box.