A preset is six dials, not a category.
The six workload types you pick from are archetypes over six underlying cost dimensions. Studio keys the estimate off the dimensions and treats the type as a preset, so a hybrid that fits no bin degrades gracefully: you override one dial and keep the preset. This page shows the dials, the defaults each preset sets, and for every default whether it rests on a published source, a derivation, or an estimate we badged as one.
Each dimension gates a pricing mechanism the providers themselves made structural.
| dimension | what it determines | pricing mechanism it gates |
|---|---|---|
| D1 · Latency tolerance | Whether work can move off the synchronous path. | Batch eligibility; provisioned-throughput economics. |
| D2 · Context growth | Whether input cost is constant, linear, or super-linear over a session or task. | Context management as a lever; long-context pricing tiers. |
| D3 · Calls per task | The multiplier between tasks (what you count) and requests (what the invoice counts). | Estimate variance; retry cost; where spend guardrails belong. |
| D4 · Output : input ratio | Which price, input or the dearer output, dominates the bill. | Whether max_tokens discipline or context trimming moves the total. |
| D5 · Prefix stability | The fraction of the prompt that is identical across requests. | Prompt caching, which stacks with everything else. |
| D6 · Non-inference infra share | How much of the total lives outside the token bill. | Vector stores, sandboxes, seats, compliance: the costs domain-cut data never sees. |
No two presets share a profile. That is what makes them presets.
The last column is what the dimensions imply about delivery, described rather than recommended. Studio's picker still shows only the tiers your chosen provider actually sells.
| workload | D1 · Latency tolerance | D2 · Context growth | D3 · Calls per task | D4 · Output : input ratio | D5 · Prefix stability | D6 · Non-inference infra share | delivery modes that fall out |
|---|---|---|---|---|---|---|---|
| RAG · support copilot | SecondsUser-facing, synchronous | Flat, assembledContext rebuilt per request from retrieval; 8k to 32k in | 1 to 3Main call, embed, optional rerank | about 1 : 20200 to 800 out against 8k to 32k in | Medium to highSystem prompt stable; retrieved chunks churn | HighVector store, embeddings, corpus storage | On-demand; caching on the prefix; provisioned only at sustained high utilization |
| Chatbot/General | SecondsStreaming expected | Accumulates per turnSuper-linear; 15k to 40k by turn 20 unmanaged | 1 per turnSessions of 5 to 30 turns | about 1 : 10, fallingRatio worsens as history grows | HighPersona and guardrails identical across sessions | Low to mediumSession store, chat UI | On-demand; caching first; provisioned only at very high sustained concurrency |
| Agentic workflow | MixedInteractive trigger, async execution | Accumulates per stepScratchpad and tool outputs until the task completes | 5 to 30+The multiplicative dimension; high variance | VariableReasoning and structured output inflate both sides | HighTool schemas and instructions fixed across steps | Medium to highSandboxes, orchestration, external APIs, tracing | On-demand or async queue; caching on the schema prefix; spend guardrails |
| Batch summarization | Hours (up to 24h)No user in the loop | Flat, per documentIndependent jobs; 8k to 50k in each | 1 per documentVolume from document count, not call fan-out | about 1 : 40200 to 500 out; the most predictable of the six | Very highInstructions identical; only the content slot varies | LowStorage and pipeline orchestration | Batch by default; caching stacks; provisioned is not warranted |
| Code assistant | Sub-second to asyncCompletions strict; review and test generation tolerant | Large, semi-stableRepository snapshot reused across many requests | 1 (completion) to 30+ (agent mode)Spans the widest range of any type | about 1 : 1520 to 300 token completions; reviews longer | Highest of any type90%+ cache hits on active codebases | Distinct axisSeats against raw API; the break-even is a seat question | On-demand; repository-context caching is the lever; async for PR review |
| Classification / extraction | Hours (mostly)On-demand only for live routing | Flat, per itemSchema prompt constant, content varies | 1 per itemPlus validation retries, tracked as a KPI | about 1 : 50+10 to 200 token structured output | near 100%Cache hit rate approaches the ceiling | LowPipeline only; the strongest fine-tune candidate | Batch plus caching; a fine-tuned small model at volume |
Coding agenthighest
Presents as Code assistant. Behaves as Agentic on D3 (1 to 20+ calls per task) and D1 (sub-second to minutes).
A code-assistant preset under-estimates by the full call multiplier. Override D3 and keep the preset.
RAG-backed chatmost common
Presents as Chat assistant. Behaves as Chat on D2 (per-turn growth) plus RAG on D6 (vector store, embeddings) and per-request chunk injection.
Neither preset alone captures both. Chat misses the infrastructure lines; RAG misses the history growth.
Extraction inside an agentcontained
Presents as Agentic. Behaves as A classification step within D3.
Harmless on the orchestrator's context. If it fans out to a cheap fine-tuned model, the blended per-task cost has two model prices.
Long-form generationno bin exists
Presents as Batch summarization. Behaves as Output-dominated (D4 inverts to about 5 : 1), batch-tolerant, high prefix stability.
The nearest preset carries input-heavy assumptions and puts the optimization on the wrong side. Content generation exists as its own preset for this reason.
The defaults, graded by what stands behind them.
The value in each cell is what Studio prefills today, read from the same module Studio reads. The grade and the note say where that kind of number comes from. Cache-eligible share is the best grounded: it follows from the caching pages. Token shapes are derivable from first principles. Usage volumes are the weakest, with no public benchmark, and your own telemetry replaces them first.
| unit economics | RAG | Chat | Agentic | Batch | Code | Classification |
|---|---|---|---|---|---|---|
| Monthly volume (users or items) | Estimated 500users· engineering default No public source. Sized as a mid-market contact center (about 500 active agents). Validate against headcount or session analytics. | Estimated 2,000users· engineering default No public source. Sized for a mid-size internal deployment (about 2,000 employees). Validate against auth or active-session logs. | Estimated 150users· engineering default No public source. Sized for a power-user persona (about 150 users). Validate against orchestration or product analytics. | PartialE 100,000docs/mo· engineering default Pipeline workload, no MAU. Monthly document volume from a first-principles cadence: 5,000 documents a day over 20 working days. | PartialC 120users· engineering default Engineering-org AI adoption rates consistent with about 120 active developers at a mid-size software company. A rough proxy, not a direct figure. | Estimated 500,000records/mo· engineering default Pipeline workload, no MAU. About 25,000 records a day from typical CRM or ticket ingestion. No published workload-level benchmark. |
| Requests per user per month (or calls per item) | Estimated 400req/user· engineering default No public source. From a work pattern: 20 working days times 20 queries a day per agent. Validate with contact-center analytics. | Estimated 60req/user· engineering default No public source. From a session model: 3 sessions a week, 5 turns each, 4 weeks. Validate against session logs. | Estimated 40req/user· engineering default No public source. From a work pattern: 2 agentic tasks a day over 20 days. Validate with workflow run history. | PartialE Monthly volume expressed directly. First principles from a typical ingestion cadence. Validate against job submission logs. | PartialD 800req/user· engineering default Copilot's fill-in-the-middle studies imply about 30 completions a day for active developers. Review frequency estimated separately. The best public proxy available. | PartialE 1.03calls· engineering default Monthly record volume from a first-principles pipeline rate. No workload-level public benchmark. Validate against upstream event counts. |
| Input tokens per request | PartialEG 4,200tokens· engineering default Cookbook RAG examples show the prompt structure. System prompt 800, three chunks 2,700, query 150, history 550. First-principles word-count math. | PartialEG 3,500tokens· engineering default Multi-turn docs show history accumulation. Averaged across session depth at word count times 1.33. Grows super-linearly without compression. | PartialFG 3,200tokens· engineering default Benchmark traces publish step-level context. Agentic docs show scratchpad accumulation. Averaged across early and late steps. | PartialE 8,500tokens· engineering default An average business document of about 6,000 words times 1.33, plus a system prompt of about 500 tokens. Highly variable by document type. | PartialDE 1,400tokens· engineering default Copilot's completion context is about 6,000 characters. Blended with review and explanation requests that carry a full diff. | PartialE 750tokens· engineering default Schema definition of about 350 tokens, fixed, plus about 400 tokens of variable content. Much shorter than summarization. |
| Output tokens per request | PartialE 380tokens· engineering default A support answer of about 280 words times 1.33. Validate by sampling response lengths from your logs. | PartialE 420tokens· engineering default An assistant response of about 315 words times 1.33. Varies widely by use case; set max_tokens from real sessions. | PartialF 480tokens· engineering default Benchmark traces publish reasoning-trace lengths per step. Structured tool calls add format overhead. Averaged across orchestrator steps. | PartialE 350tokens· engineering default A structured summary of about 260 words (three bullets and a paragraph) times 1.33. Constrain with an explicit format and max_tokens. | PartialDE 165tokens· engineering default Completions run 20 to 100 tokens; reviews and explanations are much longer. Blended at a 75 : 25 completion-to-review ratio. | PartialE 45tokens· engineering default Structured JSON with two to five fields, about 45 tokens. Highly predictable with JSON mode. |
| Cache-eligible share of inputStudio's input is the cache-eligible share; the research graded the resulting hit rate. Same page, different quantity. | StrongA 55% of input· engine default (`defaultPerformance`) The caching page states cache read at 0.1 times and write at 1.25 times the base input price. The share is derived from the system prompt's part of total input with proper prefix placement. | StrongA 55% of input· engine default (`defaultPerformance`) Automatic caching for multi-turn conversations is documented. The share is the static prefix's part of total input at average session depth. | StrongA 50% of input· engine default (`defaultPerformance`) Tool definitions and system instructions cache as a fixed prefix; the scratchpad rotates. Varies by task length, so a share rather than a fixed number. | StrongA 45% of input· engine default (`defaultPerformance`) The schema prompt is identical across all documents in a job, so most of the input is cacheable. The most favorable caching profile of the six. | StrongA 70% of input· engine default (`defaultPerformance`) A fixed repository-context prefix enables high hit rates. The pre-warming pattern is described on the page. | StrongA 75% of input· engine default (`defaultPerformance`) Schema and label definitions are identical across records in a batch. Placing them before the per-record content is what makes the share high. |
| LLM calls or actions per request | N/A One LLM call per user request: embed, retrieve, generate. | N/A One LLM call per user turn. | PartialFG 10calls· engineering default Benchmark traces publish LLM call counts per task; multi-agent docs show orchestration patterns. The default of 10 LLM and 8 tool calls is a midpoint. | N/A One LLM call per document, or two to four for very long documents that need chunking passes. | N/A One LLM call per completion or review request. | N/A Effectively 1.03 LLM calls per record: 97% first-pass success plus a 3% retry rate. |
| business-unit conversion | RAG | Chat | Agentic | Batch | Code | Classification |
|---|---|---|---|---|---|---|
| Primary business unit | Estimated Support tickets a month. No industry benchmark for contact-center query volume; ranges span orders of magnitude. Validate against your ticketing system. | Estimated Conversations a month. Derived from MAU times sessions per user, both themselves ungrounded. Validate against session logs. | Estimated Tasks a month. No published benchmark, and a task is workload-specific: a lookup or a multi-hour research run. Validate against run history. | PartialE Documents a month, first principles from a typical ingestion cadence. Validate against scheduler or platform logs. | PartialC Developers. The CIO survey gives context on org sizes and adoption; a proxy, not a per-developer token figure. | Estimated Records a month. No published benchmark. Validate against upstream event counts: CRM, support platform, transaction database. |
| Primary conversion ratio | Estimated Queries per ticket, default 3.5. No published benchmark; simple tickets resolve in one or two queries, escalations need five to ten. A practitioner estimate. | Estimated Turns per conversation, default 5. Varies widely: Q&A about 2, research about 15, support bots about 4. Validate against session analytics. | PartialF LLM calls per task, default 10. Benchmark tasks run 8 to 15; 10 is a centre estimate. Validate against orchestration logs. | PartialG Calls per document, default 1. One request per document is the documented batch pattern; multi-pass chunking for very long documents raises it. A design parameter, not an estimate. | PartialD Completions per developer per day, default 30. Consistent with published activation and acceptance patterns. Validate with IDE telemetry. | PartialAG Effective calls per record, default 1.03, from JSON-mode reliability: 97% first-pass success. Track your own retry rate; above 5% signals a prompt or schema problem. |
| Secondary conversion ratio | N/A One LLM call per query; the query-to-ticket ratio is the whole conversion. | N/A One LLM call per turn; turns per conversation is the whole conversion. | PartialF Tool calls per task, default 8. About 80% of LLM steps invoke a tool. Each tool result re-enters the context, compounding input cost beyond the call count. | N/A One call per document. No secondary factor. | Estimated Review requests per PR, default 5. No published benchmark; estimated from typical PR interactions: diff review, tests, docs, refactor, explanation. | N/A One call per record; the retry factor sits in the primary ratio. |
| A | Anthropic · Prompt caching | Cache read and write multipliers against the base input price; per-model minimum token thresholds; the pre-warming pattern. | page undated · read 2026-09-12 |
| B | Andreessen Horowitz · Welcome to LLMflation – LLM inference cost is going down fast | The cost trajectory for equivalent-performance inference. Informs how long a quoted rate is likely to hold. | page dated 2024-11-12 · read 2026-09-12 |
| C | Andreessen Horowitz · How 100 Enterprise CIOs Are Building and Buying Gen AI in 2025 | Adoption rates by use case and engineering-org deployment patterns. A proxy for org sizes, not a token figure. | page dated 2025-06-10 · read 2026-09-12 |
| D | GitHub · How GitHub Copilot is getting better at understanding your code | Fill-in-the-middle, the character budget a completion model sees, and the acceptance-rate deltas from more context. | page dated 2023-05-17 · read 2026-09-12 |
| E | First-principles derivation | Word count times 1.33 tokens per word; session-depth models; typical document-length distributions by type. Any practitioner can reproduce it. | a method, not a page |
| F | AgentBench, SWE-bench, WebArena task traces | Published academic benchmarks whose task traces include LLM call counts, reasoning-trace lengths and tool-invocation patterns. Named as methods; individual papers are not cited to a customer. | a method, not a page |
| G | Anthropic · Claude Developer Platform documentation | RAG, multi-agent and agentic workflow examples with token counts and prompt structure. | page undated · read 2026-09-12 |
// the one default the overview discloses in full, with all six fields: #assumptions · what is described but never printed: #never