AI Daddy › Model Landscape
Pricing and Costs · Model Landscape
Understanding the cost structure of LLM systems is essential for production planning. This chapter covers pricing models, cost optimization strategies, and…
Pricing and Costs
Understanding the cost structure of LLM systems is essential for production planning. This chapter covers pricing models, cost optimization strategies, and total cost of ownership analysis.
Table of Contents
Pricing Models
Token-Based Pricing
Most LLM APIs charge per token:
Cost = (input_tokens × input_rate) + (output_tokens × output_rate)
Key observations:
- Output tokens cost 2-5x more than input tokens
- Pricing varies significantly by model tier
- Some providers offer batch discounts
Tiered Pricing
Some providers offer volume discounts:
| Tier | Monthly Spend | Discount |
|---|
| Standard | 0−5K | 0% |
| Growth | 5K−50K | 10-20% |
| Enterprise | $50K+ | Custom negotiation |
Commitment-Based Pricing
Pre-purchase tokens at discounted rates:
Standard: $2.50 / 1M input tokens
Committed (1-year): $2.00 / 1M input tokens (20% savings)
Current API Pricing
August 2026 Pricing
Last verified: August 15, 2026. Prices change frequently. Always re-check: OpenAI, Anthropic, Google, xAI, DeepSeek
August 2026 price moves (the two that matter most): Claude Sonnet 5's introductory 2/10 per 1M became permanent on August 10, 2026, and the scheduled September 1 increase to 3/15 was canceled, so Sonnet 5 is now permanently cheaper than the Sonnet 4.6 it replaced. Going the other way, DeepSeek raises V4 prices 3x to 12x effective August 16, 2026 at 16:00 UTC and switched from flat rates to peak and off-peak billing, ending its run as the unambiguous cheap option: at peak, V4-Flash output (1.32per1M)nowcostsmorethanGPT−5.6Luna′s1.20. Also new: GPT-5.6-Cyber at 12.50/75 (August 10, restricted access), Gemini 3.7 Flash at a half-price 0.75/3.75 through December 31 2026, and Grok 4.6 at 2/6 with a long-prompt tier that applies the higher rate to every token in the request once the prompt reaches 200K.
August 2026 retirements and sunsets: Claude Opus 4.1 retired from the Claude API on August 5, 2026 (the last 15/75 Opus tier; still live on Bedrock and Google Cloud on their own schedules). The OpenAI Assistants API sunsets August 26, 2026, replaced by the Responses API plus the Conversations API, with no automated migration for Threads. OpenAI is also shutting down its Evals Platform, Agent Builder, and Reusable Prompts on November 30, 2026 (evals go read-only October 31; OpenAI points eval users to the third-party Promptfoo). Anthropic's legacy Workbench and experimental prompt-tools APIs shut down August 17, 2026.
Deprecations effective in 2026: OpenAI retired GPT-4o, GPT-4.1, GPT-4.1-mini, o4-mini from ChatGPT on Feb 13, 2026; gpt-5.2-chat-latest and gpt-5.3-chat-latest deprecated May 8, 2026; Realtime API Beta removed May 12, 2026; Sora app shut down April 26, 2026 (API EOL Sep 24, 2026). Anthropic retires Claude Sonnet 4 and Claude Opus 4 on June 15, 2026, and Claude Opus 4.1 on August 5, 2026. Google Vertex retired gemini-3-pro-preview Mar 26, 2026; Project Mariner shut down May 4, 2026. Gemini 2.5 Pro/Flash deprecated June 17, 2026.
Price moves: Anthropic released Claude Fable 5 on June 9, 2026 at 10/50 per 1M: its most capable widely released model (Mythos-class with safeguards), priced at 2x Opus 4.8 but less than half of Claude Mythos Preview. Claude Mythos 5 (same model, safeguards lifted, Glasswing-only) shares the 10/50 price. Anthropic released Claude Opus 4.8 on May 28, 2026 at the same 5/25 per 1M as Opus 4.7, with an optional fast mode at 10/50 per 1M (about 2.5x faster and 3x cheaper than the Opus 4.7 fast mode, which was 30/150). DeepSeek made its 75% V4 Pro discount permanent on May 22, 2026: from June 1, 2026 the new list price drops to 25% of the original (0.435/0.87 per 1M input/output), and the cache-hit input price for all DeepSeek models was cut to 1/10 of the launch price on April 26, 2026. DeepSeek V4 Flash (0.14/0.28 per 1M, 1M context) is the cheapest frontier-class API by a wide margin.
OpenAI (GPT-5.x Generation)
| Model | Input / 1M | Output / 1M | Notes |
|---|
| GPT-5.6 Sol ⭐ NEW | $5.00 | $30.00 | GA July 9, 2026. Flagship of the three-tier GPT-5.6 line. 1M context, 128K max output. |
| GPT-5.6 Terra ⭐ NEW | $2.00 | $12.00 | Cut 20% on July 30, 2026 from 2.50/15. GPT-5.5-class quality at roughly half the price; the general production default. |
| GPT-5.6 Luna ⭐ NEW | $0.20 | $1.20 | Cut 80% on July 30, 2026 from 1/6. Priced against open-weight competition; the volume tier for classification, extraction, and routing. |
| GPT-5.6-Cyber ⭐ NEW | $12.50 | $75.00 | August 10, 2026. Cached input $1.25. 400K context. Daybreak Red tier only: identity verification, legal attestations, approved use cases, Responses API only. Hardware security keys mandatory on individual accounts from September 1, 2026. |
| GPT-5.5 | $5.00 | $30.00 | Released April 23, 2026. 1M context. New class of multimodal flagship. |
| GPT-5.5 Instant ⭐ NEW | check latest | check latest | Default in ChatGPT and chat-latest since May 5, 2026. 52.5% fewer hallucinations on high-stakes prompts. |
| GPT-Realtime-2 ⭐ NEW | $32.00 (audio) | $64.00 (audio) | Released May 7, 2026. GPT-5-class realtime voice. |
| GPT-Realtime-Translate ⭐ NEW | (audio pricing) | (audio pricing) | 70+ input → 13 output languages. |
| GPT-5.4 Pro | $30.00 | $180.00 | Maximum reasoning; long-context doubles to 60/270 |
| GPT-5.4 | $2.50 | $15.00 | Flagship; native computer use; cached input $1.25 |
| GPT-5.4-mini | $0.75 | $4.50 | Best cost/performance in GPT-5 tier |
| GPT-5.4-nano | check latest | check latest | Smallest GPT-5.4 variant; released March 2026 |
| GPT-4o | $2.50 | $10.00 | Retired from ChatGPT Feb 13, 2026; API access varies |
| GPT-4o-mini | $0.15 | $0.60 | Legacy; check API availability |
Anthropic (Claude Fable + 4.x Generation)
| Model | Input / 1M | Output / 1M | Context | Notes |
|---|
| Claude Opus 5 ⭐ NEW | $5.00 | $25.00 | 1M | Released July 24, 2026 (claude-opus-5). Unchanged from Opus 4.8. Optional Fast mode at 10/50 per 1M, about 2.5x faster. New default on Claude Max. |
| Claude Sonnet 5 ⭐ NEW | $2.00 | $10.00 | 1M | Released June 30, 2026 (claude-sonnet-5), default across products. Introductory pricing made permanent August 10, 2026; the scheduled September 1 rise to 3/15 was canceled. Cache write 2.50(5min)/4.00 (1 hr); cache hit 0.20;Batch1 / $5. Permanently cheaper than Sonnet 4.6. |
| Claude Fable 5 ⭐ NEW | $10.00 | $50.00 | 1M | Released June 9, 2026 (claude-fable-5) on Claude API, Claude Platform on AWS, Bedrock, Vertex AI, Microsoft Foundry. Most capable widely released Anthropic model (Mythos-class with safeguards; sensitive queries fall back to Opus 4.8 in under 5% of sessions). Adaptive thinking always on; 128K max output; 30-day data retention applies. |
| Claude Mythos 5 ⭐ NEW | $10.00 | $50.00 | 1M | Same underlying model as Fable 5 with safeguards lifted in some areas. Limited availability: Project Glasswing partners and select biology researchers. Succeeds Mythos Preview at less than half its price. |
| Claude Opus 4.8 | $5.00 | $25.00 | 1M | Released May 28, 2026 on API, Bedrock, Vertex AI. Dynamic Workflows research preview with parallel subagents. Optional fast mode at 10/50 per 1M (about 2.5x faster, 3x cheaper than the Opus 4.7 fast mode). SWE-bench Verified 88.6%; SWE-Bench Pro 69.2%; OSWorld-Verified 82.3%. |
| Claude Opus 4.7 | $5.00 | $25.00 | 1M | Released April 16, 2026 on API, Bedrock, Vertex, Microsoft Foundry. Higher-resolution vision, improved SWE. Fast mode is no longer offered on this model: a fast-speed request returns an error. |
| Claude Opus 4.6 | $5.00 | $25.00 | 1M | 128K max output; adaptive thinking at standard rates. |
| Claude Sonnet 4.6 | $3.00 | $15.00 | 1M | Superseded by Claude Sonnet 5 (June 30, 2026), which is both newer and cheaper at 2/10. |
| Claude Haiku 4.5 | $1.00 | $5.00 | 200K | Fastest Anthropic model; cache hit input $0.10 / 1M. |
| Claude Mythos Preview | n/a | n/a | - | Restricted research preview (~11 Glasswing partners); succeeded by Claude Mythos 5 on June 9, 2026. |
NOTE
Claude 1M context at standard pricing: Fable 5, Opus 4.8, Opus 4.7, Opus 4.6, and Sonnet 4.6 include the full 1M token context window at standard rates with no premium tier for long context. Batch API offers a 50% discount. Cache hits cost 10% of the standard input price. Fast mode is available on Opus 5 and Opus 4.8 at 10/50 per 1M and is no longer offered on Opus 4.7 (which errors) or Opus 4.6 (which runs at standard speed and standard rates); the historical Opus 4.7 fast tier was 30/150. Fast-mode pricing stacks with caching multipliers but is not available on the Batch API or Claude Platform on AWS. No Fable-tier fast mode at launch.
Google (Gemini 3.x Generation)
| Model | Input / 1M | Output / 1M | Context | Notes |
|---|
| Gemini 3.7 Flash ⭐ NEW | $0.75 | $3.75 | 1M | GA August 13, 2026. Half-price introductory rate through December 31, 2026, then 1.50/7.50. Context caching 0.075/1M;Batch0.375 / $1.875. Model cards and cost plans should use the January 2027 numbers for anything long-lived. |
| Gemini 3.1 Pro | $2.00 | $12.00 | 1M | 200K+ context: 4.00/18.00 |
| Gemini 3.1 Flash | $0.10 | $3.00 | 1M | Best price/performance; high-volume |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | 1M | Deprecated June 2026 |
WARNING
Gemini 2.5 deprecation: Gemini 2.5 Pro and 2.5 Flash are scheduled for deprecation on June 17, 2026. Migrate to Gemini 3.x models.
xAI (Grok)
| Model | Input / 1M | Output / 1M | Context | Notes |
|---|
| Grok 4.6 ⭐ NEW | $2.00 | $6.00 | 500K | Released August 12, 2026. Cached input 0.50(upfrom0.30 on Grok 4.5, so cache-heavy loops do not get cheaper). At or above a 200K prompt the rate doubles to 4/12 for every token in the request, not just the excess. Fast variant is 2x. |
| Grok 4 | $3.00 | $15.00 | 256K | Native tool use; real-time search |
| Grok 4.1 Fast | $0.20 | $0.50 | 2M | High-volume, low-cost |
| Grok 3 mini | check latest | check latest | - | Faster, less accurate |
Open-Weight and Value-Tier Models via API (August 2026)
| Model | Input / 1M | Output / 1M | Context | Provider Examples |
|---|
| DeepSeek V4 Pro ⭐ REPRICING AUG 16 | 1.32peak/0.66 off-peak | 3.96peak/1.98 off-peak | 1M | Effective August 16, 2026 at 16:00 UTC, DeepSeek moved to peak/off-peak billing and raised prices 3x to 12x depending on token type (cache-hit input went from 0.003625to0.044, a 12.1x rise). Off-peak is exactly half peak. GA as build 0813 on August 13 with MIT weights. |
| DeepSeek V4 Flash ⭐ REPRICING AUG 16 | 0.44peak/0.22 off-peak | 1.32peak/0.66 off-peak | 1M | Same August 16 repricing (was 0.14/0.28). GA as build 0731 on July 31 with MIT weights. At peak, output now exceeds GPT-5.6 Luna's $1.20 per 1M. |
| Qwen3.8-Max ⭐ NEW | check latest | check latest | 262K (to ~1M) | Alibaba API; open weights August 12 under a commercially gated license. |
| Tencent Hy3 ⭐ NEW | ~$0.13 | ~$0.53 | 256K | Via OpenRouter. 295B / 21B-active MoE, Apache 2.0, global from August 5, 2026. Among the cheapest frontier-adjacent rates available. |
| MAI-Code-1.1-Flash ⭐ NEW | $0.20 | $1.20 | check latest | Microsoft, August 11, 2026. A 73% list-price cut versus MAI-Code-1-Flash; shipped into GitHub Copilot. |
| Muse Spark 1.2 ⭐ NEW | $1.25 | $4.25 | 1M | Meta Model API. A muse-spark-1.2-contributor tier costs 0.10/0.20 in exchange for permission to train on your prompts and completions: check policy before enabling. |
| DeepSeek-V3.2 | $0.28 | $0.42 | 128K | DeepSeek API. 98% cache-hit discount. Effective rates can drop 10–30× via routing. |
| Mistral Medium 3.5 ⭐ NEW | $1.50 | check latest | 256K | Mistral API. Unified chat/reasoning/coding/vision; 77.6% SWE-Bench Verified. |
| Kimi K2.6 ⭐ NEW | check latest | check latest | - | Moonshot API. 1T MoE / 32B active; agent swarm to 300 sub-agents. |
| Qwen 3.6-35B-A3B ⭐ NEW | check latest | check latest | - | Apache 2.0 weights; self-host or via API providers. |
| Llama 4 Scout | $0.11 | $0.34 | 10M | Together AI, Groq, Fireworks. Note: effective context degrades fast past 32K. |
| Llama 4 Maverick | $0.27 | $0.85 | 1M | Together AI, Groq, Fireworks. MoE-aware serving required. |
| DeepSeek-V3 | $0.25 | $1.10 | 128K | DeepSeek API, Together AI |
| DeepSeek-R1 | $0.55 | $2.19 | 128K | DeepSeek API |
| Mistral Large 3 | $0.50 | $1.50 | 256K | Mistral API, AWS Bedrock |
| Llama 3.3 70B | ~$0.10–0.20 | ~$0.30–0.60 | 128K | Groq, Together AI |
| Qwen2.5-Coder-32B | ~$0.50 | ~$1.00 | 32K | Together AI |
| Gemma 4 (31B / 26B-A4B MoE / E4B / E2B) ⭐ NEW | self-host | self-host | 256K | Apache 2.0. 140+ languages; native vision/audio; function calling. |
Embedding Models
| Model | Cost / 1M tokens | Dimension |
|---|
| Cohere Embed 4 ⭐ NEW | $0.10 | 256 / 512 / 1024 / 1536 (Matryoshka) |
| text-embedding-3-large | $0.13 | 3072 |
| text-embedding-3-small | $0.02 | 1536 |
| Voyage-3 | $0.06 | 1024 |
| Cohere embed-v3 | $0.10 | 1024 |
IMPORTANT
Inference-time Compute Costs: For models with "Extended Thinking" or reasoning modes (GPT-5.4 Pro, Claude Opus 4.6), you are charged for internal thinking tokens even if not shown to the user. This can increase total request cost by 2x-10x for logic-heavy tasks. Always set a budget_tokens cap in production.
Cost Calculation
def calculate_request_cost(
input_tokens: int,
output_tokens: int,
model: str
) -> float:
pricing = {
"gpt-5.4": {"input": 2.50, "output": 15.00},
"gpt-5.4-mini": {"input": 0.75, "output": 4.50},
"claude-sonnet-4.6": {"input": 3.00, "output": 15.00},
"claude-opus-4.6": {"input": 5.00, "output": 25.00},
"gemini-3.1-flash": {"input": 0.10, "output": 3.00},
}
rates = pricing[model]
cost = (
(input_tokens / 1_000_000) * rates["input"] +
(output_tokens / 1_000_000) * rates["output"]
)
return cost
Example Cost Calculations
Scenario 1: RAG Chatbot
Per request:
- System prompt: 500 tokens
- Retrieved context: 2,000 tokens
- User message: 100 tokens
- Response: 300 tokens
Input: 2,600 tokens, Output: 300 tokens
GPT-5.4 cost: (2600 × $2.50 + 300 × $15) / 1M = $0.0110 per request
At 10,000 requests/day:
Daily: $95
Monthly: $2,850
Scenario 2: Document Summarization
Per document:
- Document: 8,000 tokens
- Summary: 500 tokens
GPT-5.4 cost: (8000 × $2.50 + 500 × $15) / 1M = $0.0275
1,000 documents: $27.50
10,000 documents: $275
Monthly Cost Projection
def project_monthly_cost(
requests_per_day: int,
avg_input_tokens: int,
avg_output_tokens: int,
model: str
) -> dict:
per_request = calculate_request_cost(
avg_input_tokens, avg_output_tokens, model
)
daily = per_request * requests_per_day
monthly = daily * 30
yearly = monthly * 12
return {
"per_request": per_request,
"daily": daily,
"monthly": monthly,
"yearly": yearly
}
# Example
costs = project_monthly_cost(
requests_per_day=50000,
avg_input_tokens=2000,
avg_output_tokens=400,
model="gpt-5.4"
)
# Output: ~$18,750/month
Cost Optimization Strategies
Strategy 1: Model Routing
Route requests to appropriate model tiers:
class ModelRouter:
def __init__(self):
self.classifier = load_complexity_classifier()
def route(self, query: str, context: str) -> str:
complexity = self.classifier.predict(query)
if complexity < 0.3:
return "gpt-5.4-mini" # Simple queries
elif complexity < 0.7:
return "gpt-5.4-mini" # Medium, try cheap first
else:
return "gpt-5.4" # Complex queries
def route_with_fallback(self, query: str, context: str) -> str:
# Try cheap model first
response = self.try_model("gpt-5.4-mini", query, context)
if self.is_quality_sufficient(response):
return response
# Fallback to expensive model
return self.try_model("gpt-5.4", query, context)
Potential savings: 50-70% with minimal quality impact
Strategy 2: Prompt Optimization
Reduce token count without losing quality:
# Before: 2,500 tokens
system_prompt = """
You are a helpful customer support assistant for Acme Corp.
You have access to our product documentation and should answer
questions accurately and helpfully. Always be polite and professional.
If you don't know something, say so rather than making things up.
Format your responses clearly with bullet points when listing items.
[... more verbose instructions ...]
"""
# After: 800 tokens
system_prompt = """
You are Acme Corp's support assistant.
Rules:
- Answer from provided context only
- Admit uncertainty
- Use bullet points for lists
- Be concise
"""
# Savings: 1,700 tokens × $2.50/1M = $0.00425 per request
# At 10K requests/day: $42.50/day = $1,275/month
Strategy 3: Caching
Cache responses for repeated or similar queries:
class ResponseCache:
def __init__(self, ttl_seconds: int = 3600):
self.exact_cache = TTLCache(maxsize=10000, ttl=ttl_seconds)
self.semantic_cache = SemanticCache(threshold=0.95)
def get_or_generate(self, query: str, context: str) -> tuple[str, bool]:
# Check exact cache
cache_key = self.make_key(query, context)
if cache_key in self.exact_cache:
return self.exact_cache[cache_key], True # Cache hit
# Check semantic cache
similar = self.semantic_cache.find_similar(query)
if similar:
return similar.response, True # Semantic hit
# Generate new response
response = self.generate(query, context)
self.exact_cache[cache_key] = response
self.semantic_cache.add(query, response)
return response, False # Cache miss
# With 30% cache hit rate:
# Baseline: $3,000/month
# With caching: $2,100/month
# Savings: $900/month
Strategy 4: Batch Processing
Process multiple requests together for efficiency:
# Real-time: pay full price
for query in queries:
response = model.generate(query)
# Batch API (OpenAI offers 50% discount):
batch_responses = model.batch_generate(queries)
# Cost: 50% of real-time pricing
Strategy 5: Output Length Control
Limit response length appropriately:
# Reduce unnecessary output
response = model.generate(
prompt=prompt,
max_tokens=300, # Limit output
stop=["\n\n"] # Stop at natural break
)
# Cost impact:
# Before: avg 500 output tokens = $0.0075 per request (GPT-5.4)
# After: avg 250 output tokens = $0.00375 per request
# Savings: 50% on output costs
Cost Optimization Summary
| Strategy | Effort | Potential Savings |
|---|
| Model routing | Medium | 50-70% |
| Context Caching | Low | 60-90% (Input) |
| Prompt optimization | Low | 20-40% |
| Response caching | Medium | 20-40% |
| Batch processing | Low | 50% (OpenAI/Anthropic) |
Context Caching Economics
The "Golden Rule" for RAG (still true in 2026).
If you have a fixed system prompt or a shared knowledge base (prefix) larger than 10,000 tokens, Context Caching is mandatory.
Break-even Analysis (Claude Sonnet 4.6):
- Standard Input: $3.00 / 1M tokens
- Cached Input: $0.30 / 1M tokens (90% discount)
- Cache Write Fee: 3.75/1Mtokens(5−minTTLat1.25x);6.00 (1-hour TTL at 2x)
Break-even = (Write Fee) / (Standard Rate - Cached Rate) ≈ 1.4 requests (5-min) or 2.2 requests (1-hour)
If your long prefix is used by more than 2 users, caching it is strictly cheaper than sending it raw every time. Both OpenAI and Anthropic now offer batch API discounts (50% off) that stack with caching.
Self-Hosting & GPU Cloud Arbitrage
The Reserved vs. Serverless Tradeoff:
| Model Size | Serverless (RunPod/Together) | Reserved (Lambda/AWS) |
|---|
| Burst Capacity | Infinite (cold starts) | Fixed |
| Utilization | Pay only for compute time | 24/7 fixed cost |
| TCO Break-even | Cost-effective < 40% util | Cost-effective > 40% util |
Principal-level Nuance:
"GPU Cloud Arbitrage" involves moving production workloads between providers based on spot instance availability. Tools like Skypilot automate this, saving up to 60% on self-hosting costs by following "low-demand" regions globally. The rise of MoE models (Llama 4 Scout fits on a single H100, Maverick on ~2x H100, DeepSeek V4 Flash on 4x H100) has further reduced self-hosting GPU requirements compared to dense models.
When Self-Hosting Makes Sense
Break-even analysis:
API cost at scale:
- 1M requests/month
- 2,500 tokens average
- GPT-5.4: ~$37,500/month
- Claude Sonnet 4.6: ~$30,000/month
Self-hosted equivalent (Llama 4 Maverick via MoE):
- 2x H100 80GB: ~$6/hour × 730 = $4,380/month
- Engineering time: $5,000/month (0.5 FTE)
- Ops overhead: $2,000/month
- Total: ~$11,380/month
Savings vs GPT-5.4: $26,120/month = 70%
Savings vs Claude Sonnet 4.6: $18,620/month = 62%
Self-Hosting Cost Components
| Component | Monthly Cost | Notes |
|---|
| GPU compute | $5K-20K | Depends on model size |
| Storage | $200-500 | Model weights, logs |
| Networking | $100-500 | Egress, load balancing |
| Engineering | $5K-15K | Partial FTE for ops |
| Monitoring | $100-500 | Observability tools |
GPU Requirements by Model Size
| Model Size | GPU Config | Estimated Cost/Month |
|---|
| 7B (INT4) | 1x A10G | $500-800 |
| 7B (FP16) | 1x A100 40GB | $1,500-2,500 |
| 70B (INT4) | 2x A100 80GB | $5,000-8,000 |
| 70B (FP16) | 4x A100 80GB | $10,000-15,000 |
| 405B (INT4) | 8x H100 | $20,000-30,000 |
Decision Framework
Choose API when:
- Volume < 100K requests/month
- No ML ops expertise
- Need highest quality (frontier models)
- Fast iteration needed
Choose self-hosting when:
- Volume > 500K requests/month
- Have ML infrastructure team
- Data privacy requirements
- Predictable, stable workload
- Custom fine-tuning needed
Total Cost of Ownership
TCO Components
def calculate_tco(scenario: dict) -> dict:
# Direct costs
api_or_compute = scenario["monthly_api_cost"]
# Engineering costs
development = scenario["dev_hours"] * scenario["engineer_rate"]
maintenance = scenario["maintenance_hours"] * scenario["engineer_rate"]
# Infrastructure
vector_db = scenario["vector_db_cost"]
monitoring = scenario["monitoring_cost"]
# Indirect costs
downtime_risk = scenario["expected_downtime_hours"] * scenario["revenue_per_hour"]
monthly_tco = (
api_or_compute +
development / 12 + # Amortized over year
maintenance +
vector_db +
monitoring +
downtime_risk
)
return {
"monthly_tco": monthly_tco,
"yearly_tco": monthly_tco * 12,
"breakdown": {
"llm": api_or_compute,
"engineering": development / 12 + maintenance,
"infrastructure": vector_db + monitoring,
"risk": downtime_risk
}
}
Example TCO Comparison
Scenario: Customer Support Bot (50K requests/month)
| Cost Component | API-Based | Self-Hosted |
|---|
| LLM costs | $5,000 | $3,000 |
| Vector DB | $70 | $200 |
| Engineering (monthly) | $500 | $3,000 |
| Monitoring | $100 | $200 |
| Monthly Total | $5,670 | $6,400 |
At this scale, API is cheaper due to engineering overhead.
Scenario: Large-Scale RAG (2M requests/month)
| Cost Component | API-Based | Self-Hosted |
|---|
| LLM costs | $50,000 | $15,000 |
| Vector DB | $500 | $1,000 |
| Engineering (monthly) | $1,000 | $8,000 |
| Monitoring | $200 | $500 |
| Monthly Total | $51,700 | $24,500 |
At this scale, self-hosting is significantly cheaper.
Interview Questions
Q: How would you optimize costs for a high-volume RAG application?
Strong answer:
I would approach cost optimization in layers:
1. Architecture optimization:
- Model routing: Use cheap model for simple queries
- Caching: 30-40% of queries may be cacheable
- Prompt compression: Minimize system prompt tokens
2. Model selection:
Simple queries (60%): GPT-5.4-mini at $0.003/request
Complex queries (40%): GPT-5.4 at $0.011/request
Weighted avg: $0.0062/request (vs $0.011 all GPT-5.4)
Savings: 44%
3. Infrastructure:
- Batch embedding updates (50% cheaper)
- Right-size vector DB
- Use spot instances where possible
4. Monitoring:
- Track cost per query type
- Alert on anomalies
- Regular cost reviews
Q: When would you recommend self-hosting vs using APIs?
Strong answer:
Decision depends on multiple factors:
Volume threshold:
- Below 100K/month: Almost always API
- 100K-500K: Evaluate case by case
- Above 500K: Often self-hosting wins
Team capabilities:
- No ML ops: API regardless of scale
- Strong infra team: Consider self-hosting earlier
Quality requirements:
- Need absolute best: APIs (frontier models)
- Good enough works: Self-hosted open models
Other factors:
- Data privacy: May force self-hosting
- Latency control: Self-hosting gives more control
- Fine-tuning needs: Self-hosting enables more customization
My recommendation process:
- Start with APIs for fastest iteration
- Build abstraction layer for model switching
- Evaluate self-hosting when spend exceeds $10K/month
- Pilot with shadow deployment before committing
References
Previous: Capability Assessment | Next: Model Selection Guide