LLM API Pricing Comparison
Compare input/output/cached prices, context windows and blended cost across Claude, GPT, Gemini and DeepSeek
TL;DR Summary & Citation: [LLM API Pricing Comparison] is used to Compare input/output/cached prices, context windows and blended cost across Claude, GPT, Gemini and DeepSeek. Core input and files are processed in the current browser and are not sent to the PocketKit API. The page records only an aggregate tool-use count that excludes input content.
LLM pricing comparison and blended price
The tool gathers input, output and cached prices plus context window for Claude, GPT, Gemini and DeepSeek into one table, and weights the two prices into a blended figure at the input:output ratio you set for one-dimensional sorting. Price data is built in locally, editable, and never fetched online.
Usage Guidelines & Tips
- Estimate your real input:output ratio first (chat/RAG about 3:1, pure generation about 1:1) so the blended sort matches actual spend.
- Non-Claude prices marked reference can change; verify on each provider site before production.
Frequently Asked Questions
1. How is the "blended" price calculated?
It is the weighted average of input and output price at the input:output ratio you pick, giving one comparable per-million price for sorting. Chat/RAG is often about 3:1 and pure generation closer to 1:1; your real cost depends on your actual mix.
2. Are these prices accurate?
Claude prices are official references; OpenAI, Google and DeepSeek are marked as reference values that can change, so verify on each provider site.
3. For the same job, which model is actually cheapest?
First decide whether your workload is input-heavy or output-heavy. Classification, extraction and format conversion produce short responses, so input and cached-input rates dominate; long-form writing and code generation are output-heavy, where a cheap output rate matters far more. The table compares input, output and cached pricing alongside context windows side by side. In practice the cheapest setup is usually not one model at all but tiered routing — send easy tasks to a cheap model and reserve the strongest one for genuinely hard reasoning.
4. What is Prompt Caching and how can it cut input bills by up to 90%?
Anthropic, OpenAI, and DeepSeek support prompt caching on repetitive prefixes. When identical system prompts or reference documents exceed minimum thresholds (1024 tokens for Claude, 64 tokens for DeepSeek), cached tokens are billed at ~10% of standard input rates with substantially reduced Time-to-First-Token (TTFT).
5. How can I leverage Batch APIs for 50% discounts?
For asynchronous workloads that do not require immediate responses (e.g. bulk catalog tagging, synthetic dataset generation, offline extraction), OpenAI and Anthropic offer Batch API endpoints. Jobs completed within a 24-hour SLA receive a 50% price cut with separate rate limit pools.