The first DeAI Inference Price Index snapshot is dated 2026-09-07 UTC. It is a list-price table, not a blended "cheapest API" ranking. On that date, OpenRouter's public model catalog listed openai/gpt-oss-120b at $0.037 input / $0.17 output per million tokens, the lowest open-weight row in this basket. DeepSeek V4 Flash 0731 listed at $0.14 / $0.28, matching DeepSeek's official API anchor recorded 2026-08-20.
Key takeaways
- Snapshot date: 2026-09-07 UTC. Observation surface: OpenRouter
GET /api/v1/modelslist pricing, converted from dollars-per-token to dollars per million tokens. Official DeepSeek V4 Flash remains the only first-party anchor carried forward from launch sourcing. - Lowest open-weight row in the basket: gpt-oss-120b at $0.037 / $0.17. DeepSeek V4 Flash 0731 at $0.14 / $0.28 matches the official API. A separate OpenRouter slug,
deepseek/deepseek-v4-flash(0423), listed cheaper at $0.089 / $0.177 — different checkpoint, different price. - Frontier closed APIs in the same catalog: GPT-5.5 at $5 / $30, Claude Sonnet 4.6 at $3 / $15, Gemini 2.5 Flash at $0.30 / $2.50, Grok 4.6 at $2 / $6.
- This is not a 14-host matrix. Together, Fireworks, Groq, DeepInfra, Cerebras, Hyperbolic, Nebius, and Morpheus are named as hosts to re-check on their own pages; they are not invented here.
- A 1 billion input / 250 million output mix costs ~$80 on gpt-oss-120b, ~$210 on official DeepSeek V4 Flash, and ~$12,500 on GPT-5.5 at these list rates, before cache and batch.
How to read this table
Every row is USD per million tokens, input / output, list price, on-demand. OpenRouter exposes pricing.prompt and pricing.completion as dollars per token; the numbers below are those fields × 1,000,000, rounded to three decimals. Batch variants (:batch) are omitted from the table — they are typically about half the on-demand rate on this aggregator and would mix two products.
List price is not your invoice. Output tokens usually cost several times input. Prompt caching, context-tier jumps, and minimums live on each host's pricing page. Switching OpenAI-compatible hosts is usually a base-URL and model-slug change; slugs are not portable. deepseek/deepseek-v4-flash and deepseek/deepseek-v4-flash-0731 priced differently on the same day.
The August 2026 issue named in the launch backlog was never fetched: the weekly sweep was not running. This September table is issue one. The hub at /prices holds the same snapshot as JSON.
The 2026-09-07 basket
Sorted by input list price. Source: OpenRouter models API unless noted.
| Model slug | Publisher | Input $/M | Output $/M | Notes |
|---|---|---|---|---|
| openai/gpt-oss-120b | OpenAI (open weights) | 0.037 | 0.170 | Apache 2.0 weights; aggregator list |
| z-ai/glm-5.3-flash | Z.ai | 0.075 | 0.250 | Flash tier, not GLM-5.3 full |
| deepseek/deepseek-v4-flash | DeepSeek | 0.089 | 0.177 | 0423 listing — not the 0731 checkpoint |
| meta-llama/llama-4-scout | Meta | 0.100 | 0.300 | Llama 4 Scout |
| qwen/qwen3.6-35b-a3b | Alibaba | 0.100 | 0.900 | MoE 35B-A3B |
| nousresearch/hermes-4-70b | Nous Research | 0.130 | 0.400 | See Hermes 4 run-guide |
| deepseek/deepseek-v4-flash-0731 | DeepSeek | 0.140 | 0.280 | Matches official API as of 2026-08-20 |
| meta-llama/llama-4-maverick | Meta | 0.200 | 0.696 | Llama 4 Maverick |
| google/gemini-2.5-flash | 0.300 | 2.500 | Closed; reference row | |
| qwen/qwen3-coder | Alibaba | 0.300 | 1.000 | Qwen3 Coder 480B A35B listing |
| moonshotai/kimi-k2.5 | Moonshot AI | 0.450 | 2.250 | See Kimi K2.5 run-guide |
| deepseek/deepseek-v4-pro | DeepSeek | 0.955 | 1.911 | 0423 listing |
| z-ai/glm-5.3 | Z.ai | 1.400 | 4.400 | Full GLM-5.3, not Flash |
| x-ai/grok-4.6 | xAI | 2.000 | 6.000 | Closed; reference row |
| moonshotai/kimi-k3 | Moonshot AI | 3.000 | 15.000 | See Kimi K3 run-guide |
| anthropic/claude-sonnet-4.6 | Anthropic | 3.000 | 15.000 | Closed; reference row |
| openai/gpt-5.5 | OpenAI | 5.000 | 30.000 | Closed flagship; reference row |
Official DeepSeek V4 Flash, recorded 2026-08-20 from DeepSeek's API docs and still matching the 0731 OpenRouter row on 2026-09-07: $0.14 / $0.28. That is the only first-party number in this issue. Kimi K2.5's launch-sprint official-API anchor was ~$0.60 / $3.00 as of 2026-08-20 (approximate); OpenRouter listed K2.5 at $0.45 / $2.25 on 2026-09-07. Treat that gap as two products (official vs aggregator), not a ranking.
Worked mix: 1B in, 250M out
Same traffic on four list-price rows:
| Route | Math | List cost |
|---|---|---|
| gpt-oss-120b (OpenRouter) | 1,000 × $0.037 + 250 × $0.170 | $79.50 |
| DeepSeek V4 Flash official / 0731 | 1,000 × $0.14 + 250 × $0.28 | $210 |
| Kimi K2.5 (OpenRouter) | 1,000 × $0.45 + 250 × $2.25 | $1,012.50 |
| GPT-5.5 (OpenRouter listing) | 1,000 × $5 + 250 × $30 | $12,500 |
Capability is not in this table. A cheaper row that fails your evals is not cheaper. For the eval side of that decision see GPT-5.5 vs open models and the cheapest LLM API roundup. Self-hosting changes the unit from tokens to GPU-hours; that math is in self-hosting vs API cost.
What this snapshot does not cover
Other hosts. The launch roundup named fourteen provider families: DeepSeek, OpenAI, Anthropic, Google, Mistral, Groq, Together, Fireworks, DeepInfra, Cerebras, OpenRouter, Hyperbolic, Nebius, and Morpheus. This issue prices the comparable public catalog we could fetch on one date. Direct Together, Fireworks, Groq, and marketplace routes are not filled with guessed numbers. Check those pricing pages before routing production traffic.
Morpheus. Morpheus is one decentralized marketplace among peers. It has no row here because this snapshot did not fetch a Morpheus list price. Absence is not a ranking. When a dated marketplace rate exists, it lands in this table on the same criteria as everyone else.
Quality, retention, refusals. Price without policy is half a decision. Retention pages live under the Trust Tracker. Refusal scores are not published until Cycle 1 of the Refusal Index completes — see this week's companion report.
How the index will update
Weekly, dated, append-only. A new snapshot does not rewrite this one. The JSON at /data/prices.json is the machine copy; this article is the citable write-up. If a number here disagrees with a provider's current page, the provider's page wins for buying decisions — this page wins as a historical record of what the aggregator listed on 2026-09-07.
FAQ
What is the cheapest open-weight LLM API in September 2026?
On the 2026-09-07 OpenRouter snapshot, openai/gpt-oss-120b listed at $0.037 input / $0.17 output per million tokens. DeepSeek V4 Flash 0731 listed at $0.14/$0.28, matching DeepSeek's official API. Your invoice still depends on input/output mix, cache hits, and the host you actually call.
How much does DeepSeek V4 Flash cost per million tokens?
DeepSeek's official API listed V4 Flash at $0.14 input / $0.28 output per million tokens as of 2026-08-20, and OpenRouter's deepseek/deepseek-v4-flash-0731 row matched that on 2026-09-07. A cheaper 0423 listing also appeared on OpenRouter at $0.089/$0.177 — checkpoint slugs are not interchangeable.
Are OpenRouter prices the same as official APIs?
No. OpenRouter is an aggregator: the listed rate is what that router charged for that model slug on the snapshot date, and it can include routing markup or a cheaper backend. Official DeepSeek, OpenAI, Anthropic, and Google pages remain the source for first-party rates.
Why isn't every host in this table?
This snapshot is OpenRouter's public catalog plus the official DeepSeek V4 Flash anchor. Together, Fireworks, Groq, DeepInfra, Cerebras, Hyperbolic, Nebius, and Morpheus set their own lists and enter later issues when fetched on the same date.
Do these prices include cache and batch discounts?
No. Rows are list prompt/completion rates in USD per million tokens. Batch rows on OpenRouter are often half the on-demand rate; prompt caching is billed separately. Normalize to your traffic mix before calling anything "cheapest."
Questions
- What is the cheapest open-weight LLM API in September 2026?
- On the 2026-09-07 OpenRouter snapshot, openai/gpt-oss-120b listed at $0.037 input / $0.17 output per million tokens. DeepSeek V4 Flash 0731 listed at $0.14/$0.28, matching DeepSeek's official API. Your invoice still depends on input/output mix, cache hits, and the host you actually call.
- How much does DeepSeek V4 Flash cost per million tokens?
- DeepSeek's official API listed V4 Flash at $0.14 input / $0.28 output per million tokens as of 2026-08-20, and OpenRouter's deepseek/deepseek-v4-flash-0731 row matched that on 2026-09-07. A cheaper 0423 listing also appeared on OpenRouter at $0.089/$0.177 — checkpoint slugs are not interchangeable.
- Are OpenRouter prices the same as official APIs?
- No. OpenRouter is an aggregator: the listed rate is what that router charged for that model slug on the snapshot date, and it can include routing markup or a cheaper backend. Official DeepSeek, OpenAI, Anthropic, and Google pages remain the source for first-party rates.
- Why isn't every host in this table?
- This snapshot is OpenRouter's public catalog plus the official DeepSeek V4 Flash anchor. Together, Fireworks, Groq, DeepInfra, Cerebras, Hyperbolic, Nebius, and Morpheus set their own lists and enter later issues when fetched on the same date.
- Do these prices include cache and batch discounts?
- No. Rows are list prompt/completion rates in USD per million tokens. Batch rows on OpenRouter are often half the on-demand rate; prompt caching is billed separately. Normalize to your traffic mix before calling anything 'cheapest.'
Sources
- OpenRouter API — list models — OpenRouter
- OpenRouter models — OpenRouter
- DeepSeek API Docs — DeepSeek
- OpenAI API Pricing — OpenAI
- Anthropic Pricing — Anthropic
- Google AI for Developers — Gemini API Pricing — Google
- Together AI — Together AI
- Fireworks AI — Fireworks AI
About DeAI
DeAI is an independent publication covering open-weight AI models, private inference, and decentralized infrastructure — the tools for running AI you actually control. We test providers on price, privacy, and refusal behavior and publish the numbers, not the vibes. DeAI is powered by Morpheus (mor.org), a decentralized inference marketplace, and covers it on the same terms as every other provider.
Powered by Morpheus and StrandCMS
Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider — we rank it wherever the data lands. StrandCMS is the open-source, agent-first framework this site is built on.
