AI model price watch

China model API costs for cheap fallback work.

A practical, dated comparison of China-origin LLM APIs and a Claude baseline, focused on Matt's use case: fallback agent/coding work when premium APIs do not feel worth the bill.

As of 2026-10-04 Updates scheduled: Every Friday Prices shown in USD per 1M tokens Scenario: 10M input + 2M output
$1.76cheapest scenario cost: 01.AI Yi Lightning
$40same rough session on Claude Sonnet 5
22.7×Sonnet cost multiple versus cheapest listed
11models/providers tracked

Best picks right now

#9 · 01.AI$1.76

Yi Lightning

Tiny cheap bilingual model

Short, cheap bilingual tasks

Input
$0.147/M
Output
$0.147/M
Cached
—/M
Context
16K

Watch: Short context and weaker coding suitability make it a narrow fallback.

#8 · ByteDance / Doubao$1.8

Seed 2.0 Mini / Lite class

Cheapest lightweight tier

Bulk cheap text tasks, first-pass summaries, non-sensitive low-stakes jobs

Input
$0.1/M
Output
$0.4/M
Cached
$0.02/M
Context
128K to 256K

Watch: Not a frontier coding model; international catalog may differ from mainland Doubao.

#2 · Tencent$2.38

Hunyuan Hy3

Ultra-cheap general fallback

Very low-cost general work, triage, summaries, first-pass edits

Input
$0.132/M
Output
$0.528/M
Cached
$0.033/M
Context
256K

Watch: Quality needs task-specific testing before trusting it for code changes.

#1 · DeepSeek$2.7

DeepSeek V4.1 Flash

Cheap strong fallback

Default low-cost fallback, summaries, debugging, routine agent work

Input
$0.15/M
Output
$0.6/M
Cached
$0.003/M
Context
1M

Watch: Peak pricing can double. Treat China-hosted APIs as not appropriate for secrets or private home-lab credentials.

#4 · Alibaba / Qwen$4.96

Qwen3.7 Plus

Balanced general model

General assistant work, code review, long-context tasks where Qwen quality matters

Input
$0.276/M
Output
$1.1/M
Cached
—/M
Context
256K to 1M tiered

Watch: Pricing changes by region, context band, promo, and thinking mode.

#3 · Alibaba / Qwen$5

Qwen3 Coder 480B A35B

Best coding value candidate

Coding agents, repository edits, tool use, long code context

Input
$0.3/M
Output
$1/M
Cached
$0.1/M
Context
262K

Watch: Routed price varies by provider. Re-check route before assuming cost.

Full price table

RankModelRoleInputOutputCached inputContext10M in + 2M out
#9 Yi Lightning01.AI Tiny cheap bilingual model $0.147 $0.147 — 16K $1.76
#8 Seed 2.0 Mini / Lite classByteDance / Doubao Cheapest lightweight tier $0.1 $0.4 $0.02 128K to 256K $1.8
#2 Hunyuan Hy3Tencent Ultra-cheap general fallback $0.132 $0.528 $0.033 256K $2.38
#1 DeepSeek V4.1 FlashDeepSeek Cheap strong fallback $0.15 $0.6 $0.003 1M $2.7
#4 Qwen3.7 PlusAlibaba / Qwen Balanced general model $0.276 $1.1 — 256K to 1M tiered $4.96
#3 Qwen3 Coder 480B A35BAlibaba / Qwen Best coding value candidate $0.3 $1 $0.1 262K $5
#5 MiniMax M2.7 / M3 discounted tierMiniMax Low-cost alternate $0.3 $1.2 $0.06 512K to 1M tiered $5.4
#6 GLM-4.5Zhipu / Z.AI Stronger but pricier $0.6 $2.2 $0.11 128K $10.4
#7 Kimi K2.7 CodeMoonshot / Kimi Capable but not budget-first $0.95 $4 $0.19 256K $17.5
#10 Claude Haiku 4.5Anthropic Western baseline comparison $1 $5 $0.1 200K $20
#11 Claude Sonnet 5Anthropic Premium escalation baseline $2 $10 $0.2 200K $40

Recommendation

Fallback ladder

  1. Local model first for private home-lab work and anything containing secrets.
  2. DeepSeek Flash or Hunyuan Hy3 for cheap low-risk fallback work.
  3. Qwen3 Coder when the task is code-heavy and cheap general models struggle.
  4. Claude Haiku/Sonnet only for escalation, review, or high-stakes tasks.

Safety and budget notes

  • Do not send API keys, credentials, private home-lab logs, or personal documents to China-hosted APIs unless explicitly approved.
  • Set hard provider spend caps. For fallback experiments, a $1–$2 per-run cap and $20–$30 monthly cap is a sane starting point.
  • Reasoning/thinking modes, long context, repeated retries, and verbose outputs are the common bill multipliers.

Source notes

These public prices move often. Friday updates should re-check provider docs or live pricing pages before changing numbers.