The 'top ten' ranking was based on a composite of several leader boards. Of course it is subjective and changes all the time

Model Input Cached input / cache hit Output Notes
Claude Opus 5 $5.00 $0.50 $25.00 Fast mode: $10 input / $50 output. Batch: $2.50 input / $12.50 output. platform.claude+1
Claude Fable 5 $10.00 $1.00* $50.00 Anthropic states a 90% discount for prompt-cache hits; cache-write prices were not surfaced in the retrieved Fable-specific page. platform.claude+1
GPT-5.6 Sol $5.00 $0.50 $30.00 Cache write: $6.25/MTok. developers.openai
GPT-5.6 Terra $2.50 $0.25 $15.00 Cache write: $3.125/MTok, inferred from OpenAI’s stated 125% of input rate; confirm in your dashboard before large commitments. developers.openai
GPT-5.6 Luna $1.00 $0.10 $6.00 Cache write: $1.25/MTok, inferred from the same published schedule. “Max effort” changes reasoning budget, not the base input/output price tier. developers.openai+1
Grok 4.6 $2.00 $0.50 $6.00 Standard rate for prompts below 200K tokens; long-context pricing may increase to $4 input / $1 cached / $12 output per MTok. I could not retrieve xAI’s own pricing page directly, so this is vendor-attributed secondary reporting. mem0+1
Kimi K3 $3.00 $0.30 $15.00 Moonshot/Kimi’s official pricing page presents $0.30, $3.00, and $15.00 per MTok; these correspond to cache hit, input, and output. platform.kimi+1
Tencent Hunyuan 3 (Hy3) $0.132 promo Not separately surfaced $1.40 promo Tencent TokenHub’s official promotion page shows approximately $0.132 input and $1.40 output per MTok; this is a promotional rate, so treat it as subject to change. tencentcloud+1

Conclusion

  1. First thing I noticed is that Grok is finally breaking into the top models. It had been dragging behind. Claude Sonnet fell off my top 10 list.

  2. All these models are closed source. Can't run them locally even if you did have a couple of H200s in your lab.

  3. The two Chinese models, Kimi K3 and Tencent Hy3 are much cheaper than Claude, and some GPT models. However, GPT Luna and Grok 4.6 are close. Claude is pretty expensive.

  4. It is important to learn how to use prompt caching effectively to reduce token count. more about hat

  5. I did the research on the model ranking myself. Then I used Perplexity to fetch the cost data and used Claude Sonnet to format the table. Image generated using Gemini Pro.