A model is the underlying system doing the work inside an AI app. OpenAI makes GPT models; Anthropic makes Claude models. Changing the model can change the quality, speed, and cost of the answer—even inside the same app.
Each update will keep this same format: where the leading models stand, what the latest release adds, and whether it gives us a reason to reconsider what we use.
My rankings · three models worth discussing right now| My rank | Model / maker | Capability score | Cost per test task |
|---|---|---|---|
| 01 | GPT-6 AstraOpenAI · leader | 53 | $3.26 |
| 02 | Claude Fable 5.1Anthropic · joint leader | 53 | $7.63 |
| 03 | DeepSeek V4.1 FlashDeepSeek · new challenger | 40 | $0.27 |
About my rankings: These are my rankings of the models most worth discussing right now: flagship models from the top AI companies, plus DeepSeek’s lower-cost challenger. That doesn’t mean they’re the ones you should use. That depends on the task—and flagship models tend to be the most expensive.
What do the numbers mean?
Artificial Analysis’ Intelligence Index v4.3 combines ten tests; higher scores mean stronger performance across those tests. They are not percentages of human intelligence. Astra and Fable tie; DeepSeek scores lower, but costs much less. Figures use the highest-effort versions evaluated, including Fable’s fallback configuration. Cost is a weighted average API charge per test task, not a subscription price or a quote for your work. Checked September 12, 2026. Sources: Astra and Fable comparison · DeepSeek evaluation.
DeepSeek V4.1 Flash: how much capability do you need to pay for?
Last week we covered Astra. This week, Chinese AI developer DeepSeek released V4.1 Flash, which can work with text and images. The discussion is about delivering useful capability with less computing overhead and a much smaller usage bill.
DeepSeek says its new design reduces the memory and storage needed to retain information while working. That is a technical efficiency claim from the developer. Separately, Artificial Analysis’ evaluation shows a lower bill: about one-twelfth of Astra’s cost per test task, with a lower overall capability score. That supports a value story; it does not establish equal performance or prove a twelvefold reduction in electricity use.
What “open weight” means: DeepSeek makes the trained model files available under an MIT license. Other organizations can download, modify, and run the model on their own infrastructure or through a hosting provider. That offers more deployment choices. It doesn’t mean running it is free, that every part of its training is public, or that this large model will fit on an ordinary laptop. See the model release.
What I’m watching: whether the price advantage survives real work—after counting retries, mistakes, and human review. A cheaper model that meets the standard for a repeatable task could change what is practical to automate.
Create your own evaluation: Build a prompt asking the AI to do something - the more complex the better. Try it with different models, especially new ones as they come out and see the differences yourself. I have one and so should you.
Read the launch news — Reuters →DeepSeek’s announcement and efficiency claims · Independent performance and cost results

