Leading AI models compared: where things stand in September 2026 | WhaleBiz

The race for the best AI: where things stand in September 2026
The previous edition of this article described the balance of power as of April 2026. In five months it turned over completely: almost every model that was a flagship then is the previous generation now. That is itself the article's main conclusion, and below we explain why we stopped publishing score tables.
Here is what changed at the three main providers.
Anthropic Claude: the Claude 5 family
The Claude 4.6 and 4.8 lines discussed in the spring have been replaced by the Claude 5 family. The current line-up:
- Claude Fable 5.1 and Claude Fable 5 - the most capable widely available tier, for hard reasoning and long-horizon agentic work. 1M context window, $10 and $50 per million input and output tokens.
- Claude Opus 5 - the working flagship. 1M context window, $5 and $25 per million. The Opus 4.8, 4.7 and 4.6 line is still served at the same price.
- Claude Sonnet 5 - the balance of quality and cost. 1M context window, $2 and $10 per million, noticeably cheaper than Sonnet 4.6 at $3 and $15.
- Claude Haiku 4.5 - the cheapest tier: 200K context window, $1 and $5 per million.
More important than the price list is the change in how these models are steered. The fixed "thinking budget" (budget_tokens) is gone in the Claude 5 family, replaced by adaptive thinking and an effort level from low to max. The practical meaning for a business is simple: on the same model you can now deliberately economise on easy tasks and pay more only where the cost of an error is high.
OpenAI GPT: GPT-6 Astra and the GPT-5.6 family
In spring the flagship was GPT-5.4. By September 2026 the picture is different. July 2026 brought the GPT-5.6 family in three lines: Sol (the flagship), Terra (balancing quality and cost) and Luna (the cheapest), alongside a specialised GPT-5.6 Cyber for security work. Above them sits GPT-6 Astra, described in the official documentation as the company's most capable model for end-to-end work.
All of these run with a context window of roughly 1.05M tokens. The price movement deserves separate attention: at the end of July 2026 OpenAI cut Luna by 80% and Terra by 20%, and at the end of August it temporarily cut Sol by more than 20%. For a business that is a direct signal: a payback calculation made in spring is already wrong by autumn and is worth redoing.
Google Gemini: Pro and Flash have split apart
Something new happened at Google: the Pro and Flash lines stopped moving together. As of September 2026 the Google DeepMind catalogue holds:
- Gemini 3.8 Flash - the main workhorse for agentic tasks at scale, plus the specialised Gemini 3.8 Flash Cyber.
- Gemini 3.5 Flash-Lite - for high-volume, cost-sensitive scenarios.
- Gemini 3.1 Pro for complex tasks, and Gemini 3.1 Deep Think for science and research.
- Separate multimodal models: Gemini Omni (video), Gemini Audio, Gemini Embedding 2, and the image generator Gemini Image (Nano Banana).
Google says Gemini 3.5 Pro is coming, and Flash 3.7 and 3.8 carry introductory pricing that expires on 31 December 2026. That last point is worth remembering when planning next year's budget.
Want to consult with us?
We can help you choose, build and deploy the perfect AI solution for your business. Leave your details and we'll get back to you.
Table: what actually separates them
Only the parameters that move slowly and can be leaned on when choosing. Prices are per million tokens, input and output, at the provider's list price and before caching or batch discounts.
| Model | Context window | Price in/out | Typical role |
|---|---|---|---|
| Claude Fable 5.1 | 1M | $10 / $50 | hardest reasoning, long agentic runs |
| Claude Opus 5 | 1M | $5 / $25 | working flagship, complex dialogue and documents |
| Claude Sonnet 5 | 1M | $2 / $10 | price-quality balance at volume |
| Claude Haiku 4.5 | 200K | $1 / $5 | simple tasks at scale |
| GPT-6 Astra | ~1.05M | see OpenAI pricing | end-to-end multi-step work |
| GPT-5.6 Sol / Terra / Luna | ~1.05M | see OpenAI pricing | flagship / balance / economy |
| Gemini 3.8 Flash | see Google catalogue | introductory price to 31.12.2026 | agentic tasks at scale |
| Gemini 3.1 Pro / Deep Think | see Google catalogue | see Google pricing | complex reasoning, research |
Why we removed the benchmark scores
The previous edition carried a detailed table of percentages on SWE-bench, GPQA Diamond, AIME and Humanity's Last Exam. We removed it deliberately, and here is why.
In the five months between editions each of the three labs shipped a new generation. Any number frozen into the text lives for weeks; the article lives in search for years. The result is that a reader makes a decision on last year's data believing it is current. For material people find in Google eighteen months after publication, that is not a small inaccuracy, it is straightforward misinformation.
What lasts longer and is worth leaning on: the line-ups, the context window, the published price, and the task profile a model was built for. Those are in the article. For current scores, go to a live leaderboard such as LMArena and to the providers' own documentation, linked at the end.
Open models: when they make sense
Open weights closed a lot of ground during 2026. DeepSeek shipped the V4 family, Meta continues with Llama 4, and Qwen and Mistral are active. The practical case for a business is not about matching the flagship on scores, it is about two things: data never leaves your infrastructure, and the cost per request does not depend on another company's price list.
The cost of that is equally clear. Hardware, operations, updates and security become your job rather than the provider's. For a clinic, a law firm or a financial company with strict privacy requirements that trade is often justified. For a business that needs a working WhatsApp agent in two weeks, almost never.
How to pick a model for your task
The right question is not "which model is best" but "which model gets this task to a result most cheaply". A practical breakdown follows.
For customer support and chatbots
Take a fast, cheap tier: Gemini Flash, GPT-5.6 Luna, Claude Haiku 4.5 or Sonnet 5. At conversational volume the token-price difference turns into real money, while the difference in "intelligence" on routine questions about prices and opening hours is invisible. More on how such a solution is built on the AI for support page.
For complex consultation and document work
Here a flagship is justified: Claude Opus 5, GPT-5.6 Sol or GPT-6 Astra, Gemini 3.1 Pro. A scenario where the model parses a contract, cross-references several documents and is accountable for the conclusion is exactly the case where the cost of an error exceeds the difference in token price.
For coding and agentic work
Claude Opus 5 and Claude Fable 5.1, GPT-6 Astra, Gemini 3.8 Flash. The key practical parameter here is not a benchmark score but the model's ability to carry a long chain of steps without losing context, and the effort level you are willing to pay for.
For strict privacy requirements
Open models on your own infrastructure: DeepSeek, Llama, Qwen, Mistral. Beyond that the question stops being about the model and becomes one about storage architecture and who legally owns the database.
For maximum economy
The cheapest tier at any of the three providers, plus discipline on the application side: cache the stable part of the prompt, keep answers short, and stop running a flagship on the question "what time do you open".
What from "the 2026 trends" actually happened
- Autonomous coding agents. Happened. No longer a demo but a working tool, and all three providers sell dedicated lines for agentic work.
- Multi-agent systems. Happened. Orchestrating several specialised agents instead of one universal one became the standard architecture.
- Native multimodality. Happened. Every provider has separate models for video, audio and images.
- Adaptive thinking instead of a fixed budget. Happened during 2026: control over reasoning depth moved from "how many tokens to spend" to "how hard to try".
- Million-token context windows. Happened, and stopped being a differentiator: it is now a baseline property of the top tiers rather than an advantage.
A separate thread is what all this means for an economy in which agents start transacting with each other. We covered that in our piece on virtual agent economies.
Why we do not commit to one model
The practical conclusion from the whole article is one thing. Between two editions of this material the balance of power turned over completely, and that will happen again. A solution hard-wired to a particular model from a particular provider goes stale with it and has to be rewritten.
At WhaleBiz we build agents so that the model behind a given task can change without reworking the solution: cheap and fast for volume conversation, strong for complex reasoning and documents. The client does not end up paying for one provider's weak spots in order to get their strengths.
If you are deciding what to build your agent on and want to discuss a concrete task rather than an abstract ranking, get in touch.
Frequently asked questions
Which AI model is most cost-effective for customer support and chatbots?
For high-volume conversational traffic you do not count who is smarter, you count what one handled conversation costs at acceptable quality. As of September 2026 that job usually falls to the fast, cheap lines: Gemini 3.8 Flash and 3.5 Flash-Lite from Google, GPT-5.6 Luna from OpenAI, and Claude Haiku 4.5 or Claude Sonnet 5 from Anthropic. A flagship goes only on the scenarios where a mistake costs more than the difference in token price.
What is the practical difference between a 200K and a 1M context window?
The window sets how much material the model holds in front of it in a single request, without retrieval. 200K tokens fits a large knowledge base or a long conversation; a million fits several whole documents, months of message history, or a sizeable code repository. For a support agent the window is rarely the bottleneck: what matters there is the quality of retrieving the right passage, not the ability to load everything at once.
Why would a business want open models like DeepSeek, Llama or Qwen?
Open weights can run on your own infrastructure, so data never goes to an external provider. That is the main argument wherever privacy requirements are strict: healthcare, law, finance. The gap to the closed flagships narrowed considerably during 2026, but there is a cost of its own: hardware, operations and updates become your job.
Why is there no benchmark score table in this article?
Because it goes out of date faster than the article can ship. The large labs release new generations every few weeks, and a number frozen into the text misleads the reader a month later. Line-ups, context window sizes and published prices last longer, so those are what the article carries. For current scores, go to a live leaderboard.
Why does WhaleBiz not commit to a single AI model?
Because the optimal model depends on the task, and the balance of power shifts every few months. We build agents so the model behind a given task can be swapped without rewriting the solution: fast and cheap for volume conversation, stronger for complex reasoning and document work. Committing to one vendor means paying for their weak spots along with their strengths.
Sources
- Anthropic, Claude models overview and the official Anthropic pricing page.
- OpenAI, model catalogue in the developer documentation.
- Google DeepMind, Gemini model catalogue.
- LMArena - a live leaderboard for checking current results.

Michael Romm
Michael is the founder and CEO of WhaleBiz, leading business and marketing strategy. An expert in data (SQL, Python) and developing automation and AI solutions for businesses.