AI rating — September 2026
Which AI models are the strongest right now — by the composite score: QueryWise's own rating plus all publicly available ratings.
The strongest AI today is Claude Fable 5
★ “Maximum” mode lineup
In Maximum mode your question goes to the 4 strongest AIs from different companies for independent cross-checking. The list refreshes automatically every hour as new models come out.
Leaderboard
The score (0–10) is a composite assessment: QueryWise's own rating combined with all publicly available AI model ratings.
| # | AI model | Overall score |
|---|---|---|
| 1 | Claude Fable 5 Anthropic |
9.7
|
| 2 | GPT-5.6 Sol OpenAI |
9.7
|
| 3 | Claude Opus 5 Anthropic |
9.5
|
| 4 | Kimi K3 Moonshot |
9.5
|
| 5 | Claude Opus 4.8 Anthropic |
9.1
|
| 6 | GPT-5.6 Terra OpenAI |
9.1
|
| 7 | GPT-5.5 OpenAI |
9.0
|
| 8 | Claude Sonnet 5 Anthropic |
8.9
|
| 9 | Grok 4.5 xAI |
8.6
|
| 10 | DeepSeek V4 Pro DeepSeek |
8.6
|
| 11 | GLM 5.2 Z.ai |
8.6
|
| 12 | GPT-5.6 Luna OpenAI |
8.5
|
| 13 | DeepSeek V4 Flash DeepSeek |
8.5
|
| 14 | Gemini 3.6 Flash Google |
8.5
|
| 15 | Gemini 3.1 Pro Google |
8.0
|
| 16 | MiniMax M3 MiniMax |
7.6
|
| 17 | Kimi 2.6 Moonshot |
7.6
|
| 18 | GLM 5.1 Z.ai |
7.1
|
| 19 | GPT-5.4 mini OpenAI |
7.1
|
| 20 | Grok 4.3 xAI |
6.8
|
| 21 | Gemini 3.5 Flash Lite Google |
6.7
|
| 22 | Claude Sonnet 4.6 Anthropic |
6.7
|
| 23 | Qwen3 Max Qwen |
5.2
|
| 24 | Claude Haiku 4.5 Anthropic |
5.0
|
| 25 | GPT-OSS 120B OpenAI |
5.0
|
| 26 | GPT-4.1 OpenAI |
4.5
|
| 27 | Mistral Large Mistral |
4.1
|
Scores 0–10 on the QueryWise scale: a composite per-dimension estimate factoring in our speed and reliability measurements.
Methodology: the score is compiled automatically from QueryWise's own rating and all publicly available AI model ratings. Values update automatically and represent QueryWise's composite assessment.
What’s happening in the rankings in July 2026
The top spot is currently split on quality: Claude Fable 5 and GPT-5.6 Sol both scored 9,8. That doesn’t mean they perform the same way. Fable 5 has a median response time of 30 seconds and costs about 43 tenge per question. Sol is noticeably faster — 17 seconds — and cheaper, at around 25 tenge. For users handling a high volume of requests, that price gap quickly turns into real spending.
That’s why the current table leader won’t necessarily be the best choice for a specific task. Fable 5 makes sense when the highest possible answer score matters most and the latency is acceptable. GPT-5.6 Sol looks like the more practical all-purpose option: the same score, nearly half the wait, and 18 tenge less per question. It’s a strong fit for work chats, drafting text, and regular checks.
The chasing pack is close behind. Kimi K3 scored 9,4, costs about 13 tenge, and has a median response time of 29 seconds. Claude Opus 4.8 also scored 9,4, but costs roughly 22 tenge; its response time is unavailable in the current measurements. The gap between second and third place is just 0,4 points, while the gap between third and sixth is 0,2. The rankings are tight: a small change in quality or speed could reshuffle several models.
The low-cost options are particularly interesting. GPT-5.6 Terra scored 9,2 and costs around 13 tenge. Grok 4.5 also scores 9,2, responds in 16 seconds, and costs about 10 tenge. For quick everyday requests, it’s one of the most rational choices on the list. Claude Sonnet 5, with a score of 9 and the same estimated price, suits users who want a more affordable Claude without aiming for the top spot.
There’s an important contrast at the other end of the table. GPT-5.5 ranks eighth, costs around 25 tenge, and has a median response time of 77 seconds — nearly five times slower than GPT-5.6 Sol. Gemini 3.5 Flash is the fastest among models with a stated median: 11 seconds at a price of about 10 tenge, although its score is 8,7. GPT-5.6 Luna received the same score at the same price, but its speed isn’t listed in the data.
The measurements make one thing clear: in Kazakhstan, the cost per question now matters just as much as the ranking position. If you’re watching every request, look at Kimi K3, Grok 4.5, Terra, or Gemini 3.5 Flash. For demanding work, start with Fable 5 or Sol, then compare the results on your own examples. These are assistants, not medical, legal, or financial advice: critical decisions should be checked separately.