🔍 Read the full analysis: Choosing The Right AI Role In My September 2026 Workflow on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Thorsten Meyer’s 29 September 2026 analysis argues that with six frontier models clustered within roughly 20 points on the Artificial Analysis index but differing about 100x in cost per task, model selection is now a cost-per-task decision. His workflow pairs Claude Opus 5.5 for building with the newly released GPT-6.1 Sol for review at a fraction of the price.
GPT-6.1 Sol launched on 29 September 2026 into a frontier AI market where, according to Thorsten Meyer’s analysis on ThorstenMeyerAI.com, six leading models now sit within about 20 index points of each other on the Artificial Analysis Intelligence Index while their cost per task differs by roughly 100x. His conclusion: the practical question has shifted from “which model is smartest?” to “which model clears a quality bar at the lowest cost per task?” — and his answer pairs Claude Opus 5.5 as the main builder with the new Sol model as a routine, cheap reviewer.
Meyer’s framework, published the day Sol launched, assigns each model a specific role rather than treating them as interchangeable. Opus 5.5 (released 22 September, index score 58 at max effort, $5.98 per task) is his main model for development work. GPT-6.1 Sol, released the same day as the article, scores 51 at its xhigh effort setting but costs $0.39 per task — making it his choice for detailed investigation and independent review. GPT-6 Astra and Claude Fable serve as occasional second opinions, Sonnet 5.5 and the budget GPT-6 Luna handle scoped subtasks and bulk classification.
Three findings anchor the analysis. First, Opus 5.5 outscores its more expensive sibling Fable 5.1 by 5 points while costing less per task. Second, Sonnet 5.5 at maximum effort costs more per task than Opus at maximum for 2 fewer points, which Meyer argues makes that setting hard to justify. Third, Sol costs roughly one-eighth of Astra and one-twentieth of Fable per task for a score only 1 to 2 points lower.
The effort setting, not the model choice, is the largest cost lever. On Opus 5.5, moving from xhigh to max effort adds 2 index points and 73% more cost per task; moving from medium to max multiplies cost by 4.46x for 7 points. Meyer therefore runs Opus at high (54 points, $1.82) for most development, reserving xhigh (56 points, $3.46) for architecture, migrations and trust boundaries, and treating max effort as rarely worthwhile.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
Why Cost Per Task Now Drives Model Choice
The analysis documents a practical shift for anyone paying for AI-assisted work: when capability gaps between leading models shrink to single index points, the cost difference becomes the deciding factor. Sol’s review pass at $0.32 to $0.39 per task is, in Meyer’s words, cheap enough to run on every meaningful change, which turns independent model review from an occasional luxury into routine practice.
The cross-family pairing matters as much as the price. Meyer argues that a different model family reviewing Opus’s output is a stronger check than Opus reviewing itself, and the low cost of Sol makes that discipline affordable. He also cautions against false economies: halving model price saves only 12.5% of real cost in his illustrative example, and a single extra minute of human review can erase the saving — a figure he presents as illustrative, not measured.
A Month of Consecutive Frontier Releases
September 2026 saw near-continuous releases: Claude Fable 5.1 on 1 September, GPT-6 Astra on 3 September, Opus 5.5 and GPT-6 Luna on 22 September, Claude Sonnet 5.5 on 28 September, and GPT-6.1 Sol on 29 September — one day after the earlier GPT-6 Sol, whose score of 48 Sol matches even at its medium setting at one-fifth of the earlier model’s $1.06 per-task cost.
All capability figures in the analysis come from the Artificial Analysis Intelligence Index v4.3.x, which Meyer describes as a map of general capability, not a verdict on any specific workload. List prices per 1M tokens range from Opus 5.5 at $4/$20 (input/output, with cache reads at $0.20) and Fable and Astra at $10/$50, down to Luna at $0.10/$0.50. Sol is priced at $2/$10, the same as its week-old predecessor.
“In four weeks, the AI frontier stopped being a leaderboard and became a price curve. Six models now sit within about 20 index points of each other, while their cost per task differs by roughly 100x.”
— Thorsten Meyer, ThorstenMeyerAI.com
Benchmark Noise and Missing Settings
Several limits are acknowledged in the source. Artificial Analysis has not yet published low or max effort settings for GPT-6.1 Sol, and Meyer notes that a single index point falls inside measurement noise — meaning Sol’s 1-to-2-point deficit against Astra and Fable may not be meaningful. The human-cost comparison (a single extra minute of review erasing a price halving) is explicitly illustrative rather than measured.
Sol also has confirmed drawbacks at higher effort: 57 to 69 seconds to first token at high and xhigh, which Meyer says rules it out as an interactive model at those settings. All conclusions are based on one benchmark index and one practitioner’s workload; generalising to other tasks, or to teams with different cost structures, is untested in the source.
Shadow Tests and the Full Sol Matrix
The immediate open item is the completion of Sol’s benchmark picture: Artificial Analysis still needs to publish Sol’s low and max effort scores, which could shift where the model sits on the price-performance curve. Meyer’s own stated practice is to shadow-test any candidate model against real workload before switching roles, implying his stack will be re-evaluated as each new release lands — a cadence that September’s near-weekly launches make likely to continue into October.
Key Questions
Which AI model does the analysis recommend as the main working model?
Claude Opus 5.5 at high effort (54 index points, $1.82 per task) for features, APIs, multi-file work and refactors, stepping up to xhigh (56 points, $3.46) for hard problems such as architecture and migrations. Max effort is described as rarely worth the cost.
What is GPT-6.1 Sol’s role, and why was it chosen over similar models?
Sol, launched 29 September 2026, is used for detailed investigation and independent review. At $0.39 per task at xhigh it scores 1 to 2 points below Astra and Fable 5.1 while costing roughly one-eighth and one-twentieth as much, making routine review passes affordable.
Does a higher effort setting make a model smarter?
No, according to Meyer: “effort is not capability.” Higher effort buys a few benchmark points at steep cost — on Opus 5.5, medium to max multiplies cost 4.46x for 7 points — and cannot compensate for missing requirements.
Are the benchmark scores a guarantee of real-world performance?
No. All figures come from the Artificial Analysis Intelligence Index v4.3.x, which Meyer describes as a map of general capability. He recommends shadow-testing any model on your actual workload before switching.
What are GPT-6.1 Sol’s main drawbacks?
At high and xhigh effort it takes 57 to 69 seconds to produce a first token, making it unsuitable for interactive use at those settings. Its low and max settings have not yet been benchmarked, and Opus 5.5 still leads it by 5 points at xhigh.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
