Which AI Model Should You Use in September
Sol, Luna, and Opus 5.5 shipped Sep 22 within hours. A plain routing table across coding, business work, and computer use, plus the DevDay date to watch next.
On this page
Three frontier models shipped within hours on September 22, and all three cut prices. Sol halves GPT-5.6 rates, Luna goes near free, Opus 5.5 undercuts its predecessor by 40 percent on typical work. Here is a plain routing table so you pick by task instead of by logo, plus the one date to watch next.
Takeaways
Sol at $2 and $10: complex coding and business workflows, 68.8 percent on DeepSWE within 1.1 points of Fable 5 at about 80 percent lower cost per task. Luna at $0.10 and $0.50: high volume everyday work, 66.6 percent on DeepSWE comparable to Opus 5 medium effort at 93 percent lower cost. Opus 5.5 at $4 and $20: long agentic coding with cheap cache reads at $0.20, Terminal-Bench 4.0 leader at 66.4 percent. Rule: route each task to the cheapest model that clears your eval, and re price monthly because the floor keeps moving.
Which model for which job?
Coding goes to Sol for hard refactors and to Luna for bulk passes: Sol at max effort matches Fable 5.1 xhigh on FrontierCode at much lower cost, while Luna handles review drafts and test fills where volume dominates. Business workflows across 47 tools go to Sol at xhigh: 33.2 percent on AutomationBench at $0.27 per task against Opus 5 max at 26.9 percent and over 11 times the cost. Computer use goes to Sol for OSWorld style flows at 60.5 percent near Opus 5 medium at 60.3 with about 80 percent lower cost. Long sprawling migrations go to Opus 5.5: a 200,000 line audit finished in under three hours where Opus 5 took over 20 with 2.5 times the tokens.
hard refactor: Sol xhigh -> 68.8 DeepSWE, near Fable 5, fraction of cost
bulk passes: Luna max -> 66.6 DeepSWE, Opus 5 class, 93 percent less
business flows: Sol xhigh -> 33.2 AutomationBench at $0.27 per task
long migration: Opus 5.5 -> 66.4 Terminal-Bench 4.0, $0.20 cache reads
reality check: vendor tables only, no third party head to head yetCache changes the math
Opus 5.5 cache reads at $0.20 per million are 60 percent below Opus 5, and cache dominates agentic coding bills. OpenAI counters with 90 percent off cached input reads plus a prompt caching dashboard and explicit breakpoints. Compare cost per completed task with your cache hit rate, never sticker price per token.
How do you lock in a stack without churning weekly?
Pin three tiers today
Default everyday work to Luna, professional coding to Sol, sprawling agent runs to Opus 5.5. Three pinned IDs beat one clever router nobody measured.
Gate changes on your evals
Keep a 10 case suite per workload and switch only when cost per completed outcome drops three runs in a row. Independent hands on tests already show Luna finishing simple builds in seconds at a fraction of Opus 5.5 cost with visible quality gaps on hard ones.
Watch context pricing
Sol charges higher long context rates above 272K tokens, Opus 5.5 fast mode doubles token price for 2.5 times speed. Dumping whole workspaces into every call stays poor architecture at any sticker price.
Mark DevDay on the calendar
OpenAI DevDay runs September 29 at Fort Mason in San Francisco with a free livestreamed keynote at 10am Pacific. Expect Astra wider release talk and agent platform news: re run evals after, not before.
What about Sonnet 5.5 and Haiku 5.5?
Anthropic says both follow in the coming weeks with many of the same efficiency and safety changes. If your workload is mostly mid tier chat and summarization rather than agentic coding, wait for those IDs before locking the middle of your stack. The framework above still holds: floor model for volume, middle for daily professional work, top for sprawling runs.
Do vendor benchmarks settle it?
Can I trust the benchmark tables?
Treat them as directional. Both labs report their own numbers with different harnesses and effort levels, and Anthropic itself notes benchmark margins have become a less reliable guide at these capability levels. The consistent signal across tables is efficiency: similar scores at much lower cost per task, not a new capability peak.
How is this different from the Sep 25 cost post?
That post covered why the floor fell: Jev pricing, HydraFusion routing, Unity skills cutting wasted tokens. This one is a selection guide: which pinned ID handles which task today, with the numbers to route on. Read both, then set per run caps from the budgeting post.
As of September 29, 2026: September crowned no winner, it crowned a portfolio. Pin Luna for volume, Sol for professional coding, Opus 5.5 for sprawling agents, and re check after DevDay. Next, read why AI got cheap and fast for the per run budget that makes this routing pay.