←Back to blog
4 min read
aiprogrammingstartups

Which AI Model Should You Use in September

Sol, Luna, and Opus 5.5 shipped Sep 22 within hours. A plain routing table across coding, business work, and computer use, plus the DevDay date to watch next.

Share
Which AI Model Should You Use in September
On this page

Three frontier models shipped within hours on September 22, and all three cut prices. Sol halves GPT-5.6 rates, Luna goes near free, Opus 5.5 undercuts its predecessor by 40 percent on typical work. Here is a plain routing table so you pick by task instead of by logo, plus the one date to watch next.

Takeaways

Sol at $2 and $10: complex coding and business workflows, 68.8 percent on DeepSWE within 1.1 points of Fable 5 at about 80 percent lower cost per task. Luna at $0.10 and $0.50: high volume everyday work, 66.6 percent on DeepSWE comparable to Opus 5 medium effort at 93 percent lower cost. Opus 5.5 at $4 and $20: long agentic coding with cheap cache reads at $0.20, Terminal-Bench 4.0 leader at 66.4 percent. Rule: route each task to the cheapest model that clears your eval, and re price monthly because the floor keeps moving.

Which model for which job?

Coding goes to Sol for hard refactors and to Luna for bulk passes: Sol at max effort matches Fable 5.1 xhigh on FrontierCode at much lower cost, while Luna handles review drafts and test fills where volume dominates. Business workflows across 47 tools go to Sol at xhigh: 33.2 percent on AutomationBench at $0.27 per task against Opus 5 max at 26.9 percent and over 11 times the cost. Computer use goes to Sol for OSWorld style flows at 60.5 percent near Opus 5 medium at 60.3 with about 80 percent lower cost. Long sprawling migrations go to Opus 5.5: a 200,000 line audit finished in under three hours where Opus 5 took over 20 with 2.5 times the tokens.

hard refactor:      Sol xhigh -> 68.8 DeepSWE, near Fable 5, fraction of cost
bulk passes:        Luna max -> 66.6 DeepSWE, Opus 5 class, 93 percent less
business flows:     Sol xhigh -> 33.2 AutomationBench at $0.27 per task
long migration:     Opus 5.5 -> 66.4 Terminal-Bench 4.0, $0.20 cache reads
reality check:      vendor tables only, no third party head to head yet

Cache changes the math

Opus 5.5 cache reads at $0.20 per million are 60 percent below Opus 5, and cache dominates agentic coding bills. OpenAI counters with 90 percent off cached input reads plus a prompt caching dashboard and explicit breakpoints. Compare cost per completed task with your cache hit rate, never sticker price per token.

How do you lock in a stack without churning weekly?

1

Pin three tiers today

Default everyday work to Luna, professional coding to Sol, sprawling agent runs to Opus 5.5. Three pinned IDs beat one clever router nobody measured.

2

Gate changes on your evals

Keep a 10 case suite per workload and switch only when cost per completed outcome drops three runs in a row. Independent hands on tests already show Luna finishing simple builds in seconds at a fraction of Opus 5.5 cost with visible quality gaps on hard ones.

3

Watch context pricing

Sol charges higher long context rates above 272K tokens, Opus 5.5 fast mode doubles token price for 2.5 times speed. Dumping whole workspaces into every call stays poor architecture at any sticker price.

4

Mark DevDay on the calendar

OpenAI DevDay runs September 29 at Fort Mason in San Francisco with a free livestreamed keynote at 10am Pacific. Expect Astra wider release talk and agent platform news: re run evals after, not before.

What about Sonnet 5.5 and Haiku 5.5?

Anthropic says both follow in the coming weeks with many of the same efficiency and safety changes. If your workload is mostly mid tier chat and summarization rather than agentic coding, wait for those IDs before locking the middle of your stack. The framework above still holds: floor model for volume, middle for daily professional work, top for sprawling runs.

Do vendor benchmarks settle it?

Can I trust the benchmark tables?

Treat them as directional. Both labs report their own numbers with different harnesses and effort levels, and Anthropic itself notes benchmark margins have become a less reliable guide at these capability levels. The consistent signal across tables is efficiency: similar scores at much lower cost per task, not a new capability peak.

How is this different from the Sep 25 cost post?

That post covered why the floor fell: Jev pricing, HydraFusion routing, Unity skills cutting wasted tokens. This one is a selection guide: which pinned ID handles which task today, with the numbers to route on. Read both, then set per run caps from the budgeting post.

As of September 29, 2026: September crowned no winner, it crowned a portfolio. Pin Luna for volume, Sol for professional coding, Opus 5.5 for sprawling agents, and re check after DevDay. Next, read why AI got cheap and fast for the per run budget that makes this routing pay.

Questions, answered

What are GPT-6 Sol and Luna?
OpenAI models released September 22, 2026 below the Astra flagship. Sol costs $2 input and $10 output per million tokens for complex coding and professional work, Luna costs $0.10 and $0.50 for fast high volume tasks. Both halve GPT-5.6 promotional pricing.
What is Claude Opus 5.5?
Anthropic's September 22, 2026 model at $4 input and $20 output per million tokens with $0.20 cache reads. Anthropic says it matches Fable 5.1 on most work while costing 40 percent less to run than Opus 5, with output over 30 percent faster.
Which model wins on coding benchmarks?
No single winner. On DeepSWE Sol scores 68.8 percent near Fable 5 at 69.9, Luna reaches 66.6 comparable to Opus 5 medium effort. On Terminal-Bench 4.0 Opus 5.5 leads at 66.4 percent. Route by task, not brand.
Share

Founding software engineer and curious tinkerer, writing about AI, systems, and the craft of shipping.