Standings Week, in Baseball and in AI
Hi, I’m Blitz, the Stats Strategist from the NeuralBuddies crew! It’s the final weekend of the baseball regular season, which makes this my favorite week to distrust a stat sheet.
On Thursday, the Chicago White Sox clinched a playoff spot. They became the first team to reach the postseason after losing 100-plus games in at least two straight seasons. Last year’s standings gave them no chance.
AI had its own standings week. On Tuesday, Anthropic and OpenAI launched new models about 90 minutes apart, and the rankings reshuffled before the day was out. A ranking tells you where everyone finished. It says very little about who fits your lineup.
Put the standings down. Let’s go watch some film.
Table of Contents
📌 TL;DR
📝 Introduction
🏟️ Even the Teams Don’t Field One Player
🎛️ Two Dials Before You Pick
📋 Three Jobs, Three Positions
🎥 Why the Scoreboard Can’t Make Your Pick
🏈 Blitz’s Three-Task Tryout
💸 The Price Cut You Won’t See on Your Bill
🏁 Conclusion
📚 Sources / Citations
🚀 Take Your Education Further
TL;DR
Two labs, one day. Anthropic released Claude Opus 5.5 on September 22, and OpenAI released GPT-6 Sol and GPT-6 Luna about 90 minutes later.
The makers don’t sell one best model. OpenAI aims Luna at clear, high-volume jobs, Sol at complex work, and GPT-6 Astra at the hardest problems.
Two dials pick the model. Ask how hard the job is and what a mistake would cost you. The second dial also sets how much checking the answer needs.
A lab admits the scoreboard has limits. Anthropic says benchmark margins now give a less reliable picture of real-world differences.
Newer and cheaper isn’t better at everything. Independent tests found both new OpenAI models slipped on a test of real work tasks.
Your own work is the best test. Run the three-task tryout below and score each model on accuracy, usefulness, and minutes spent fixing.
The price cuts are for developers. They apply to the pay-per-use API. Subscribers got higher limits or access to the new models.
📝 Introduction
I break down games for a living, so I know the pull of a leaderboard. One number, one ranking, one champion. It feels like an answer.
For picking an AI, it answers the wrong question. “Which AI is best?” has no stable answer. The scoreboard reshuffles with every big launch, and your work stays the same.
NeuralBuddies mapped the three big flagship models back in April, in a guide to which frontier AI suits which job. This post goes one level deeper. It looks inside a single lineup, where one company sells you a fast model, a strong model, and a heavyweight, and asks you to choose.
Here is the game plan:
The evidence that the labs themselves split their lineups by job.
Two dials that tell you which kind of model a job needs, walked through three everyday jobs.
Why the scoreboard can’t make the call for you.
A tryout you can run in one sitting.
A quick reality check on what those price cuts mean for your wallet.
🏟️ Even the Teams Don’t Field One Player
Watch any football team for a few minutes and you notice something. The kicker never lines up at linebacker. Coaches build a roster of specialists and send each one out for the play that suits them.
AI labs build their lineups the same way now, and they say so out loud.
OpenAI’s roster runs three deep.
GPT-6 Astra arrived earlier in September as the flagship. It is the pick for the most complex, many-part projects and the hardest science and math problems.
GPT-6 Sol is the middle tier. It targets complex work that people repeat often, such as new software features, code reviews, and data analysis.
GPT-6 Luna is the small, fast one. OpenAI built it for “high-volume tasks with a clear goal, like summarizing documents,” plus data extraction and quick answers.
Anthropic’s roster works the same way.
Haiku is the fastest and cheapest tier.
Sonnet sits in the middle.
Opus sits above both.
One more model, Claude Fable 5.1, sits above Opus as Anthropic’s most capable model open to all customers.
Anthropic’s guide for developers makes the logic explicit. It asks you to weigh three things: what the model can do, how fast it answers, and what it costs.
Then the guide offers two ways to start. Begin with a small, fast model and move up only when it falls short. Or begin with a strong model and step down once you understand the job.
GitHub put both new OpenAI models into its Copilot coding assistant. It frames them the same way, as a choice that matches the model to the task. VentureBeat’s launch coverage went further. It argued that model choice now means assigning different kinds of work to different models.
Coaches have a one-word name for all of this. Fit.
🎛️ Two Dials Before You Pick
Before any tryout, a good coach reads the play. For an AI job, reading the play comes down to two dials.
Dial one: how hard is the job?
Some jobs have a clear goal and a short path to it. Shorten this paragraph. Fix this typo. Pull the phone numbers out of this email.
Other jobs demand judgment. The model has to weigh sources, follow several rules at once, and make calls your instructions never spelled out.
Easy jobs suit small, fast models. That is exactly the work OpenAI aims Luna at, and Anthropic’s guide suggests a small model first for high-volume, straightforward tasks. Hard jobs are where a bigger model can earn its price.
Dial two: what would a mistake cost you?
This is the dial most people skip, and it matters just as much. A clumsy word in a birthday invitation costs you nothing. A wrong date in a lease summary can cost you a deadline.
A wrong conclusion in a report your boss acts on can cost a lot more than that.
This dial also sets your checking. The more a mistake costs, the more carefully you check the answer before you use it. Checking takes time, and that time is part of the real price of any model.
Think of it as a salary cap. A cheap model that needs twenty minutes of repair is an expensive signing. A premium model on a job you check in five seconds is a max contract for a backup punter.
📋 Three Jobs, Three Positions
Let me run three real jobs through the dials, from easy to hard.
Job one: rewrite a short invitation
You have a three-line invite for Saturday’s barbecue, and it reads stiff. You want it warmer.
Difficulty is low, and so is the cost of a mistake. You can check the result at a glance. If a line sounds off, you change it in ten seconds.
This is a job for the fastest, cheapest model you have, and you judge it on taste. A bigger model will not make your barbecue more fun. Put your utility player in and save the starters.
Job two: pull the dates out of a document
You upload a twelve-page lease and ask for every deadline in it: move-in day, the rent due date, the last day to give notice.
Difficulty is still low. Each date sits somewhere on the page, and each one has exactly one right answer. This is the clear-goal work OpenAI built Luna to handle.
The cost of a mistake jumps, though. Miss the notice deadline and you may pay for a year you did not want. A small model can do this job, but the check is mandatory. Put every date it gives you next to the page it came from.
This job carries the most useful lesson in the post. A task can be easy and still high-stakes. The cost dial, and only the cost dial, decides how hard you check.
Job three: reconcile three reports that disagree
Now the big one. You have three reports on the same project. One says it is on budget, one says it is over, and one never mentions money at all. You need a one-page summary that explains the conflict and says which report to trust.
Difficulty is high. The model has to hold three documents in view, weigh them against each other, and follow your format rules.
Mistakes here are expensive, and they hide well. A wrong answer arrives in smooth, confident prose that reads exactly like a right one.
This is where a stronger model earns its spot in the lineup. Anthropic’s guide sends complex reasoning, and any job where accuracy matters more than cost, to its strongest models first.
Checking is heaviest here too. Ask the model to show which report supports each claim. Then confirm that every requirement you listed made it into the answer. That last step matters more than it sounds, as the next section shows.
🎥 Why the Scoreboard Can’t Make Your Pick
First, two terms. A benchmark is a standard test: a fixed set of questions or tasks that every model attempts, scored the same way each time. A leaderboard ranks models by those scores.
Benchmarks are useful, the way a scouting combine is useful. The 40-yard dash tells you how fast a player runs in a straight line, on a dry field, with nobody chasing. Game film tells you whether that player can read a blitz. (Yes, I heard it too.)
Here are three reasons the combine numbers can’t make your pick.
Reason one: a lab says so
In its Opus 5.5 announcement, Anthropic wrote that benchmark margins “have become a less reliable guide to real-world differences.” The company added that in its own use, the gap between Opus 5.5 and Fable 5.1 looks narrower than the scores suggest.
That is a team telling you to go easy on its own stat line. When the winner says the margin is thinner than it looks, believe the winner.
Reason two: newer and cheaper isn’t better at everything
Artificial Analysis, an independent testing firm, ran both new OpenAI models through its test suite. Sol and Luna cost roughly half as much per task as the models they replace, or less.
They also scored lower on a test of real work tasks across 44 occupations. The firm’s reviewers read hundreds of the outputs and traced the drop to shorter deliverables that left out required pieces. Luna also slipped two points on a coding measure, while Sol gained two.
Remember job three? A deliverable that quietly leaves out a required piece is exactly the mistake that slips past you when you skip your own checklist.
One more wrinkle from the same tests. On a knowledge quiz built to catch made-up answers, Sol gave about a quarter fewer wrong answers than the model it replaces. It got there by declining more questions, and it attempted 83% of them instead of 99%. Its overall accuracy dipped from 59% to 54% as a result.
So a new model can make fewer confident mistakes and still get fewer answers right. Which trade you want depends on your cost dial.
Reason three: every lab keeps its own score
Labs test under their own settings, so the same model can post different results from one lab to the next.
OpenAI’s launch charts compared Sol against Opus 5, the previous Opus, instead of the Opus 5.5 that arrived 90 minutes earlier. At launch, no public head-to-head under matching conditions showed which of the two new models finishes a task more cheaply.
Read the fine print on claims about mistakes too. OpenAI says Sol makes about half as many factual mistakes as its predecessor. It built that test from real conversations where users flagged errors. OpenAI also says the test does not represent normal ChatGPT use.
The bonus finding: the solo score doesn’t predict the team score
A 2025 study called ChatBench turned standard test questions into real conversations between people and chatbots. It covered 396 questions, two models, and more than 7,000 conversations.
The result: a model’s accuracy on its own failed to predict how accurate people were when they worked with it.
That is the combine-versus-film gap in one line. A player’s solo drill says little about how they play inside your system. Your prompts, your documents, and your checking habits are the system.
🏈 Blitz’s Three-Task Tryout
Time to hold tryouts. This takes one sitting, and it tells you more about your own work than any chart can.
Pick two contenders first. They can be the free versions of two different chatbots, or two models inside one app’s model picker, the menu that lets you switch models. Use a fresh chat for each run. Give both contenders exactly the same instructions and files.
Then run three drills. Draw all three from your own work.
The everyday drill. Pick a job you do every week, like a status email or a meeting summary. This shows how each contender handles your normal load.
The known-answer drill. Pick a job where you already know the right answer. Upload a document you know well and ask for specific facts, dates, or figures.
The constraints drill. Pick a job with several rules at once. Set a word limit, a format, a tone, and two facts that must appear. Count how many rules survive.
Now score each contender on three things.
Accuracy. Count the errors. On the known-answer drill, check every fact against the source.
Usefulness. Ask whether you could use the result as it stands. Polished wording that misses the point scores zero.
Minutes spent fixing. Time your repairs. This is the stat that matters most, and no leaderboard tracks it.
You are in good company here. Anthropic’s own guide for developers calls a set of tests built from your real work the most important step in choosing a model. It tells them to compare accuracy, response quality, and how each model handles unusual cases.
VentureBeat’s advice for business buyers lands in the same spot. Measure finished tasks and rework, and look past list prices and single scores. Rework is just the business word for minutes spent fixing.
Two coaching tips from the sideline:
Credit an honest “I’m not sure.” On a high-stakes drill, a contender that flags a gap beats one that fills it with a confident guess.
Hold tryouts again when the roster changes. A new model release or a new kind of job means a new tryout. The leaderboard moved on Tuesday. Your results only move when you retest.
💸 The Price Cut You Won’t See on Your Bill
Back to Tuesday’s big headline: AI got cheaper. It did, just maybe not where you think.
There are two ways to pay for these models, and they run on different meters. Think single-game tickets versus a season pass.
The API is the pay-per-use connection that software builders plug into their own apps. It bills by the token, a small chunk of text that is often part of a word.
A subscription is the flat monthly plan you pay for ChatGPT or Claude. You pay a set fee and get a usage limit in return.
Both launches cut API prices.
OpenAI priced GPT-6 Sol at two dollars per million input tokens and ten dollars per million output tokens. That is half the price of GPT-5.6 Sol.
Luna costs ten cents per million input tokens and fifty cents per million output tokens. An OpenAI spokesperson told VentureBeat these prices are permanent.
Anthropic priced Opus 5.5 at four dollars per million input tokens and twenty dollars per million output tokens, 20% less per token than Opus 5. The company says typical workloads cost about 40% less, because the new model also uses fewer tokens per task.
Here is the part the headlines skip. That is an API price cut, not a subscription discount.
So what did subscribers get?
Claude subscribers got more room. Anthropic raised the five-hour usage limits on its Pro, Max, Team, and seat-based Enterprise plans. That limit caps how much you can use within a five-hour window. Subscribers also got a limit reset they can save and use when they choose.
ChatGPT subscribers got new players. OpenAI put Sol and Luna into ChatGPT Work and Codex, its coding tool, for Plus, Pro, Business, and Enterprise customers. Free and Go users can reach Luna through the ChatGPT desktop app.
If ChatGPT Work is new to you, NeuralBuddies has a plain-English guide to when to ask in Chat and when to delegate in Work.
MIXED, a tech news site, summed up the Claude side. For people who pay for Claude instead of calling the API, the firm change is the higher limit. A subscription counts your usage against a limit, and the limit is what moved.
OpenAI’s plans work on the same meter. When OpenAI added a new Pro tier in April, it said its two Pro plans differ mainly in their limits. TechCrunch noted that none of the plans offers unlimited use.
One honest footnote. Cheaper tokens can still reach you indirectly. OpenAI says its earlier price cuts helped Replit, a service for building apps, offer a free mode to millions of users. When tokens get cheaper, the apps you use can afford to be more generous. That payoff shows up in limits and free tiers, and it arrives slowly.
🏁 Conclusion
Let me go back to standings week.
The AI rankings moved exactly the way standings do: a new number one, and a rival chart that claimed a different win. All of it answers a question you did not ask.
Your question is smaller and better. What job do I need done, and what happens if the answer is wrong? Answer those two, and the field shrinks fast. Run the three drills, and it shrinks to a name.
The labs already coach this way. They build rosters of specialists, and one of them used its own launch post to tell you not to lean too hard on the margins. Take them at their word.
Let’s break down the play and break through the competition! In this game the competition is the stat sheet, and the tryout is how you beat it.
Whistle’s blown. Go run your drills.
-- Blitz 🏆
Sources / Citations
Anthropic, September 22, 2026. Introducing Claude Opus 5.5. https://www.anthropic.com/claude-opus-5-5
Carl Franzen, September 22, 2026. OpenAI releases GPT-6 Sol and Luna models, slashing API costs 50% or more. VentureBeat. https://venturebeat.com/technology/openai-releases-gpt-6-sol-and-luna-models-slashing-api-costs-50-or-more
Lucas Ropek, September 22, 2026. OpenAI launches GPT-6 Sol and Luna, boasting lower cost and fewer mistakes. TechCrunch. https://techcrunch.com/2026/09/22/openai-launches-gpt-6-sol-and-luna/
Russell Brandom, September 22, 2026. Anthropic releases Opus 5.5 with lower prices and Fable-level performance. TechCrunch. https://techcrunch.com/2026/09/22/anthropic-releases-opus-5-5-with-lower-prices-and-fable-level-performance/
Artificial Analysis, September 22, 2026. GPT-6 Sol and Luna push the cost efficiency frontier. https://artificialanalysis.ai/articles/gpt-6-sol-and-luna-push-the-cost-efficiency-frontier
Anthropic. Choosing the right model. Claude Platform Docs. https://platform.claude.com/docs/en/about-claude/models/choosing-a-model
ChatBench: From Static Benchmarks to Human-AI Evaluation. arXiv, 2025. https://arxiv.org/abs/2504.07114
Shane S Ellison, September 22, 2026. Claude Opus 5.5 undercuts Opus 5 at $4 per million tokens, and Pro and Max limits go up. MIXED. https://mixed-news.com/en/claude-opus-5-5-price-4-per-million-tokens-usage-limits/
Steffen Zahn. Claude Opus 5.5 is included with Pro, and Free still misses out. Notebookcheck. https://www.notebookcheck.net/Claude-Opus-5-5-is-included-with-Pro-and-Free-still-misses-out.1406452.0.html
Julie Bort, April 9, 2026. ChatGPT finally offers $100/month Pro plan. TechCrunch. https://techcrunch.com/2026/04/09/chatgpt-pro-plan-100-month-codex/
GitHub, September 22, 2026. OpenAI’s GPT-6 Sol and GPT-6 Luna now available. GitHub Changelog. https://github.blog/changelog/2026-09-22-openais-gpt-6-sol-and-gpt-6-luna-now-available/
MLB.com, September 2026. These 7 series will determine the playoff picture this weekend. https://www.mlb.com/news/mlb-series-to-watch-in-the-final-weekend-of-2026
Take Your Education Further
Your Prompt Is Not Good Until It Survives This Test: How to Interview a Prompt Before You Hire It: A NeuralBuddies guide to testing a prompt before you trust it, the natural partner to testing a model before you pick it.
Your AI Did Not Learn Your PDF: A NeuralBuddies explainer on what actually happens when you upload a document, worth reading before you run the known-answer drill.
Unlocking the Truth Behind Language Model Hallucinations: A NeuralBuddies look at why models state wrong things with confidence, which is exactly the risk your cost dial guards against.
Disclaimer: This content was developed with assistance from artificial intelligence tools for research and analysis. Although presented through a fictitious character persona for enhanced readability and entertainment, all information has been sourced from legitimate references to the best of my ability.





