Two new models in one week — how to test them on your own work before you switch
Anthropic shipped Claude Fable 5.1 on 1 September and OpenAI followed with GPT-6 Astra two days later. Both are described by their makers as the most capable models they have built. Neither claim tells you whether your month-end gets better. Here is a two-hour test that does, and a note on what each launch actually gives a business paying US$20 a month, which is less than the headlines suggest.
Two of the three AI tools this column has covered since June got new flagship models within two days of each other. Anthropic released Claude Fable 5.1 on 1 September, describing it and its restricted sibling Mythos 5.1 as its “most advanced models for coding and knowledge work.” OpenAI announced GPT-6 Astra on 3 September, calling it “the world’s most intelligent and aligned model,” and it reached Plus subscribers over the following days. Both companies published benchmark scores that mean nothing to anyone running a business, and both new models now appear in the model picker of the app you already use, with conditions attached that I will come to.
If you have spent the past few months building prompts, templates and small routines around one of these tools, this is the week they may have changed underneath you. That is not a reason to panic and it is not a reason to switch. It is a reason to test, and the test is simpler than it sounds.
Benchmarks are not a purchasing decision
The launch pages read like exam results. OpenAI reports that GPT-6 Astra scores 98 per cent on an advanced mathematics test and finishes computer tasks in 47 per cent less time than its predecessor, GPT-5.6 Sol. Anthropic reports that Mythos 5.1, the same model as Fable 5.1 with fewer restrictions, designed protein binders with a hit rate near 50 per cent against a typical 10 to 15 per cent. Impressive, and irrelevant to whether the model can reconcile a supplier statement without inventing a line.
The claim that matters to a finance team is buried further down in both announcements. OpenAI says Astra produces polished documents, spreadsheets and presentations and handles multi-step workflows; Anthropic says Fable 5.1 is built for knowledge work and long, multi-step tasks. That is a testable claim. The vendors have tested it on their material. You should test it on yours.
The reason is that every business has its own dialect. Your supplier names, your chart of accounts, your habit of writing “GCT incl.” in a description field, the way your bank statement exports with the date in the wrong column. A model that is brilliant on a public benchmark can still be wrong on the fourth line of your aged receivables. The only benchmark that counts is the one you build from your own files.
Same prompt, same files, two models. Five tasks you already do, scored on correctness,
instruction-following and editing needed. (Branded graphic by PGH Consulting, LLC)
The two-hour bake-off
Pick five tasks you already do with an AI tool. Not five things you would like to do one day. Five things you did last week. For most finance teams the list looks something like this.
First, variance commentary. Give the model a budget-versus-actual table for one month and ask for a paragraph per line over a threshold. You already know what the real explanations are, so you can grade it.
Second, a reconciliation. A bank export and a ledger export for the same period, with three known differences you have planted deliberately. Ask it to find the differences and list them.
Third, a customer email. A late-payment reminder to a customer who has been with you for years, in your tone, with the invoice details from a small table.
Fourth, a cash forecast prompt. Thirteen weeks of receipts and payments and a question: where does the balance fall below zero, and by how much?
Fifth, a contract or policy summary. A supplier agreement, an insurance schedule, a lease. Ask for the five clauses that carry money or dates.
Run each task through the model you use today and through the new one, with exactly the same prompt and exactly the same files. Use copies with names changed if the data is sensitive; the point is structure, not secrecy. Then score each answer on three things: was it correct, did it follow the instruction, and how much editing did it need before you could use it. A three-point scale is plenty. Write the scores in a table with ten cells.
You will have your answer in an afternoon, and it will be your answer, not a press release.
A simple test table beats a benchmark: five tasks, three questions each, on a 1 to 3 scale. (Branded graphic by PGH Consulting, LLC)
What the scores usually show
A few things to expect before you start, so the results do not surprise you.
New models are usually better at the open-ended tasks: the email, the summary, the commentary. They read tone more accurately and they follow long instructions more faithfully. They are often no better, and occasionally worse, at the arithmetic tasks, because arithmetic was never the model’s strength and a bigger model does not fix that. The reconciliation is the task most likely to produce a confident, tidy and wrong answer, and it is the one where you should look hardest.
The other consistent finding is that a new model changes the shape of its output. Longer answers, different headings, a summary paragraph where there was none. If you have a routine that pastes model output into a template, a bigger model can break it without ever being wrong. That is the thing to check before you switch anything a colleague relies on.
ChatGPT Plus: Astra included, in the Work and Codex modes, chat still defaults to GPT-5.6 Sol.
Claude Pro: Fable 5.1 visible but needs prepaid usage credits. Per OpenAI and Claude help centres, 10 September 2026. (Branded graphic by PGH Consulting, LLC)
What the launches cost you
This is the part the headlines skip, and it matters for a business budgeting US$20 a month per person, because on both tools the US$20 plan gets the new model with strings attached.
On ChatGPT, OpenAI says Astra is included in Plus, Pro, Business and Enterprise at no change in price, within existing usage allowances. Read the help pages, though, and the Plus entitlement is narrower than the announcement. On Plus, Astra is available in the ChatGPT Work and Codex modes, while the ordinary chat window still defaults to GPT-5.6 Sol. The chat version, called GPT-6 Pro, is reserved for the US$100 and US$200 Pro plans and for Business and Enterprise, with weekly or monthly message caps. On Enterprise it is switched off by default and a workspace owner has to enable it.
On Claude, Fable 5.1 is not included in the standard Pro plan at US$20. A Pro subscriber will see it in the model picker but has to buy prepaid usage credits to run it. It is included in the Max plans at US$100 and US$200 a month, up to half of the weekly usage allowance, and on the premium seats of Team and Enterprise plans on the same terms. Free accounts do not get it at all. For a small business on Claude Pro, the practical position is that the model you had in August is still the model you have, and the new one is a separate purchase.
Neither arrangement is unreasonable, but neither is what a quick read of the launch coverage suggests. A business comparing the two tools on price should compare what US$20 actually buys this month, in the mode it actually uses, rather than what the headline says.
What a model upgrade will not do
Three limits, and they apply to both launches.
A more capable model does not know your business any better than the last one did. Everything it produces is still built from what you give it, and a vague prompt gets a fluent, confident and generic answer from any model at any price.
A model upgrade does not change your data obligations. If you were not putting client data into a personal chat account in August, the same rule holds in September. Both companies still offer business plans with different retention terms from the consumer ones, and the safer place for company work is still a company account.
And both launches carry new safety machinery that can get in the way of ordinary work. OpenAI says Astra’s misalignment monitoring can pause or stop a task it misreads, and that it refuses some advanced security tasks outright. Anthropic says Fable 5.1 restricts categories of cybersecurity and life-sciences work and routes those queries to its Opus models instead. Neither should touch a finance workflow. If one does, a refusal on an innocent task is worth a support ticket, not a workaround.
Write down last week’s five AI tasks, run them through both models, plant three errors in the
reconciliation, check what your plan includes, and check admin settings on a company account. (Branded graphic by PGH Consulting, LLC)
What to try this week
1. Write down the five AI tasks you actually ran last week, with the prompt and the file you
used. If you cannot name five, that is useful information too.
2. Run each one through your current model and the new one with identical inputs, and
score both on correctness, instruction-following and editing needed. Ten cells, one
afternoon.
3. Plant three known errors in the reconciliation test before you run it. A model that finds
two of three is telling you something a benchmark cannot.
4. Check what your subscription actually includes this month. On ChatGPT Plus the new
model is in the plan but only in the Work and Codex modes; on Claude Pro it is a separate
credit purchase.
5. If you administer a company account, find out where the new model sits for your people
before they ask why they cannot see it: an admin toggle on ChatGPT Enterprise, the seat
type on Claude Team and Enterprise. Then decide deliberately whether to turn it on.
Model capabilities, plan inclusions and prices are as published by OpenAI and Anthropic at the time of writing and change frequently; confirm on the vendor’s own pages before relying on them. A new model changes what the tool can do, not the need to check what it did.
Peta-Gaye Hardy is the founder of PGH Consulting, LLC, where she helps finance and operations teams adopt AI in practical, low-risk ways. She writes the weekly AI in Finance & Business column and is based between Jamaica and the United States. Learn more at www.pghconsultinggroup.com. Follow on Instagram and YouTube @pghconsultinggroup, and connect on LinkedIn at linkedin.com/in/peta-gaye-hardy.
Disclosures: This article is informational and does not constitute investment, tax, legal, or accounting advice. Readers should consult a qualified professional before acting. Model capabilities, plan inclusions and prices are as published by OpenAI and Anthropic at the time of writing and are subject to change; readers should confirm the current position on the vendors’ own pages before subscribing. Prices quoted are the published United States figures; local billing may differ. AI tools can produce errors, and every figure, clause, or claim they produce should be verified against a source before being shared or acted upon. The author has no commercial relationship with OpenAI, Anthropic or any product mentioned and was not compensated by them. All product names, logos, and trademarks are the property of their respective owners and are referenced for editorial purposes. The examples described are illustrative and do not depict any real business.