AI revenue forecasting sounds like a job for the biggest model you can find. Give it your numbers, ask what next month looks like, get an answer. We tested that idea against Scenario's own forecast engine, lost the first round to a free model, rebuilt the engine, and then beat OpenAI's GPT-6-Astra on the measure that had beaten us. This is the whole story, including the bit where we lost.
How the test works
Forecasts are easy to make and hard to check. Anyone can say "next month will be £12,000". The question is how often that kind of statement turns out right, and you only find out by testing on months the forecaster hasn't seen.
So the benchmark replays businesses month by month. At each point, everyone gets only the data that existed at that moment: the subscriptions, the invoices, the payments so far. Nobody can peek at what happened next. Then everyone forecasts revenue 30 days ahead, and we compare each forecast with what really happened.
The businesses are simulated: ten made-up subscription businesses with realistic patterns (growth, cancellations, failed payments, the odd bad month), generated from a fixed seed so the test can be repeated exactly. Simulated data is a limitation, and we'll say so every time we quote a figure. It's also the only way to run a fair test before you have years of real customer history.
Two simple yardsticks run alongside, because any forecast worth having should beat them:
- Nothing changes: next month equals this month.
- The recent trend carries on: whatever the last few months did, keep doing it.
Round one: a free AI beat us
On 27 September we ran the first version of the engine against gpt-oss-120b, a free, open model, given exactly the same figures.
Some of it went well. Across three runs and 528 forecasts, the real revenue landed inside Scenario's forecast range 84% of the time, against 50% for the AI. And Scenario gave a usable forecast every time, while the AI failed to produce one 13% of the time.
But on the number most people care about, how close the forecast lands, the free model won. It landed within 5% of the real figure 54% of the time. Our engine managed 49%.
That stung, and it was useful. A forecast engine that can't beat a general-purpose chatbot on its own job needs work, whatever else it does well.
What we changed
Scenario's forecast doesn't ask anyone to guess a total. For a subscription business it works from the bottom up. For every subscription it asks two questions: will this invoice get paid, and will this customer renew? It learns both chances from the business's own history. Then it adds new customers at the pace the business has recently been winning them, and blends the result with the two yardsticks above, in proportions learned from what has worked best for that business.
Version two tightened how those chances are learned and how the blend is chosen, so the engine leans on whichever approach has actually predicted that business well, month by month. Same principle, better calibrated.
Round two: Scenario vs GPT-6-Astra
On 29 September we tested the new engine against GPT-6-Astra, OpenAI's frontier model, at medium reasoning. Same ten simulated businesses, same replay, same figures for both. One run, 104 forecasts.
| Measure | Scenario | GPT-6-Astra |
|---|---|---|
| Within 5% of the real figure | 56.7% | 49.0% |
| Closer to the real figure, head to head | 55.8% | 44.2% |
| Real figure inside the stated range | 86.5% | 80.8% |
| Average miss | 7.9% | 8.4% |
In plain English: Scenario's forecasts landed within 5% of the real figure 57% of the time, vs 49% for GPT-6-Astra given the same data. Head to head, Scenario's forecast was the closer one in 56% of cases.
It wasn't a landslide, and we won't pretend it was. The average miss was close: 7.9% against 8.4%. GPT-6-Astra is a genuinely capable model, and it did far better than the "nothing changes" yardstick. But on the measure that beat us two days earlier, the engine went from 49% to 57%, and the frontier model came in at 49%.
One run, 104 forecasts on simulated businesses, 29 September 2026. Your results will vary.
Why a calculator can beat a frontier model
It isn't that the AI is bad at maths. It's that forecasting revenue isn't really a language problem.
A language model reads your figures and writes the most plausible continuation. It's very good at that, and it will often land somewhere sensible. But it isn't learning each customer's chance of renewing from your history, it isn't testing itself on your past months, and it gives a slightly different answer each time you ask.
A forecast engine does the boring parts properly. It counts every subscription. It measures how often invoices really get paid. It checks which method would have worked on your last few months before trusting it with the next one. And it gives the same answer for the same data, every time, which matters when you're about to make a decision on it.
That's why Scenario is built the way it is: the AI understands the question, and the engine does the maths. You can ask in plain English, "what does next month look like if I lose my two biggest customers?", and the numbers in the answer still come from the calculation, not from a guess.
What the numbers don't say
A few honest limits:
- Simulated businesses aren't your business. Real businesses have launches, bad weeks and one customer who pays annually in March. The engine can't foresee a change nobody told it about, which is why you can tell Scenario about launches, holidays and quiet seasons.
- One run is one run. The GPT-6-Astra result is a single run of 104 forecasts. We'll rerun it, and we'll publish the result either way.
- Better isn't perfect. Hitting within 5% more than half the time is good for a 30-day revenue forecast. It still means a range matters more than any single number.
The takeaway
If you're forecasting revenue with a chatbot, it will give you a sensible-sounding number. Ask it for a range, ask how it would have done on your last three months, and ask it again tomorrow to see if the answer moves. Those three questions tell you more than the forecast itself.
If you'd rather see it done on a real set of numbers, the live demo runs Scenario's forecast on a sample business, no sign-up needed. And if a price rise is the decision on your mind, the price increase calculator is free.
See it on your own numbers.
Scenario connects to Stripe or your accounting software, forecasts next month and answers what-if questions in plain English. Free to start.
Scenario’s figures are estimates from your data and stated assumptions, not financial advice. Scenario is not affiliated with Stripe.
