The AI race has shifted again

Sep 15 / Nikodem Szumilo
A year ago, it felt like momentum in AI was moving in one direction: away from ChatGPT and towards Claude. It happened surprisingly slowly.

For months, ChatGPT remained the default because everyone knew it. Then more people started using Claude seriously and realised that, for many types of professional work, it was simply better. It was better at reading long documents, better at writing, often better at strategy and commercial judgement and particularly good when you gave it a complicated task and let it work.

For much of the last year, Anthropic felt like it was leading the AI race. It doesn't feel like that anymore.

I don't mean that Claude has suddenly become bad. It hasn't. Anthropic still makes exceptional models. Indeed, its latest model, Fable 5.1, is very good.
But that is almost the point.

Fable 5.1 is better than Fable 5. It is somewhat more capable, behaves better on long-running tasks and improves the economics of some agentic workloads. But for the sort of professional work we test at VARi, it feels much more like marginal progress than another leap forward.
A year ago, a new Claude release could materially change which model I recommended people use. Fable 5.1 doesn't really do that.
And I think that tells us something important about where the AI market is going.

The smartest model isn't necessarily the model people want to use

This is something I increasingly hear from people who use AI heavily. Claude models are very smart, but people don't always enjoy working with them. They can overthink relatively simple problems. They sometimes make too many independent decisions. They can head surprisingly far down a path before you realise that an early assumption wasn't quite what you wanted. And because the work can be extremely sophisticated, checking what they have done becomes a substantial job in itself.

The original Fable 5 was perhaps the clearest example. It was arguably the most capable widely available model when it launched and was explicitly designed for very difficult reasoning and long-horizon agentic work. It was also expensive.
And businesses did not rush to use it.

Ramp's latest model-level data is fascinating. In its first month, Fable 5 represented only 6% of tokens purchased from Anthropic and 11.4% of Anthropic model spending. GPT-5.6 Sol, by comparison, represented 25% of OpenAI tokens and 23% of spending.

The new Fable 5.1 has only just arrived, so it is too early to know whether adoption will look different. But I suspect the basic problem remains.
We spent the last few years assuming that the most intelligent model would win.

Perhaps that assumption was wrong. For most business work, we don't necessarily want maximum intelligence. We want enough intelligence, combined with predictability, speed, cost efficiency and control. That is particularly true in real estate.

You don't need a frontier model capable of working autonomously for ten hours to take a number from one spreadsheet, check it against a lease and put it into another spreadsheet. In fact, giving that task to a model determined to demonstrate how clever it is can sometimes make things worse.

We increasingly don't use the smartest models

This has probably been the biggest change in our own use of AI over the last few months. The models have become so good that, for the vast majority of real-estate tasks we test internally and with clients, we simply don't need the frontier models anymore.

We increasingly use smaller models. For example, ChatGPT Terra at medium reasoning is remarkably good for a huge amount of professional work. Even when we use more capable models such as Astra or Fable, we often don't want them operating at maximum reasoning effort.

More thinking is not always better.
Give a very capable model an enormous reasoning budget and it sometimes starts solving problems you didn't ask it to solve. It explores alternatives you don't care about, adds complexity you don't need and makes additional judgement calls that then have to be reviewed.
In other words, the extra intelligence can create extra supervision. For many tasks, a slightly less intelligent model with excellent instructions gives us a better result. That changes the economics of AI quite dramatically.

ChatGPT has made a comeback for us

At VARi, we definitely use ChatGPT more than we did a few months ago. Possibly as much as we did a year ago. The GPT-5.6 generation has been excellent. In our real-estate testing, ChatGPT has become very good at combining market research, modelling, asset-level analysis and reasonably sophisticated judgement. On underwriting tasks, for example, it can now identify lease issues, find relevant market evidence, build cash flows, analyse risks and turn the whole thing into an asset-management plan.
Not perfectly. But extremely well.

More importantly, the cost-to-quality ratio is phenomenal. And increasingly I think that is what matters. OpenAI has just released GPT-6 Astra as well. It looks extraordinary on the published benchmarks, but on our real-estate workflows it looks to be offering limited progress. Once exception is 3D and visual reasoning where it offers a very significant advantage over other models (see my LinkedIn post here). It has much better commercial judgement then Claude but it is still not perfect and if you want things done your way, you need to draft detailed instructions (I don’t see how this would ever not be the case) otherwise known as Skills. 

Skills may matter more than another smarter model

One of my favourite developments in ChatGPT is the ability to create reusable skills. That sounds much less exciting than a giant new reasoning model, but I think it is probably more important for business adoption. 

Rather than repeatedly asking an extraordinarily intelligent model to work out how to perform a task, we can tell a cheaper model exactly how we want the task done. Once you've worked out how your investment memo should be structured, how your DCF should treat lease events, how your risk register should work, how your market report should cite evidence or how your due-diligence process should escalate uncertainty, you can encode much of that. Then the problem becomes repeatable.

This is exactly what companies should want. You probably don't want an AI investment analyst reinventing your underwriting process every Monday morning. You want it applying your process consistently and flagging the unusual things that require judgement. The value starts moving away from the model and towards the system around the model.

Google's latest model makes this even more obvious

Google has just released Gemini 3.8 Flash. It is a good model. It is fast, relatively inexpensive, multimodal and very capable at coding, reasoning and agentic tasks. In isolation, it is impressive technology.
But I have to admit that I find the release slightly underwhelming. Why?
Because on broad independent intelligence comparisons, Gemini 3.8 Flash is roughly on par with some of the best open-weight models. That is an extraordinary achievement for the open-source ecosystem.
For Google, one of the world's largest technology companies with effectively unlimited access to talent, data, infrastructure and capital, it is less impressive.

The best Chinese and open-weight models have caught up extraordinarily quickly. Qwen, Kimi, GLM and others are now producing models that are genuinely competitive with the leading proprietary systems across many tasks. Sometimes they are much cheaper. Sometimes you can host them yourself. Sometimes they perform better.

A few years ago, the assumption was that frontier AI would remain concentrated among a tiny number of American labs because nobody else could afford to compete. That assumption is starting to look questionable too. If Google's latest proprietary model is approximately as intelligent as a leading model whose weights you can download, the moat is clearly not simply model intelligence.

This is very good news for open-source AI

I think open-source models may be one of the most important developments for large companies. More organisations we speak to are exploring them, including models developed in China. Some are testing them through model-serving platforms; others are looking seriously at hosting models themselves.

The attraction is obvious.
You control the infrastructure.
You control the data.

You can determine exactly when the model changes—or decide that it shouldn't change at all.
Nobody can suddenly retire the model your workflow depends on.
And increasingly, you aren't accepting a dramatic performance penalty to get those benefits.

For much of real-estate work, these models are already easily capable enough. There is a tendency in AI discussions to focus on the hardest possible tasks: solving new mathematics, autonomous scientific research or extremely complicated software engineering.

Most office work isn't like that.
A very large proportion of real-estate work involves reading documents, extracting information, comparing things, applying reasonably clear rules, producing standard calculations, updating spreadsheets, conducting research, preparing first drafts and checking consistency.
You don't need the world's smartest AI to do that.
You need reliable AI with good instructions.

Microsoft suddenly makes much more sense

The other change is Microsoft. I have complained about Copilot often enough that I'm still slightly uncomfortable writing this, but Copilot Cowork has become a genuinely interesting product. This isn't really because Microsoft suddenly built the world's best model.
It did something smarter. It stopped pretending that one model should do everything. Depending on what your organisation enables, Cowork now gives access to different GPT and Claude models and can choose between them according to the task.

That makes enormous sense.
You don't really care which model drafts the email, searches your files, reconciles spreadsheets or prepares a report. You care that the work gets done. And for many companies, having this inside the Microsoft environment is a huge advantage. Your emails are already there. Your files are there. Your calendar is there. Your permissions, governance and enterprise controls are there. Cowork can execute multi-step tasks across that environment rather than merely drafting a response, and Microsoft has now made it generally available to Microsoft 365 Copilot customers.

If you’d like to learn more about Copilot Cowork – see our starter guide here (use code NEWS40 for a 40% discount). 

The AI race is becoming much less interesting at model level

This leads to what I think is the bigger point. We still spend an enormous amount of time discussing which model is better. Claude or ChatGPT? Fable or Astra?

I increasingly think these are interesting questions for people like us who test models, but less important questions for most real-estate organisations. The best models we have today are capable of doing the vast majority of the information-processing work we do in real estate. They can analyse investment opportunities, build financial models, review contracts, prepare IC papers, conduct market research, analyse portfolios, draft reports, inspect project documentation and help design strategies. Not perfectly. But neither do humans.

The bigger differences now come from what you give the model, how you instruct it, which tools it can access, which workflow surrounds it and how you check the result. That brings us to the real bottleneck.

AI can produce work much faster than humans can check it

Our rough internal experience is something like this:
  • 95% of AI output is essentially flawless.
  • Around 4% is wrong in a way that doesn't really matter.
  • Around 1% is wrong in a way that matters a lot.

Those aren't scientific statistics. They're simply a useful rule of thumb from the workflows we have been building and testing. Unfortunately, you don't know in advance which 1% is the important 1%.

So you review.

Imagine AI drafts a 400-page contract in several hours. Brilliant. A human lawyer still has to read 400 pages.
And not simply skim them. They need to understand them, trace important clauses, think about interactions, identify omissions and decide whether the document actually reflects the commercial agreement. AI generation can happen in hours. Human comprehension still takes days. The same applies to investment.

We have seen workflows where someone preparing an IC memo becomes several times more productive using AI. That is transformational. 

But the Investment Committee doesn't become several times faster at understanding investments.
If you suddenly produce three times as many beautifully researched investment papers, the committee still has to read them, challenge the assumptions and understand each deal in the context of the portfolio.

You haven't removed the bottleneck. You've moved it.

The next frontier is not intelligence. It is review.

This is why I think some of the most important improvements in AI over the next couple of years will be much less glamorous than another enormous model.

We need better ways of checking AI.

Every important number should be traceable to a source. Workflows need systematic testing. Models should flag uncertainty rather than bury it underneath confident prose. Consequential actions need approval checkpoints. We need records of what the AI did, which data it used and which instructions it followed. And human review needs to be proportionate to the consequence of getting something wrong.

This is governance, but not governance in the slightly abstract sense of writing an AI policy and putting it on SharePoint.

It is operational governance
. How exactly do I know this number is right? Where did this assumption come from? Which documents were checked? Can I reproduce the result? Those questions matter much more to me than whether Claude scored two percentage points above ChatGPT on a benchmark.

The AI race is becoming an implementation race

So yes, I think momentum has shifted away from Anthropic. But not because Anthropic suddenly makes bad models. Fable 5.1 is excellent. The problem for Anthropic is almost the opposite: being excellent is no longer enough to feel exceptional.

Its latest model improves on an already extraordinarily capable model, but for normal professional work the progress feels marginal. Meanwhile ChatGPT has become extremely good again. Microsoft has built a genuinely useful multi-model work environment. Google's latest model is good but not obviously more intelligent than the best open alternatives. And Chinese and open-weight models continue to make extraordinary progress.

Model intelligence is becoming commoditised much faster than I expected. I increasingly think human review, not AI intelligence, is the real bottleneck.

Created with