Since I wrote the last newsletter, quite a lot has happened in AI.
Claude Fable came back (in a less capable form). ChatGPT 5.6 was released. Moonshot launched Kimi K3, an extremely capable Chinese model. Microsoft improved Copilot Cowork and added the latest OpenAI models.
There was also plenty of drama.
There have been arguments about whether the most advanced models are too dangerous to release, whether Chinese models represent a cybersecurity threat, whether export controls can slow their development, and whether open-weight models should be available at all.
And after all of that, where have we - the users - ended up?
More or less where we were two months ago.
If you were using Claude two months ago, you can continue using Claude. It is still a very good model. Fable 5 is somewhat more capable and considerably more expensive, but you do not need to redesign your entire AI strategy.
If you were using ChatGPT, you can continue using ChatGPT. GPT-5.6 is better, more persistent and more capable than the previous version. It’s also much cheaper than Fable.
If you were using open-weight or non-US models, you can continue doing that too. Kimi K3 is highly capable and shows that the performance gap between American frontier models and the best Chinese alternatives remains narrow on many practical tasks.
In other words, all providers have moved forward.
This is not a story in which one company suddenly discovered intelligence while everyone else was left behind. Claude improved. ChatGPT improved. Chinese and open-weight models improved.
Despite all the drama, AI development has returned to the same old pattern: the models are gradually becoming more capable, more persistent and more useful.
More capable does not necessarily mean smarter
The most interesting thing about the latest models is that they do not feel dramatically smarter to me.
They feel more relentless.
That distinction is difficult to explain, so here is my favourite analogy.
Imagine giving me a very complicated maths problem. An older AI model is like giving me one hour to solve it. I work for an hour, make some progress, get stuck and eventually come back saying: 'Sorry, I do not know how to solve this.' You might assume that the next generation of AI would be equivalent to replacing me with Monika, who is considerably better at maths. That is not quite what has happened.
Instead, the new model is still more like me, but I am now allowed to work on the problem for two weeks.I can try one approach, realise it does not work, read more material, write some code, check the result, find an error, start again and eventually make a lot more progress. I may still not have Monika's mathematical horsepower, but persistence compensates for a surprising amount of the difference. For the record, this has happened to me with real maths problems. I have spent weeks trying to solve something that Monika understood in a day.
The latest models do not feel like a jump in intelligence. They feel like the same intelligence with much more time, more tools and an unusual willingness to keep going. That may sound like a small distinction, but it changes what AI can do. A model that can work reliably for five minutes can draft an email. A model that can work for an hour can research a market, analyse documents, run calculations and prepare a report.
A model that can work for hours or days can inspect a project folder, reconstruct a timeline, compare the facts with company policies, test different explanations, correct its own work and produce several finished outputs.
It is not necessarily having a brilliant insight. It is simply refusing to stop. This is why the models feel much more capable even when they do not feel obviously more intelligent.
Persistence is expensive
There is, of course, a catch. The longer a model works, the more it reads, writes, searches, calculates and checks. All of that consumes computing resources. That means the cost of completing the task can rise even if the cost of an individual token does not.
Claude Fable 5 is currently priced at $10 per million input tokens and $50 per million output tokens. GPT-5.6 Sol costs $5 and $30 respectively. A good Chinese model can sometimes cost 1-5% of Claude.
This matters because agents may use millions of tokens and make hundreds of tool calls while completing a single piece of work.
Kimi K3 does not need to be better than Claude at every task. It needs to be good enough for a large proportion of tasks while being cheaper, more accessible or easier to control.
For many organisations, that is a much more important competitive advantage than winning another benchmark by two percentage points.
Copilot Cowork is becoming even more useful
The other important development is Microsoft Copilot Cowork. I have been critical of Copilot for a long time. Cowork is different.
It can work across the Microsoft environment, access the files and context it is permitted to use, draft documents, respond to emails, manage calendars and complete longer, multi-step tasks.
GPT-5.6 is now available across Microsoft 365 Copilot, including Word, Excel, PowerPoint, Chat and Cowork. That integration is tremendously helpful. A marginally weaker model connected to your emails, documents, calendar and company systems will often be more useful than the smartest model sitting alone in a separate browser tab.
Unfortunately, Microsoft has also moved Cowork towards usage-based charging through Copilot Credits. Microsoft provides dashboards, budgets, alerts and spending limits, but the basic problem remains: organisations can monitor spending after or during use, yet it is still difficult to know what a particular task will cost before asking Cowork to do it.
The same task can take a different route on different runs. The model may search more sources, make more tool calls, encounter an error, repeat a step or decide that additional analysis is required. It may also use a different underlying model.
That makes budgeting frustrating.
Imagine asking a human analyst to complete a task and being told: 'It might cost £10, it might cost £100, and if you ask me to do exactly the same thing tomorrow it may cost something different.'
The analyst may still be extraordinarily useful, but the finance team will not enjoy the conversation. Honestly, the pricing aspect of working with AI feels like a scam.
We like Cowork and use it. Many of our clients see substantial benefits from it. But the inability to forecast task-level costs is a genuine problem, particularly when companies want to scale successful workflows across hundreds or thousands of employees.
Most real estate work does not need the best model
The model is no longer the main problem
For most real estate companies, model intelligence is no longer the binding constraint. The models are capable enough.
This is where the real competitive advantage increasingly
sits.
The harder questions are now:
- Do they have access to the right information?
- Are they connected to the systems where the work happens?
- Have you designed a sensible process around them?
- Do employees know what to delegate and what not to delegate?
- Can the output be checked efficiently?
- Who remains accountable?
- What happens when the model is wrong?
- Can you explain how the work was produced?
- Can you control the cost?
This is where the real competitive advantage increasingly
sits.
A mediocre model with the right data, context, workflow and supervision can produce excellent work. A frontier model with poor instructions, fragmented information and no governance can produce expensive nonsense.
That is particularly important in real estate.
Our data is messy. Documents are inconsistent. Assets are heterogeneous. Market knowledge is often tacit. Regulations vary by location. A conclusion that is correct for one property may be completely wrong for the building next door.
The solution is not to wait for an even smarter chatbot. It is to build better systems around the intelligence we already have.
The boring phase is the useful phase
After several weeks of model launches, export controls, cybersecurity warnings and arguments about national competitiveness, my conclusion is surprisingly undramatic.
Claude users can keep using Claude.
ChatGPT users can keep using ChatGPT.
Organisations using Chinese or open-weight models can keep using those.
All three routes have improved.
The latest models are more capable, but much of that improvement comes from allowing them to work longer, use more tools and persist through complex tasks. That makes them more useful, but it can also make them more expensive.
For real estate organisations, the priority should therefore move away from endlessly asking which model is number one.
The more valuable work is connecting AI to the right information, building repeatable workflows, allocating different models to appropriate tasks, training people to supervise them and establishing proper governance.
The model race will continue.
But for most businesses, the technology is already good enough.
The question is no longer whether the next model will transform your organisation. It is whether your organisation has learned how to use the models it already has.

