We did not query a database of European property deals, because there was not one. So overnight, the system built one, by reading the news.
Five years of European dealflow does not exist as a table anywhere. It exists as roughly 46,000 news stories, written in ordinary prose for people to read over coffee. By morning that had become about 170,000 structured facts: who bought what, from whom, and at what yield. It had even taken a single “£700m portfolio” headline and split it back out into the individual buildings inside it.
A pivot table cannot read English. That is the part of this that only a language model can do, and it is worth being precise about why, because it is not the part people usually point at.
Extraction is the boring half
Pulling a price out of a sentence is not hard. What is hard is everything around it. The same asset appears under three names across five years. A vendor is described once by its fund vehicle and once by its parent. A figure quoted as a headline price in one story turns up as a net initial yield in another. Half the stories are about deals that never completed.
Structuring that means resolving entities, reconciling contradictions, and knowing when two stories are about one transaction and when they are about two. Every one of those decisions is a judgement made from context, which is exactly the thing a rules engine cannot hold and a language model can.
Then it reasoned over what it had built
One example, anonymised. A London logistics estate, EPC rated E, bought at the top of the market in 2021. Peak pricing meeting a forced energy retrofit bill.
Two separate pressure signals on one building, and both of them were pulled out of a single old sentence that nobody had reason to reread. Neither was a secret. Neither was in any dataset. They were sitting in a trade story from four years ago.
“The deals were always public. They were unreadable at scale, which is not the same thing as unavailable.”
Teddy James, Tercero Analytics
What it refused to do
This is the part we care about most. Where the ownership trail was not watertight, the system declined to name an owner rather than picking the most likely one. And every figure it produced cites the sentence it came from, so any number in the output can be walked back to the story that carried it.
That constraint costs coverage. There are deals in that corpus we could have attributed with reasonable confidence and did not. We would rather hand over a smaller set that holds up under a challenge than a complete one that quietly guesses in the gaps, because the moment one attribution is found to be invented, every other row in the file becomes suspect.