News AI in real estate

The data that makes AI work in property isn’t clean, and that’s actually the point.

The best signal in any deal is locked inside inconsistent spreadsheets and ageing PDFs. An AI system that only works on tidy data is an AI system that only works in demos.

Teddy James · · 6 min read

Every property team we talk to has the same spreadsheet. It lives somewhere on a shared drive, it has been touched by six different people over three years, and no two columns use the same naming convention. The rent roll for one asset calls it “Tenant Name.” The one from the asset across the road calls it “Occupier.” The one inherited from the previous manager just says “T.”

This is not an unusual situation. It is the normal situation. And it is exactly where an AI engagement either proves its worth or falls apart entirely.

The clean-data fantasy

There is a version of the AI pitch that goes: “Once your data is clean, the AI will unlock serious value.” We hear this framing a lot. It sounds reasonable until you ask the obvious follow-up: when, exactly, does the data get clean? Who is cleaning it? At what cost, and to what end?

In practice, data cleaning is either already happening as a continuous operational task, or it is not happening at all. The firms that have been running disciplined data hygiene for years tend not to need much convincing about AI. They already have a fairly good picture of their portfolio. The firms that most need an AI system to help them see across their assets are, almost by definition, the ones whose data is messy.

The interesting question is not “how do we clean the data first?” It is “what can we extract from the data as it actually exists?”

“The data is never ready. The question is whether your AI system is.”

Teddy James, Tercero Analytics

What the data actually looks like

We did a data review for a firm managing eleven assets across two geographies. Between them, they had rent rolls in four different formats. Some were Excel files with colour-coded cells carrying meaning (which no extraction tool could read). Some were exports from a property management system that had been discontinued. One was a PDF that had been scanned rather than exported, so it was essentially a photograph of a table.

The occupancy figure looked different in every source. One system reported it by unit count. Another by square footage. A third used a blend that nobody could quite explain but that appeared to be weighted by estimated market rent. The same building, queried across three sources, gave three different occupancy numbers.

None of this was anyone’s fault. It was the accumulated result of sensible local decisions made by different people over time. The previous asset manager used the system they were trained on. The data analyst built the Excel model that answered the questions the investment committee was asking at the time. The fund administrator exported what the reporting template required.

The data was not corrupt. It was rich. It just required a system capable of reasoning about inconsistency rather than one that assumed it away.

Why messiness is actually signal-rich

Here is the part that took us some time to articulate properly: the inconsistency in the data is itself information. When an occupier’s name is spelled four different ways across four systems, that is a sign that the occupier has been around long enough to accumulate history. When a column suddenly changes format halfway down a spreadsheet, that is often a sign that ownership or management changed, which is exactly the kind of transition event that matters for understanding an asset’s trajectory.

An AI system built to extract clean numbers from clean data will fail on this kind of input. An AI system built to reason about the structure of the data, asking why a column is the way it is and cross-referencing across sources rather than accepting any single source as authoritative, can pull genuinely useful signal from material that would otherwise be binned.

The practical difference is not just technical. It changes what kind of question you can ask. Instead of “what was our occupancy rate last quarter?” you can start to ask “which assets have had consistent occupancy despite inconsistent reporting, and what does that tell us about the underlying covenant?”

What a data-first engagement actually looks like

Before we scope any engagement, we ask to see the data. Not cleaned data, not prepared data: whatever actually exists. We run a short review: what formats are in play, where the meaningful inconsistencies are, what can be extracted with high confidence and what needs human validation.

This tells us two things. First, it tells us what the AI system needs to be able to handle. Second, it tells us where the genuine value is, because the questions worth asking are usually the ones that the existing data makes hardest to answer.

The firms that get the most from AI in property are not the ones with the cleanest data. They are the ones who stopped waiting for the data to be clean and started building systems that could work with what they had.

If you are managing a portfolio and your first instinct when someone says “AI” is to think “but our data is a mess”, that instinct is not a blocker. It is the starting point. The shape of the mess usually tells you more than the clean version ever would.