Rubbish In, Rubbish Out: Why AI Is Only as Good as Your Data Layer
    Back to BlogMarket Analysis

    Rubbish In, Rubbish Out: Why AI Is Only as Good as Your Data Layer

    Luke LiplijnAugust 20, 2026

    Everyone wants to do something with AI, yet almost no one has their data in order. Why the data layer beneath the model is the scarce asset — in property markets and far beyond.

    Walk into ten boardrooms today and nine of them have an AI agenda. Copilots, agents, generative models in client contact. The ambition is there, and usually so is the budget.

    What is usually missing is a foundation strong enough to carry it.

    Data sits scattered across systems that were never designed to talk to each other. Definitions differ per department — "client" means one thing in sales and something else in finance. History is incomplete or impossible to reproduce. Ownership is unassigned. Data quality is measured when something breaks, not as routine. Governance exists as a document, not as a practice.

    Then the model arrives. And the model does exactly what it always does: it amplifies whatever you feed it.

    The oldest law in model building

    We have been building with data for more than twenty years. In every model we have seen — statistical, machine learning, generative — the same law holds: rubbish in, rubbish out; diamonds in, diamonds out.

    A mediocre model on excellent data beats an excellent model on mediocre data. Almost always. That is not an opinion; it is the lived experience of anyone who has ever had to explain a production model to a client who did not trust the answer.

    The consequence is badly underestimated. Models have become a commodity: available to everyone, at falling cost, with shrinking differences between them. What is not interchangeable is the data layer underneath — the infrastructure, the quality, the definitions, the history, the ownership. That is where the scarcity is. That is where a defensible position lives.

    Why data management is a strategic asset, not a cost line

    In many organisations data management sits under IT operations. It is a cost line, not value creation. That is a historical mistake, and it is now becoming painfully visible.

    Three reasons we treat it as an asset:

  1. It cannot be copied. A competitor can buy the same model tomorrow. They cannot buy fifteen years of clean, linked, documented history tomorrow.
  2. It sets the ceiling on every AI investment. Skip the data layer and you pay twice: first for the project that does not work, then for the foundation that has to be built anyway.
  3. It has become a governance requirement. Traceability, explainability and demonstrable quality used to be good practice. Under today's supervisory reality they are a condition for operating at all.
  4. Data management is not a precondition for an AI strategy. It *is* the AI strategy — in the only part you actually control.

    What this means for property data

    Property is one of the clearest examples of the problem. Price indices are published by different agencies on different lags, with different baskets and different revision policies. Transaction records are fragmented across land registries. Rental evidence is thin, and yields are quoted without stating whether they are gross or net. Anyone building an AI assistant on top of that mess will get confident answers that are quietly wrong.

    Our approach is the opposite order. First the linked data layer: sourced, versioned, date-stamped series from national statistics agencies, central banks and land registries, harmonised into one coherent set of definitions. Only then the decision layer on top — rankings, comparisons, indexation and chat-based analysis. Phase one without phase two is infrastructure without margin. Phase two without phase one is exactly the mistake the rest of the market is making right now.

    That is also why every figure on this platform carries a source, a reference period and a release date, and why we would rather leave a value blank than fill it with a plausible estimate.

    Part of a wider data group

    Global Property Search is part of Liplyn Group, which invests in established data businesses: companies with real revenue today, long-standing enterprise clients, and services in data management, data infrastructure, data quality and decision information. Several of those operating companies have been running for fifteen years or more — that is not nostalgia, it is accumulated data, accumulated client relationships and accumulated craft that cannot be recreated in two years.

    The market currently pays a great deal for the model and very little for unique data. That is precisely the kind of asymmetry that returns come from. We buy the undervalued part: the craft, the infrastructure and the datasets that every AI application ultimately has to land on. More about the group and its portfolio is available at LiplynGroup.com.

    What to ask before your next AI project

  5. Can you reproduce the number your model used, six months from now?
  6. Does every key metric have one owner and one definition?
  7. Do you know the release lag and revision policy of every external source you rely on?
  8. Would you be comfortable showing a regulator how an answer was produced?
  9. If the answer to any of these is no, the highest-return investment available to you is not another model. It is the layer underneath it.

    Share this article:
    View all articles