A few years ago we changed how loyalty worked. The old setup was a physical card, messy by nature: people lost them, forgot them, and a household would end up with several, one per family member, each capturing a slice of the same weekly shop. The new setup was a single digital account, shared across the household through one login.
The data got more accurate. It also got much harder to use.
Customers who had looked infrequent became frequent. Spend per customer went up. Loyalty metrics rose across the board. None of it was real in the way it looked. In most cases behavior had not changed, the measurement had. Physical cards had scattered one household's spending across several low-value IDs, and the digital account pulled it back into one. A year-over-year read showed customers getting more loyal. What actually happened was that we had finally started counting them properly.
There was a further wrinkle that only someone who was there would know. Alongside the personal cards there was a general store card, the one a cashier would scan when a customer had forgotten theirs but still wanted the loyalty promotion. That single ID was not a customer. It was two different populations wearing the same face: loyal, frequent shoppers who had left their card at home, and one-off shoppers who never had a strong enough reason to sign up for their own. Treat the general store card as one customer in your model and you have averaged together the two groups you would most want to keep apart.
For the analysts and data scientists who lived through the change, none of this was a problem. They knew the migration had happened, they knew when, and they knew the transition was not a single day but a long tail that ran for months. Ask them to compare this year to two years ago and they would correct for it on instinct, or tell you which slices were safe to trust. For the people who joined afterward, that knowledge was invisible. The numbers gave them no hint. They had to go and ask the ones who were there.
That is the whole essay in one example. The context that decided whether a piece of analysis was sound or misleading was never in the data. It lived in people.
I have written before about the kinds of context a dataset carries: what the data is, where it came from, and what it is fit for. This is a different question. Not what kinds of context exist, but how much any given use actually needs. And that is not a property of the dataset. It is a property of the dataset and the person, or system, using it.
Context requirement scales along two axes. The first is distance from where the data was made. A small team that runs an operational process and reports on its own numbers needs almost no written context, because the people involved are the source: they know the quirks, the limits, and the cases where the data lies. Move that data to another team, another business unit, another entity, and the need climbs with every step, because each step is one more group that was not there when the data got its meaning. Time works the same way. An analyst who joined after the migration is as far from it as a colleague in another market who never saw it. You were not in the room when the number came to mean what it means.
The second axis is depth of use. Simple reporting on data from its own source needs little. Reconciling data across brands needs a great deal, and it rarely goes the way the integration plan assumed. If one brand ran plastic loyalty cards and another ran a digital system from the start, combining the two is not a matter of matching columns. It is knowing that the two numbers were produced by different worlds, and what that does to any metric you build on top of them. Get that context wrong and the combined view is worse than no view at all.
One more thing raises the bar, and it cuts across both: what the decision costs if it is wrong. A team reporting its own numbers needs little context, until those same numbers feed a regulatory filing or a strategic decision that cannot be undone. Low distance and shallow use do not make a number safe. High stakes pull the context requirement up on their own, however close the data sits or however simple the use looks.
For as long as I have worked with data, that missing context was quietly filled in by people. Someone always knew, or knew who to ask. It rarely had to be written down, because the organization carried it in the heads of the people who had been there. That worked, though it was never optimal, as long as a person sat between the data and the decision.
AI removes that person. A model inherits the numbers and none of the memory. It does not know the migration happened. It cannot tell which ID was the general store card and that it carried two populations. And unlike a new analyst, it will not think to ask, because nothing in the data or the metadata hints that something is missing or off. A capable agent can flag a gap it can detect. It cannot flag context that was never written down anywhere at all. The context that lives in people is exactly what a model needs, and exactly what nothing in the system will prompt it to ask for.
I am not the only one landing here. Anindita Misra made a version of this point recently: AI does not invent enterprise knowledge, it inherits it, and an agent working without that inherited judgment is confident without institutional memory. She frames what follows as a question of accountability: if your agent makes a decision that matters tomorrow, can you name the person responsible for the knowledge behind it? That is a fair question. It also sits on top of an earlier one that most of these conversations skip. Before you ask who owns the context, ask how much context this use actually needs. The answer is not the same for every use, and building as if it were is the expensive part.
This is why every vendor now has a semantic layer, an ontology, or a knowledge graph to sell you. The instinct behind the pitch is right: AI does raise the bar on context, sharply. The conclusion drawn from it is usually wrong. You do not answer the question by building the graph for everything up front.
The level of context a number needs is set by who reads it, and why. Take a customer lifetime value score: same field, same euros. For a grocery marketing team it is close to fine on its own, given a clear definition and how current it is, because they fill in the rest themselves. A marketer knows the same value means opposite things for a one-person household and a family of five. For the single person it is a loyal, high-spend customer. For the family it is a household buying most of its groceries somewhere else. That reading is already in their head, and they know instinctively which use-cases need that context and which can go without.
Give the same score to an AI system and none of it comes along. It has no idea what business specifics decide how that number should be read. Left to its own assumptions it treats the score as a generic value metric, an e-commerce-style number where bigger is simply better, and ranks the family of five equal to the single-person household every time. Same data, same field, two readers, two entirely different context needs.
Encoding that household rule, so the agent applies it every time, is exactly what a semantic layer or an ontology is for. The challenge isn't the tools. It's reaching for them before you know which use-case needs which context.
So the foundation and the frills are not the same job. The foundation is worth building once and for everything: technical and business metadata on every meaningful dataset. What the field is, what it means in business terms, where it came from, how fresh it is, who owns it. That floor is universal and it is not optional.
Everything above it is where the discipline comes in. A semantic layer, an ontology, and a knowledge graph are not three words for the same thing, and not three grades of the same thing either. They solve different problems and cost very different amounts to run. A semantic layer, which maps your tables to agreed business metrics, sits close to the metadata floor. An ontology, a formal model of how your business concepts relate, is a serious and separate piece of modeling. A knowledge graph is that model actually populated and kept current, and it is the heaviest commitment of the three, because someone owns its accuracy for as long as it runs. You take them on use-case by use-case, only as far as the business value justifies. Most of your data never needs more than the floor. A few high-stakes, deeply combined use-cases justify the rest. Building all of it for the whole organization in advance means paying graph-level cost for reporting-level needs, and pretending you can predict every use-case before you have it. You cannot.
The loyalty migration is usually remembered as a data quality win, and it was one. We finally saw what customers were really doing. But the win came with a liability nobody logged: every analysis that spanned the change was now a trap, and the only thing standing between the organization and a bad decision was someone who remembered. That safeguard was never in a system. It walked a little further out the door with every person who left, and stood a little further away with every person who joined.
No knowledge graph would have caught that. But someone writing down what the general store card actually was, for the one use-case that needed it, would have. The question worth asking is not which context system to buy. It is the context your organization takes for granted, because the people who hold it have always been there to ask. That is the context that leaves when they do. And it is the context your AI was never going to think to ask for.