Every data platform now promises to learn your business by reading the traces it leaves behind: the schemas, the queries, the dashboards. The tension I keep coming back to is that a system of record captures a representation of reality, and a representation is not a definition. The pitch is seductive, and I feel the pull as much as anyone, because it offers the thing most of us want, which is a semantic layer without the meetings. As discovery, it genuinely works. Where we slip is the promotion, the moment we treat what the machine observed as what the business means.

Schemas are decisions, not definitions

A relational schema tells you how an application chose to persist something, not what that something means. Surrogate keys, junction tables, status codes, audit columns: those are implementation decisions, made under a deadline, often by a vendor who never met your business. Kimball drew the line decades ago, gathering business requirements from the business and treating source systems as a separate question of feasibility. Inference collapses the two, and most of us have helped it along, usually because the schema was sitting right there and the business stakeholder was not. Collapse them and three things go missing.

Meaning fragments. Customer lives in sales, billing, support, and fulfillment, and no single system holds the whole concept. An engine reading four representations finds four concepts, or worse, fuses four deliberately different meanings into one.

The business definitions that matter most were never columns. What makes a customer active, which orders count toward revenue, who is eligible for what: the business runs on definitions and rules no application ever stored, and we cannot extract what was never persisted.

Old workarounds calcify into knowledge. Infer an ontology from twenty years of applications and you institutionalize twenty years of compromises, including the ones we made ourselves. The most-used dashboard in the company can carry a revenue definition somebody invented locally years ago, and heavy usage only proves that the dashboard is popular. It does not prove that the definition is right.

Agents raise the price of being wrong

This mattered less when the consumer of inferred semantics was a chart, because a person reads a chart, and a person can smell a wrong number. An agent does not read. It acts, at machine speed, and it acts with confidence, applying a revenue definition somebody invented locally years ago as if it were policy. So the cost of an unexamined definition has changed: it used to be an argument in a meeting, and now it can be a refund issued, a forecast filed, or a customer contacted, all downstream of a concept nobody ever agreed on.

How DeltaVault lets evidence propose and the business decide

None of this argues against inference. It argues about direction. Observation is a superb way to discover candidates for your business ontology and a terrible way to legislate it. Your systems are witnesses, and good ones, but they don’t get to hold the pen.

That distinction has shaped how we’re approaching things in DeltaVault, and I offer it here as one attempt at the pattern rather than the answer to it. The business model gets authored deliberately: entities, attributes, and relationships on a live diagram, with definitions that move from draft to approved under an audit stamp. Discovery then works for that model instead of replacing it. An AI-led interview proposes candidate entities onto the diagram as ghost nodes, and you accept or dismiss them one at a time. The Map & Match workbench reads your source tables and suggests which business entity each one represents. Every suggestion carries a confidence score and a written rationale, and nothing changes until a person accepts it. For the common case where the sources arrive before the modelers, an accepted match creates a skeleton entity, marked as derived, that carries data governance immediately and can be deepened into a full model later. In each case the evidence flows up, and the meaning is decided rather than detected.

So the question underneath the feature announcements is not which platform infers the best semantic layer. It is who gets to define what your business means: the business itself, or an algorithm reading its exhaust. Your systems have plenty to say, so the move I keep landing on is to treat them as evidence, cross-examine them, and map what survives into a model someone accountable has signed off. Here is where to start.

  1. Start with the ten concepts the business actually runs on.
  2. Point discovery at the ugliest source and let it argue with you, one scored suggestion at a time.

What I’m less sure about is where the accept-or-dismiss gate holds and where it turns into rubber-stamping. If you’ve run human review over hundreds of source tables, what kept the decisions real?