In June, at TDWI in Munich, I watched an AI agent build a data product in about 18 minutes. That was pretty impressive, for me and for everyone in the room. During the talk it quickly became clear to me what has to be in place before something like that works.
The plan for this quarter had actually been set since January: data products, as the final part of my three series this year on AI in data modeling, temporal data, and data modeling fundamentals. Then came the talk "Open Standards for Data Products" by Simon Harrer, and his demo fit the series so well that it is now the hook.
Simon told me later on my podcast what the agent needed. First, a data contract describing which tables get created, how they are filled, and what the data in them means, linked to an ontology. Then skills holding the company's conventions, from naming rules to how personal data is handled. And a marketplace where it could look up what already exists. That is where it found the three data products it builds on, and it requested access on its own. For the demo, approval was automatic, otherwise the 18 minutes would not have been possible.
Interestingly, all three are descriptions, and none of them is AI. Descriptions like these are exactly what my three series this year were about. This article is the map for a three-part series on how they lead to better data products. You can hear the full conversation with Simon on November 24 in episode #005 of my podcast [antifragile data] (in German).
Part 1 · out October 28
What is a data product? A star schema with 15 tables: one, fifteen, or no data product at all — and what belongs to it besides the interface.
Part 2 · out November 18
Where does the data come from, where does a data product end, and how do you provide it so it survives the first breaking change?
Part 3 · out December 9
Automation through information models: how the model turns into the semantics that contracts and agents rely on.
What is a data product?
Ask clients and you get very different answers. So I brought Simon an example from my Dimensional Modeling Fundamentals Class: a star schema for accounting, with cost allocation across several stages, around 15 fact and dimension tables. Is that one data product, fifteen, or none at all?
His answer: one data product. The tables belong to one team, they change together, and they are only useful together. For him, the product includes the pipeline that fills the tables, not just the interface — he calls it the factory.
I like to explain this with a drive-thru. When you order there, you want the finished product, and you don't ask where the bun or the lettuce came from. All of that is still part of it. A physical table handed out without any further description, on the other hand, is just a table for now.
→ More in part 1 on October 28.
Where does a data product end?
The other extreme is a reference table with 15 rows and four columns. Simon counts that as a data product too, because someone owns it, a master data team, for instance. He would approve access automatically, since lots of people need it and there is little to protect.
Making every single table its own data product is something he considers wrong. His yardstick comes from software engineering: high cohesion, low coupling. What users buy together, own together, and version together belongs in one product, like orders and order line items. The first breaking change shows why. Version 2 of the orders next to version 1 of the line items gets complicated fast.
→ More in part 2 on November 18.
From model to value
Over the past months I have seen physical tables published as data products more than once, with no further description. If you know where the data comes from, you can get by for a while. An agent like the one in the demo finds nothing in there.
What we used to simply call the interface is now covered by three open standards: ODCS describes the data contract, ODPS the data product, and OSI (now Apache Ossie) the semantics. There is a reason for the split. An order number shows up in many contracts, sometimes as order number, sometimes as order ID, and what it means in consumer sales is something you only want to describe once. The contract links to that description, and always upward, because the semantics can't know every implementation.
That technology-independent description of the business is what I call an information model, and it is what the fundamentals series built over the summer. The semantics can be described from it. Whatever is defined there, every contract that links to it inherits, the classification from public to top secret, for example. An agent then finds the order ID even in a product where it is called order number.
That is the value I mean in the series title. Data products described from an information model are worth more, to people and agents alike. Why AI depends on that foundation is in The AI Paradox; what an information model is worth, in Data Modeling as Competitive Advantage.
→ More in part 3 on December 9.
Data products have a timeline
Simon has a nice image for this. You don't buy a canister of data at the supermarket and carry it home. You buy access, a subscription to data that keeps changing and flows through the product.
That makes time part of the product. A subscriber needs to know which view they get: what reality looks like today (As-Is), what it looked like back then (As-Was), or what we knew back then (As-Of). Those three views were the subject of the temporal data series in the spring.
How data is always there on time
The subscriber sees the interface. Making sure the data product is available as planned is the job of the engine of the data solution behind it, meaning the data logistics processes. In Data Engine Thinking, Roelant Vos and I summarized the principles these processes should follow, among them idempotent, deterministic, and independent. What they have in common is that every process can be run and re-run at any point in time without doing damage. The book has a simple test for it: run the process twice in a row and see whether that causes issues.
For availability, two things matter most. Processes run independently of each other, and every run processes all pending changes. Then you can change the loading frequency when a subscriber needs the data sooner, and nobody has to rebuild the engine.
How to recognize the value
That leaves the question of how to tell whether a data product is any good. I look at qualitative metrics first, and the most important one is reuse. Do others build on it, or does every team build its own version?
Then there are two sides to check: how well a product is described and how heavily it is used. Lots of documentation with little use points to a product built for the wrong audience. Little documentation with heavy use is a risk, because then the knowledge sits in the users' heads and nowhere else.
From storage to outcomes
Bill Reynolds put the whole path into five lines back in September:
Data without structure is storage
Data with shared meaning is an asset
Data with context is information
Information drives decisions
Decisions drive outcomes
For me, the second line is the core of this series. The best data products happen where an information model, a clean timeline, and an engine that delivers reliably come together. At that point it doesn't matter whether a person orders at the drive-thru or an agent does.
So long,
Dirk
About this series: "Data Products: From Model to Value" brings this year's three series together. Part 1 (October 28) clarifies what a data product is. Part 2 (November 18) looks at where the data comes from and how to cut and provide a data product. Part 3 (December 9) explores automation through information models. The conversation with Simon Harrer goes live on November 24 as episode #005 of my podcast [antifragile data] (in German).
Building the information model properly
Clarifying terms, negotiating definitions, modeling conceptually and logically: that is the foundation for well-described data products and the core of the Data Modeling Master Class. New dates are in preparation.
