Data Products: Context Engineering for AI

Source: Starburst•

Data Products: Context Engineering for AI

AI agents are experiencing a context crisis. AI context is the talk of the industry …

AI agents are experiencing a context crisis.

AI context is the talk of the industry right now. It’s not surprising to hear so many companies struggle with this. One survey by Deloitte found that 48% of respondents named searchability of data as a challenge to their AI automation strategy. Another 47% said reusability was the blocker.

Most companies know that the solution involves context engineering to create the context layer needed for AI.

The question is how?

The answer is data products. Data products aren’t a new approach to organizing data. But they’re getting renewed attention because they provide both a unit of packaging and a methodology for creating context.

In this article, we’ll look at why data products are the unit of context for AI. Even more importantly, we’ll see why they’re the ideal unit of work for context engineering.

The need for context engineering

AI agents don’t fail for lack of model power. Rather, they fail for lack of context.

A general-purpose model without appropriate grounding in the context layer of your specific organization will confidently fill gaps, often with hallucinated information. Overcoming this isn’t easy, because not all context currently exists as high-quality data. As I wrote previously, a lot of it isn’t written down and exists only as institutional knowledge.

This information, along with its accompanying metadata, needs to be accessed by AI. This is where data products come in. They serve both as the unit of context for AI agents and as the methodology for putting context engineering into practice.

In a core sense, data products are context engineering. Before you can feed context to AI, you need to conduct the important upstream work of curating and packaging your context as data products.

Data products make context accessible to AI

It’s one thing to identify the context you need. The next challenge is turning it into data and ensuring it’s discoverable and usable by all authorized users across the company.

Data products are uniquely suited to fill this role. A data product is a curated, accessible wrapper built around high-quality datasets and bundled with both metadata and business logic. For more information on data products in general, check out my video below.

Importantly, data products aren’t just raw data. They’re a wrapper around raw data, created with a specific intent for a specific domain team. A data product contains rich metadata, including versioning, documentation, and access control rules.

This richness gives data products multiple attributes that raw data often lacks, including being discoverable, addressable, self-describable, interoperable, trustworthy, and secure.

Data products put context engineering into practice

Data products have always provided context for analytics. Now, that role is taking on enhanced importance in the AI age. That’s because creating data products is the same as putting context engineering into practice.

Bundle data and context in one layer

The data included in a data product may exist in multiple locations. That makes it harder for AI agents to discover and use it.

Data products leverage data federation to combine data from multiple data sources located across your organization. Instead of taking a dependency on a long and expensive data centralization process, data products utilize data where it lives now, turning raw data into a single, high-quality unit of context.

Since data products are interoperable, new data products can be created as combinations of other data products, fitting together as needed. Your data products represent a federated, distributed data layer that provides not just data, but the meaning behind that data.

Enable multiple consumption models

Besides metadata, data products come bundled with the business logic that underlies their business use case. They expose this logic by providing the interfaces necessary to access their data, including the underlying semantic layer, data contracts, and programmatic interfaces such as APIs and event streams.

In this sense, data products are upstream of the rest of your agentic AI efforts. They act as the necessary building blocks required for agents to yield timely, accurate answers to customer queries.

Once that context engineering work is done, however, you can reap the benefits of economies of scale in your AI efforts. That’s because data products turn data into reusable units of context that can be consumed by many different types of data consumers.

Most data transformation efforts aim to make data available for one consumption format, such as an analytics dashboard or a data application. A data product can be consumed equally by an AI agent, by an application development team creating a new data-driven app for the customer service department, or even by end users through conversational language interfaces such as Starburst’s (AIDA).

Govern access to data

One of the challenges with AI is controlling and monitoring access to data. Part of this is ensuring the data your agents are accessing is high-quality, trusted, and authoritative.

An equally pressing concern is data security and compliance. Without proper authorization controls and governance constraints, an AI agent may attempt to access sensitive data it should never have touched. This issue has taken on increased importance given the recent headlines about rogue AI agents.

Data products act not just as a unit of access, but as a unit of data governance. This is because data products act as a contract, with field definitions, ownership, and access policies bundled together with the data.

By using a platform such as Starburst to create and manage data products, you can define granular data access rules to enforce security and compliance requirements. Role-based access control (RBAC) can grant users permissions to entire data products and catalogs. Attribute-based access control (ABAC) can enforce granular permissions, such as masking personally identifiable information (PII) from users and AI agents without the appropriate level of clearance.

Data products simplify access auditing by providing a unified access point to data that enables centralized logging and data lineage. This prevents security policies from drifting out of alignment when a data schema changes.

Companies operating in regulated industries, like government and healthcare, often need to prove that they’re properly handling sensitive customer data in accordance with strict requirements. In this context, data products provide a consistent approach to access control and logging across your entire distributed data estate, whether your data lives in fully managed cloud environments, self-managed deployments, or a mix of both.

Scaling AI with data product-driven context

If you want your AI to actually work in your organization, you need to engage in context engineering. The easiest way to do that is using data products. The only way to solve the context crisis with AI is by curating your context into high-quality data products that are accessible by multiple consumers.

Data products have been around for a while. But many people still don’t understand exactly what they are or, more importantly, how to build them as vehicles for context.

That’s why we created a free eBook that dives deep into what data products are and how to manage them. Download it today and learn how to move from a completely centralized, IT-driven approach to data to a federated and interoperable model that scales to the AI age.

What this article says