Enterprises are pointing AI agents at their data faster than they can govern them. The last generation of data systems were built for humans: someone writes a query, reads the result, and moves on. Agents work differently: they act in parallel, they get things wrong on the way to getting them right, and they generate orders of magnitude more code and operations than any human team ever did and ever will.
Today more than ever, technical leaders are under pressure to usher their engineering teams into the agentic era and to find secure and scalable solutions for agents to work with enterprise data.
Finding these solutions requires rethinking not only processes, but systems. In the words of Meta's VP of Infrastructure, Barak Yagour, once an organization embraces the concept that agents will become the primary users of its systems, the company must redesign infrastructure for AI agents.
01Why data is the hard part
Most data organizations will have to deal with some pretty serious questions about whether their data systems are a good fit for the transition to AI-first development. If your organization is undergoing AI transformation right now, you will have to find an answer to three fundamental problems regarding data: safety, trust and costs.
Safety
Agents need access to real data at scale. Imagine you have a pipeline feeding your revenue model and the pipeline is now outdated. Your engineers may very well use AI agents to rebuild it in a different system or with different assumptions. To port the business logic into the new system, an agent needs to use the output of the legacy process to determine what is correct. So it needs to run against the same inputs, compare its output against the existing one, make changes and iterate until it converges. This process cannot take place on synthetic data or in a dev sandbox.
There is no way around it: the agent needs to iterate on real production data.
At the same time, allowing agents to manipulate production data requires being extremely careful. The first problem you will face is the fact that data lives in shared locations, like cloud databases, warehouses and data lakes, so any mistake will immediately impact every other user or consumer reading from these systems.
This is one of the major differences between code and data. Code is shared in version control platforms like GitHub, but developers can work with it locally in a safe way. Instead, every iteration over your data is a production event, because every time an agent materializes its output on a shared table, dashboards, feature stores and downstream jobs read from it on their next run.
In addition, because agents iterate at a speed that is incomparable with human pace, they will build on top of each other's failures, making it extremely hard to chase down the root cause of data corruption.
Trust
Which brings us to our second point: trust. Coding agents make it possible to produce an order of magnitude more code at an incredible speed. This is great (although not without some serious side effects), but at the same time it makes it harder and harder to have clear visibility into your own system.
One of the major problems that we see in our customers is that the code bases of their projects grow quickly beyond the ability of any developer to fully understand them (a problem that we often solve in a very hands-on way with FDEs).
Every data asset, dashboard or application in your system now rests on a dependency graph that grows every time an agent contributes upstream, and a lot of that code may have reached production without a human reading it.
The main problem is that it becomes impossible to audit data processes. If six weeks after your revenue pipeline has been updated, the finance team asks where the logic comes from, the answer may be spread across orchestrator logs, warehouse query history and hundreds of thousands of lines of code. This becomes a particularly serious problem in regulated industries, where auditability and full reproducibility are a hard requirement.
Cost
Agents can be very expensive, but in the context of data processing they can be twice as expensive because they run two meters at once: you pay for the tokens they spend reasoning AND you pay again for the cloud compute they trigger.
Both meters run on work that gets thrown away, because agents do a lot of trial and error in their development process.
Keeping total spending under control becomes one of the most urgent problems, along with the ability to forecast it, since both agents and data processing are based on on-demand pricing models and demand is growing massively.