CTO guide

Redesigning your data systems for agents

Ciro Greco
Co-founder and CEO

Enterprises are pointing AI agents at their data faster than they can govern them. The last generation of data systems were built for humans: someone writes a query, reads the result, and moves on. Agents work differently: they act in parallel, they get things wrong on the way to getting them right, and they generate orders of magnitude more code and operations than any human team ever did and ever will.

Today more than ever, technical leaders are under pressure to usher their engineering teams into the agentic era and to find secure and scalable solutions for agents to work with enterprise data.

Finding these solutions requires rethinking not only processes, but systems. In the words of Meta's VP of Infrastructure, Barak Yagour, once an organization embraces the concept that agents will become the primary users of its systems, the company must redesign infrastructure for AI agents.

01Why data is the hard part

Most data organizations will have to deal with some pretty serious questions about whether their data systems are a good fit for the transition to AI-first development. If your organization is undergoing AI transformation right now, you will have to find an answer to three fundamental problems regarding data: safety, trust and costs.

01SafetyAgents need real production data, and every write to a shared table is a production event.
02TrustCode grows beyond what any developer can read, and data processes become impossible to audit.
03CostAgents run two meters at once: tokens for reasoning and compute for the work they trigger.
Figure 1. The three data problems every AI transformation has to answer.

Safety

Agents need access to real data at scale. Imagine you have a pipeline feeding your revenue model and the pipeline is now outdated. Your engineers may very well use AI agents to rebuild it in a different system or with different assumptions. To port the business logic into the new system, an agent needs to use the output of the legacy process to determine what is correct. So it needs to run against the same inputs, compare its output against the existing one, make changes and iterate until it converges. This process cannot take place on synthetic data or in a dev sandbox.

There is no way around it: the agent needs to iterate on real production data.

At the same time, allowing agents to manipulate production data requires being extremely careful. The first problem you will face is the fact that data lives in shared locations, like cloud databases, warehouses and data lakes, so any mistake will immediately impact every other user or consumer reading from these systems.

This is one of the major differences between code and data. Code is shared in version control platforms like GitHub, but developers can work with it locally in a safe way. Instead, every iteration over your data is a production event, because every time an agent materializes its output on a shared table, dashboards, feature stores and downstream jobs read from it on their next run.

AI agentIterates directly on shared tablesattempt 1attempt 2attempt 3attempt 4Revenue reportreads attempt 1Recommendationsreads attempt 2Customer billingreads attempt 3mainshared tables
main: the shared tables every application readsa wrong write, visible to everyone
Figure 2. Every iteration lands on the shared tables, and every application built on them reads it on its next run.

In addition, because agents iterate at a speed that is incomparable with human pace, they will build on top of each other's failures, making it extremely hard to chase down the root cause of data corruption.

Trust

Which brings us to our second point: trust. Coding agents make it possible to produce an order of magnitude more code at an incredible speed. This is great (although not without some serious side effects), but at the same time it makes it harder and harder to have clear visibility into your own system.

One of the major problems that we see in our customers is that the code bases of their projects grow quickly beyond the ability of any developer to fully understand them (a problem that we often solve in a very hands-on way with FDEs).

Every data asset, dashboard or application in your system now rests on a dependency graph that grows every time an agent contributes upstream, and a lot of that code may have reached production without a human reading it.

Revenue modelwhere does thislogic come from?
reviewed by a humanwritten by an agent, never read by a human
Figure 3. Every agent contribution upstream grows the dependency graph, and much of it reaches production unread.

The main problem is that it becomes impossible to audit data processes. If six weeks after your revenue pipeline has been updated, the finance team asks where the logic comes from, the answer may be spread across orchestrator logs, warehouse query history and hundreds of thousands of lines of code. This becomes a particularly serious problem in regulated industries, where auditability and full reproducibility are a hard requirement.

Cost

Agents can be very expensive, but in the context of data processing they can be twice as expensive because they run two meters at once: you pay for the tokens they spend reasoning AND you pay again for the cloud compute they trigger.

Both meters run on work that gets thrown away, because agents do a lot of trial and error in their development process.

Keeping total spending under control becomes one of the most urgent problems, along with the ability to forecast it, since both agents and data processing are based on on-demand pricing models and demand is growing massively.

attempt 1attempt 2attempt 3attempt 4attempt 5total Meter 1tokens · reasoning
280k
480k
590k
670k
750k
2.77M2.02M wasted Meter 2compute · warehouse
$3.00
$3.00
$3.00
$3.00
$3.00
$15.00$12.00 wasted ✕ wrong✕ wrong✕ wrong✕ wrong✓ kept
Figure 4. Every failed attempt is paid for twice: once in tokens, once in compute. Anonymized workload from a Bauplan deployment.

02Redesigning your data systems for agents

Evaluating a data platform used to mean asking a familiar set of questions, and all of them assumed a person at the other end: someone who reads the result and notices when it looks wrong. Agents remove that assumption. What matters now is how the system behaves when it is used faster than anyone can watch, by something that gets things wrong several times before getting them right. The three areas below are where that difference shows up first.

Make every change safe and reversible

An agent that cannot write is useless, but writing data is dangerous.

What you need is a way for an agent to write at full speed while nothing it writes reaches anyone else until you say so, and a way to always be able to audit the process.

AI agentIterates on its own branchattempt 1attempt 2attempt 3attempt 4Revenue reportreads main, unaffectedRecommendationsreads main, unaffectedCustomer billingreads main, unaffectedmergemainshared tables
main: the shared tables every application readsisolated branch: invisible until it merges
Figure 5. The same four attempts, on an isolated branch. Applications keep reading main and only see the result once it merges.

Modern table formats already brought some primitives to help with this problem. For instance, Apache Iceberg gives you snapshots, time travel and atomic commits on a single table, and the format now carries branches and tags. Similarly, zero-copy clones have given warehouses a private version of an entire database for years.

These are very powerful building blocks in a situation where we want our system to be more secure and auditable by design. The question remains how those building blocks are put to use to design a system that supports end-to-end agentic usage.

One complication comes from the fact that the unit of work of agentic workflows is almost invariably larger than one table, so they require higher-level capabilities that span across tables and systems.

Agents require systems where you can branch and isolate many data assets at once, and where a pipeline that writes dozens of tables makes them all visible together at once, rather than one at a time as each job finishes. In the second case, there is a window in which half a change is readable and nothing marks it incomplete, allowing agents to build on each other's failures.

Then ask what an agent is allowed to do inside that isolated space. Creating, renaming or dropping a table is a catalog operation, and a branch that lives inside a table cannot hold a table that does not exist yet, so a pipeline that adds an output has already left the boundary.

We envision a world where the number of agents touching the data at the same time will grow exponentially as companies become better at adopting AI automation.

This world introduces new questions that our current systems were not designed to answer. When two agents change overlapping tables, is there a defined way to resolve conflict, or does the last writer win? Can you undo the actions of agents in one simple action, even when those actions span across multiple tables? Can you read many tables as of one consistent moment?

These questions are central to our own design, in what Bauplan calls Git-for-data. Agents must work in isolation, try ideas that might fail, and roll back without an incident review.

Put the whole system in the code

“The limits of my language mean the limits of my world”, as whatshisname used to say. The most effective way to get information into the agent context is to express it directly in the code in the repository the agent has access to.

Up until now, the main question about the usability of your data platform has been how intuitive the UI was for the different users working with it: Python notebooks for data scientists, SQL scripts for data engineers, point-and-click dashboards for analysts, etc.

Moving forward, this question changes into how much of your data platform actually lives in a repository.

In the past years, frameworks like dbt helped express the transformation logic as code you can work on from the IDE. This design choice was most definitely in the right direction. However, agents push that design choice even further and require us to be more intentional about designing a system where everything is code: the data assets, the transformation logic, the control workflow in which that transformation takes place and the infrastructure that runs the whole life cycle.

repo/
├──data/types · schemas · contracts · semantics
├──transformations/Python + SQL, one output contract
├──control_flow/bronze → silver → gold, recovery paths
└──infra/pandas, polars · memory, CPU, runtime, timeout
Figure 6. Data, logic, control flow and infrastructure in one repository: an agent can read it, a human can review it.

Types and schemas are enforceable today. Descriptions can be written in dbt models and YAML and pushed down into warehouse comments, and table formats like Iceberg will reject writes that violate the schema. All of that checks the shape of a table, but there is more you can do.

The first thing is contracts. Contracts and code are typically separate artifacts, but in an ideal system for agents contracts should be baked into the transformation code itself, so they cannot drift apart and the whole DAG can be type-checked locally before a single query is issued. The second is semantics. Transformation logic does not carry what a column represents, or which of two plausible definitions of revenue it encodes.

A column can keep its name and its type and change what it counts, and every mechanism above passes it.

When both these things are expressible in code, the payoff is that documentation never drifts from data and all users can always trust the data. The only way to change what a column means is to change the code that writes it, and that change arrives in a pull request next to everything else, so a shift in the definition of revenue is something a reviewer can catch before it ships. The compounding effect is on the agents themselves. Every agent, whether it reads or writes the table afterwards, starts with a clear definition. So the same question stops being answered from scratch, differently, every time.

Make scale and mistakes affordable

“A computer in every home” made Microsoft a fortune, and for decades the digital revolution was someone sitting in front of a PC typing and clicking. As compute grew more than 11 orders of magnitude from 1975 to 2022, it is easy to forget that today the bulk of it is server-side: the digital world is mostly “invisible” infrastructure running silently in the background of our daily life.

LLM tokens will undergo the same transition. Today, most token usage is triggered by people prompting agents. Tomorrow, most tokens will be used by agents reasoning asynchronously in the background with other agents: instead of a team of data engineers checking for failed jobs every morning, you will rely on agents running in the background, maintaining and repairing broken pipelines automatically while you sleep.

As you scale data agents to cover new use cases and reason 24/7 about your data, technical leaders have two drivers to scale efficiently:

Token efficiency.

Token spend is largely determined by the interface that your agents have to work with. Platforms designed for the previous era carry large surface areas. A single operation means calling several services, reading long configuration objects, keeping cluster state in mind, etc. Their APIs are verbose, so a lot of tokens are spent just going through them.

The cost compounds inside the context window. Platform mechanics fill the context, which leaves less room for the problem and degrades the quality of the reasoning as the session goes on. Teams compensate by reaching for a bigger model, but you end up paying frontier prices for a task a smaller model could have handled on a simpler platform.

The feedback loop matters just as much. Every failed attempt is paid for twice, once in the tokens that produced it and once in the tokens spent reading the failure and trying again. A system that returns a precise error in seconds converges in a few turns. A system that takes ten minutes to fail, and reports it in a stack trace written for humans, converges in dozens.

This changes how infrastructure should be evaluated. The question to ask of a data platform is how simple it is for an agent to operate. API design becomes a purchasing criterion: how many calls a common task takes, how much has to be held in context to make one, and how readable the error is when it fails. A platform that is pleasant for a human team can still be expensive for a team of agents.

Compute efficiency.

The same goes for data processing cost. Because agents issue far more queries than a data team ever did, the cost structure of the platform running those queries matters more than its headline price. Most of that volume is speculative: exploratory scans, retries, and checks that come back clean.

The cost structure of data infrastructure where organizations pay per query, per byte scanned or per cluster will turn punitive in an agent-first world. In the evaluation of a platform, what used to be a feature of data warehouses (“start small, pay only for what you run, grow with usage”) is now a bug.

03Conclusion

Agents are not a feature you can bolt onto the data stack you already have. They change who the main user of that stack is, and most data platforms were built for a person who reads the result, notices when it looks wrong, and fixes it by hand. Safety, trust and cost are where that assumption breaks first, and each one points to a design principle. Changes should be isolated and reversible across many tables at once. The whole system (data, logic, control flow and infrastructure) should live in code that an agent can read and a human can review. Failing should be cheap enough that agents can afford to do it constantly. None of this requires waiting. You can ask any platform the questions in this guide today. The answers will tell you whether your data infrastructure is ready for the agents your teams are already deploying, or whether it's about to become what slows them down.