Founding System Engineer

About Bauplan

Bauplan is building data infrastructure for a future in which the primary users of data systems are AI agents. Agents need infrastructure that lets them experiment, inspect results, recover from mistakes, and operate autonomously on production-scale data without compromising correctness.

We are building that infrastructure around a few core principles:

  • Everything as code: agents should interact with data infrastructure through simple, composable APIs rather than dashboards and manual workflows.
  • Git for data: every operation should happen in an isolated branch that can be inspected, compared, merged, or discarded.
  • Correct by design: schemas, contracts, lineage, and other guarantees should be explicit and machine-checkable.

Bauplan combines a novel FaaS runtime for Python and SQL execution, and Git-like data operations over object storage. Customers today run tens of thousands of jobs a day on the platform, with use cases ranging from data ingestion to data pipelines and analytical queries.

Making this experience feel simple requires innovation across the entire data lifecycle: APIs, distributed execution, scheduling, query processing, storage semantics, observability, and developer tooling. While the platform itself is closed source, we've published our work at top-tier venues such as FAST, VLDB, Middleware, SIGMOD, and we routinely collaborate with institutions such as Stanford University, TogetherAI, Columbia University.

The founding team previously built a company together that was acquired in 2019 by a public AI company. Bauplan is supported by leading Silicon Valley investors (Index Ventures, Innovation Endeavours, SPC), founders, and researchers, including Chris Ré (Stanford), Ihab Ilyas (Waterloo), Spencer Kimball (Cockroach), Erik Bernhardsson (Modal), Aditya Parameswaran (Berkeley). Our recent round of funding will allow us to double our engineering team and establish the first go-to-market organization, after an initial phase of funder-led sales.

The role

We are looking for an experienced engineer to build the next generation of agent-first data infrastructure. You will join at a stage where engineers have a large influence and ownership on both the product and the engineering culture.

You will work closely with experienced systems builders, database researchers, product designers, and the founders to design and implement performant systems for running data workloads over object storage.

You will work across product abstractions, distributed systems, data management, and developer tooling. You will be expected to understand problems vertically: from what an agent is trying to accomplish, down to the runtime, storage, and correctness properties required to make it possible. The ability to learn quickly, make intelligent trade-offs, challenge conventional wisdom, and operate across blurry boundaries is essential.

We are looking for a senior hire that could (and wish to) quickly graduate to tech lead, mentoring younger engineers on how to build, maintain and scale production data systems which are core to customers' operations.

What you will do

Design

  • Design systems that allow users to safely run, inspect, verify, and modify production-scale data workloads.
  • Develop APIs and abstractions that are easy for humans to understand, and practical for agents to implement.
  • Use production traces and agent behavior to identify opportunities for new system capabilities and optimizations.

Develop

  • Implement performant systems for executing Python and SQL workloads over object storage.
  • Improve the reliability, scalability, and efficiency of Bauplan's runtime, working on scheduling, concurrency and caching.
  • Improve the quality, security, observability, and iteration speed of the overall codebase.

Collaborate

  • Help customers diagnose issues, debug workloads, and understand system behavior.
  • Collaborate closely with product and design while retaining strong ownership of technical decisions.
  • Share our work with engineering and research communities through talks, papers, blog posts, and open-source releases.

Onboarding

First 30 days

  • Get comfortable with the architecture, codebase, tooling, and deployment process.
  • Understand the main production workflows end-to-end.
  • Ship a small feature or meaningful improvement.

By 60 days

  • Take ownership of one area of the platform and contribute actively to technical decisions.
  • Debug real production issues and improve reliability, performance, or developer experience.
  • Start working with junior members of the team, establishing trust and shared engineering standards.

By 90 days

  • Make meaningful improvements to existing systems with increasing independence.
  • Lead the design of a significant new capability.
  • Become a trusted leader in both operational and technical discussions.

What we are looking for

You are an experienced engineer who enjoys building simple abstractions over complex distributed systems. You should be excellent at:

  • Rust and Python, or able to demonstrate outstanding equivalent experience in relevant systems.
  • Designing, implementing, and operating distributed systems.
  • Concurrency, asynchronous programming, fault tolerance, and failure recovery.
  • Building cloud-based infrastructure.
  • Modern software-engineering practices, including Git, containers, infrastructure as code, CI/CD, testing, security, and observability.
  • Reading and writing technical design documents.
  • Leading technical discussions and communicating trade-offs at different levels of abstraction, mentoring younger engineers through technical leadership.
  • Taking ownership of ambiguous problems from initial requirements through production operation.

You should have working knowledge of:

  • SQL and relational data systems.
  • Object storage and cloud-native storage architectures.
  • Distributed computation, whether through Kubernetes, Spark, query engines, or custom infrastructure.
  • API and developer-tooling design.

Finally, while not necessary, experience in one or more of the following areas would be especially valuable:

  • Database internals, query planning, execution engines, or optimization.
  • Apache Arrow, DataFusion, DuckDB, or related systems.
  • Transaction processing, Apache Iceberg, or data catalogs.
  • Serverless runtimes, function scheduling, or sandboxed execution.
  • Type systems, formal methods, or distributed-systems simulation.
  • Infrastructure for coding agents, agent evaluation, or autonomous software development.

We do not expect one person to have experience across all these areas. We care more about strong fundamentals, intellectual curiosity, technical judgment, and the ability to learn unfamiliar systems quickly.

Want to know more about our current stack and some of the engineering challenges we face? Check out our blog (for example, forking DuckDB and later migrating to DataFusion), our Git-for-data talk, or our latest research.

Applying

Please send an email to jacopo.tagliabue@bauplanlabs.com with your CV or LinkedIn profile, GitHub and Google Scholar (if applicable) and we will reach out to schedule a first call if there is a match.

August 23, 2026
November 23, 2026
FULL_TIME
New York
NY
US