
Maarten co-founded Soda in April 2018 alongside Tom Baeyens, after spending seven years at Collibra in data governance roles, including two years in New York. He holds a Master in General Management from Vlerick Business School and a Master in Business Information Management from KU Leuven.
Soda's platform helps organizations detect data anomalies automatically, prevent future issues through collaborative data contracts, and manage data quality across pipelines. The company has raised $27.9M across three rounds, including a $14M round in 2024 for North American expansion. Investors include Hummingbird Ventures, Singular, and Point Nine.Across nearly every post, Maarten Masschelein returns to one claim: data governance has been trapped in documents and manual processes, and AI's real contribution is turning governance into something executable rather than something written down. Early posts frame this narrowly, as a way to help data stewards implement rules without waiting on engineering . Later posts widen the frame considerably, arguing that the entire discipline is moving through stages from manual to assistive to agentic AI , and that the endpoint is "AI-ready data" that both people and agents can trust without a human interpreting it first . The most recent posts push the argument further still, arguing that catalogs and documentation are no longer enough because agents need machine-readable answers "at the moment of action" , and that metadata itself needs to become "active" rather than passive reference material . The constant underneath all of this is a strict division of labor: AI handles implementation, humans keep the authority to define what good data means .
Masschelein is consistent that automation should not touch judgment. "The steward role does not shrink in this arrangement. It sheds the waiting" . He repeats the same list of tasks AI can take over — turning plain-English rules into running checks, drafting contracts, suggesting thresholds, updating checks, surfacing anomalies — while stressing "Notice what is not on that list: deciding what good data means. That stays with you" . He extends this into a "double shift" framing for agents as data consumers: agents "act on data at machine speed" without the context a person has, so the division of labor becomes "AI agents propose... Data stewards approve" with fixes landing in staging, "never straight to production" .
Contracts are his preferred unit of executable governance. He describes them as what "define what a dataset must look like before it can move through a pipeline," generated in bulk with Contract Autopilot instead of written by hand one at a time . He frames their purpose organizationally rather than just technically: "Data contracts bring a data steward and data engineer on the same page," because stewards define meaning and ownership while engineers build pipelines, and both need a shared artifact . He poses the adoption question directly to his audience: "Would you trust a generated contract or would you still rather write it yourself check by check?" , and insists "Data contracts shouldn't start empty" .
Masschelein defines AI-ready data precisely: "data AI can consume and act on without a person interpreting it first" . He uses the same example repeatedly — a $0 order that a person can contextualize but an agent cannot — to argue that "reversing automated actions executed on bad data is vastly more expensive than fixing a broken dashboard" . He operationalizes readiness into direct diagnostic questions: coverage, ownership, contracts, and whether the last scan passed , and into a five-question checklist covering trustworthiness, machine-checkable definitions of "good," shared expectations between producers and consumers, root-cause traceability, and check costs . His prescription is explicit: "Make quality machine-readable... Scope it to one real use case, not the whole estate... Write your specs where software can check them, not in a doc nobody opens... Re-check on a schedule, because readiness decays" . He is careful to puncture the idea that this is a simple fix: "it's not a 'clean my data' button" .
He treats MCP as the mechanism that finally lets agents act on governance information rather than just read about it: "An agent needs them as a machine-readable response, at the moment of action. That is what MCP (Model Context Protocol) changes" . He frames the shift as a redefinition of what governance produces: "governance used to produce documentation. In the agentic era, g[overnance produces something else]" . Context fragmentation is his stated problem — stewards "switch between documentation, quality reports, catalogs, monitoring dashboards, and APIs just to answer a single question" — and MCP's promise is that "an agent working in Claude or Cursor can check the state of a dataset before it builds on top of it," using "the same data contract that tells your analyst the data is safe" .
In the podcast and talk appearances, Masschelein discusses material that doesn't surface in the written posts: his 12-year career in data management on the software side, including managing operational systems, data warehouses, and reporting himself, and a stated philosophy of "changing perspectives in data work by understanding different stakeholders' roles and putting himself in their shoes" 16. He also identifies Christian and Benno Verhague as influential co-founders of Caliberm and mentions a specific interest in applying data to fact-checking fake news 18, neither of which appears in the LinkedIn material. Separately, in a talk on forward-deployed engineering, he argues that the forward-deployed engineer has become "a new bottleneck replacing traditional product and engineering bottlenecks" at AI companies, illustrating it with a case where an engineer modified a machine learning matcher for deduplication edge cases and delivered results the next day, directly influencing a customer's decision to proceed 17.
From public career histories · 5 entries
Maarten Masschelein, CEO and founder of Soda Data, discusses his 12-year career in data management working on the software side, including his own experience managing operational systems, data warehouses, and reporting responsibilities. He explains his passion for technology since childhood and his philosophy on changing perspectives in data work by understanding different stakeholders' roles and putting himself in their shoes.
Maarten Masschelein discusses how the forward-deployed engineer role has become critical at AI companies, describing it as a new bottleneck replacing traditional product and engineering bottlenecks. He illustrates this with an example from a customer engagement where a forward-deployed engineer quickly modified a machine learning matcher to handle edge cases in data deduplication, delivering results the next day and directly influencing the customer's decision to proceed.
In this episode introduction, Maarten Masschelein discusses his background in data, noting his early use of Excel and his passion for the impact data can make on decision-making. He identifies Christian and Benno Verhague as influential co-founders of Caliberm, expresses interest in data's application to fact-checking fake news, and indicates his goal for the podcast is to share new insights and ideas that listeners can apply to their daily work with data practitioners and technologists.