Overview
Maarten co-founded Soda in April 2018 alongside Tom Baeyens, after spending seven years at Collibra in data governance roles, including two years in New York. He holds a Master in General Management from Vlerick Business School and a Master in Business Information Management from KU Leuven.
Soda's platform helps organizations detect data anomalies automatically, prevent future issues through collaborative data contracts, and manage data quality across pipelines. The company has raised $27.9M across three rounds, including a $14M round in 2024 for North American expansion. Investors include Hummingbird Ventures, Singular, and Point Nine.
Career Highlights
- Seven yearsCollibra in data governance before founding Soda in 2018
- Raised $27.9M for Soda including a 2024 round for US expansion
- Soda.io is a leading open-core data quality platform used by enterprise data teams globally
Career history
Insights & ideas
The through-line
Across nearly every post, Maarten Masschelein returns to one claim: data governance has been trapped in documents and manual processes, and AI's real contribution is turning governance into something executable rather than something written down. Early posts frame this narrowly, as a way to help data stewards implement rules without waiting on engineering [2][13]. Later posts widen the frame considerably, arguing that the entire discipline is moving through stages from manual to assistive to agentic AI [12], and that the endpoint is "AI-ready data" that both people and agents can trust without a human interpreting it first [10][11]. The most recent posts push the argument further still, arguing that catalogs and documentation are no longer enough because agents need machine-readable answers "at the moment of action" [6], and that metadata itself needs to become "active" rather than passive reference material [4]. The constant underneath all of this is a strict division of labor: AI handles implementation, humans keep the authority to define what good data means [2][13][7].
On the steward's role
Masschelein is consistent that automation should not touch judgment. "The steward role does not shrink in this arrangement. It sheds the waiting" [2]. He repeats the same list of tasks AI can take over, turning plain-English rules into running checks, drafting contracts, suggesting thresholds, updating checks, surfacing anomalies, while stressing "Notice what is not on that list: deciding what good data means. That stays with you" [2][13]. He extends this into a "double shift" framing for agents as data consumers: agents "act on data at machine speed" without the context a person has, so the division of labor becomes "AI agents propose... Data stewards approve" with fixes landing in staging, "never straight to production" [7].
On data contracts
Contracts are his preferred unit of executable governance. He describes them as what "define what a dataset must look like before it can move through a pipeline," generated in bulk with Contract Autopilot instead of written by hand one at a time [3]. He frames their purpose organizationally rather than just technically: "Data contracts bring a data steward and data engineer on the same page," because stewards define meaning and ownership while engineers build pipelines, and both need a shared artifact [8]. He poses the adoption question directly to his audience: "Would you trust a generated contract or would you still rather write it yourself check by check?" [9], and insists "Data contracts shouldn't start empty" [9][3].
On AI-readiness
Masschelein defines AI-ready data precisely: "data AI can consume and act on without a person interpreting it first" [10]. He uses the same example repeatedly, a $0 order that a person can contextualize but an agent cannot, to argue that "reversing automated actions executed on bad data is vastly more expensive than fixing a broken dashboard" [10][7]. He operationalizes readiness into direct diagnostic questions: coverage, ownership, contracts, and whether the last scan passed [6], and into a five-question checklist covering trustworthiness, machine-checkable definitions of "good," shared expectations between producers and consumers, root-cause traceability, and check costs [11]. His prescription is explicit: "Make quality machine-readable... Scope it to one real use case, not the whole estate... Write your specs where software can check them, not in a doc nobody opens... Re-check on a schedule, because readiness decays" [11][15]. He is careful to puncture the idea that this is a simple fix: "it's not a 'clean my data' button" [15].
On MCP and agentic context
He treats MCP as the mechanism that finally lets agents act on governance information rather than just read about it: "An agent needs them as a machine-readable response, at the moment of action. That is what MCP (Model Context Protocol) changes" [6]. He frames the shift as a redefinition of what governance produces: "governance used to produce documentation. In the agentic era, g[overnance produces something else]" [6]. Context fragmentation is his stated problem, stewards "switch between documentation, quality reports, catalogs, monitoring dashboards, and APIs just to answer a single question" [5], and MCP's promise is that "an agent working in Claude or Cursor can check the state of a dataset before it builds on top of it," using "the same data contract that tells your analyst the data is safe" [6].
From the stage
In the podcast and talk appearances, Masschelein discusses material that doesn't surface in the written posts: his 12-year career in data management on the software side, including managing operational systems, data warehouses, and reporting himself, and a stated philosophy of "changing perspectives in data work by understanding different stakeholders' roles and putting himself in their shoes" [16]. He also identifies Christian and Benno Verhague as influential co-founders of Caliberm and mentions a specific interest in applying data to fact-checking fake news [18], neither of which appears in the LinkedIn material. Separately, in a talk on forward-deployed engineering, he argues that the forward-deployed engineer has become "a new bottleneck replacing traditional product and engineering bottlenecks" at AI companies, illustrating it with a case where an engineer modified a machine learning matcher for deduplication edge cases and delivered results the next day, directly influencing a customer's decision to proceed [17].
Takeaways
- Generate a baseline of data contracts automatically rather than writing every check by hand; this is what tools like Contract Autopilot are meant to replace [3][9].
- Keep AI out of the decision of what "good data" means; use it only for implementation (drafting checks, suggesting thresholds, updating rules) and require human review before anything runs [2][13].
- Test your governance program against Masschelein's readiness questions: can you state what "good" looks like in machine-checkable form, and can you trace a broken dataset to its root cause [11][10].
- Treat agents as a distinct class of data consumer that acts without pausing on context a person would catch, and route AI-proposed fixes through staging and human approval rather than straight to production [7][10].
- Use MCP-style interfaces so agents can query coverage, ownership, contract status, and last-scan results directly, instead of relying on documentation written for humans [6][5].
- If you're building AI products, expect forward-deployed engineering work to become a bottleneck that shapes customer decisions faster than traditional product cycles [17].
Media & appearances
- SaaSiestYouTubeForward-Deployed Engineering: Why Your Product Team Must Change Now | Maarten Masschelein, CEO, SodaMaarten Masschelein discusses how the forward-deployed engineer role has become critical at AI companies, describing it as a new bottleneck replacing traditional product and engineering bottlenecks. He illustrates this with an example from a customer engagement where a forward-deployed engineer quickly modified a machine learning matcher to handle edge cases in data deduplication, delivering results the next day and directly influencing the customer's decision to proceed.
- SodaYouTubeS1 Ep01: Introducing In Conversation With, a Podcast with Host Maarten MasscheleinIn this episode introduction, Maarten Masschelein discusses his background in data, noting his early use of Excel and his passion for the impact data can make on decision-making. He identifies Christian and Benno Verhague as influential co-founders of Caliberm, expresses interest in data's application to fact-checking fake news, and indicates his goal for the podcast is to share new insights and ideas that listeners can apply to their daily work with data practitioners and technologists.
- SodaYouTubeS1 Ep02: Meet Maarten Masschelein, The Host of In Conversation WithMaarten Masschelein, CEO and founder of Soda Data, discusses his 12-year career in data management working on the software side, including his own experience managing operational systems, data warehouses, and reporting responsibilities. He explains his passion for technology since childhood and his philosophy on changing perspectives in data work by understanding different stakeholders' roles and putting himself in their shoes.
In the news
- What's the difference between a data quality dimension and a data quality metric? A metric is the number you monitor to see whether a dimension like completeness actually holds. For instance, if you want customer records to be complete, the metric might tell you that 91.5% of last names are filled in. When you're starting from zero, these are the data quality metrics I'd track first, on the tables you've classified as critical: ➨ Completeness rate: values that aren't null or empty, divided by total values. ➨ Validity rate:
- Have you ever used a RACI matrix to settle data owner vs data steward? Most teams have both roles on paper. Few can say where one job ends and the other starts. Here's how I split it for a single data domain: ➨ Data owner: Accountable. Sets policy and the quality bar, approves access, accepts the risk. ➨ Data steward: Responsible. Monitors quality, maintains definitions and lineage, triages incidents. ➨ Governance council or CDO: Consulted when a decision reaches beyond the domain or carries real risk. ➨ Data consumers: Informed
- Is a data quality check an internal control? If it protects the numbers in a financial or risk report, it should be. Most data teams just never document it as one. Most companies split risk work into three lines: - the teams that run the process - the risk and compliance team that oversees them - and internal audit that checks both A data quality check sits in the first line, with whoever runs the pipeline. Here are five regimes and what each one asks of your data. None of them explicitly tells you to buy a data quality tool. But
- Why do we need business glossaries when data dictionaries already exist? Aren't they the same thing? A data dictionary tells you what a column is. A business glossary tells you what the business means by it. The dictionary documents technical metadata: tables, columns, data types, relationships, and constraints. Engineers need that. But it cannot settle a dispute when "churn" means one thing to marketing and another to finance. The glossary standardizes business terms across departments. When a term genuinely means two things, it
- Data quality issues are found far from where they start. The error shows up in a dashboard or a report. The cause sits several steps earlier in the pipeline. The team corrects the number people can see, but the root cause is not always obvious. So here are 5 root causes every data team should care about: 𝗖𝗼𝗺𝗺𝘂𝗻𝗶𝗰𝗮𝘁𝗶𝗼𝗻 𝗴𝗮𝗽𝘀. Teams hold different assumptions about definitions, timing, or intended use, and nobody writes them down. 𝗠𝗶𝘀𝘀𝗶𝗻𝗴 𝗼𝘄𝗻𝗲𝗿𝘀𝗵𝗶𝗽. No one is accountable for the dataset, so the
- Your data governance program was written for people. Your AI agents can't read it. A governance program produces policies, glossaries, ownership models. All of it written for a human to read, interpret, and apply judgment to. An AI agent does none of those three. It doesn't open the policy page. It doesn't know that "active customer" means one thing in the glossary and another in the churn report. And it doesn't pause to ask whether the number it just read makes sense. That is a big shift. Data quality used to exist for humans
- Stop expecting your data analysts to clean data by hand. Ask an analyst what fills their week. The answers repeat across companies. - Fixing date formats. - Filling missing values. - Correcting the same address error that came back for the fourth time this month. Data cleaning is necessary work. It is also manual, repetitive, and it grows as data volume grows. The cost is specific. Companies hire analysts to find patterns, test assumptions, and support decisions with evidence. So every hour spent correcting records is an hour
- CLI, API, or MCP? Three ways to connect data quality to your stack, and most teams pick by habit rather than by the job. The question that settles it: what triggers the work, and how often? ➨ 𝗖𝗟𝗜 — the pipeline triggers it. Scheduled runs, CI checks, anything that repeats without a person in the room. Version-controlled and reviewable, in the same workflow as the rest of your code. ➨ 𝗔𝗣𝗜 — your own software triggers it. You are building a product that needs quality checks inside it, and you want the results back as data you
Related profiles
This page shows public professional information only, each fact cited. Is this you? send a correction, or ask for removal within 24 hours, no questions asked.







