Pieter De Leenheer

Founder at Collibra

Overview

Pieter De Leenheer is the founder of Collibra [1][4]. De Leenheer is an entrepreneurial technology executive with over 20 years of experience driving technological innovation and delivering enterprise software excellence in high-growth international environments, with a proven track record of scaling high-performance engineering teams and delivering data- and AI-driven enterprise SaaS products [2]. De Leenheer holds a Ph.D. in Computer Science from Vrije Universiteit Brussel [13] and served as a Software Engineer and Researcher at the AI Lab there from 2002 to 2008 [10]. De Leenheer was an AI Expert at the European Commission from 2008 to 2018 [9], an Adjunct Professor of AI Frameworks at Columbia University from 2017 to 2022 [12], and an Entrepreneur in Residence at Harvard Business School from 2020 to 2021 [11]. De Leenheer currently serves as Chief Technology Officer at DNAnexus [3] and as a Visiting Research Engineer at the San Diego Supercomputer Center [5].

Profile introduction
Source excerptLinkedIn [2]

Pieter is an entrepreneurial technology executive with over 20 years of experience driving technological innovation and delivering enterprise software excellence in high-growth international environments. He has a proven track record of scaling high-performance engineering teams with a strong focus on diversity, architecting robust cloud infrastructure, and delivering data- and AI-driven enterprise SaaS products that have fueled triple-digit business growth. Pieter also maintains strong academic credentials and frequently writes, teaches, and advises on computing and management aspects of da…

Career history

  1. Chief Technology OfficerApr 2025 to presentDNAnexus
  2. FounderJun 2008 to presentCollibra
  3. Visiting Research EngineerMay 2019 to presentSan Diego Supercomputer Center
  4. Board MemberNov 2022 to Dec 2024Spirion
  5. Chief Technology Officer2022 to 20241upHealth, Inc.
  6. Chief Technology Officer2020 to 2022ZenOptics
  7. AI ExpertJan 2008 to Jan 2018European Commission
  8. Software Engineer / Researcher at AI Lab2002 to 2008Vrije Universiteit Brussel

Education

  1. Entrepreneur in Residence2020 - 2021Harvard Business School
  2. Adjunct Professor, AI Frameworks2017 - 2022Columbia University
  3. Ph.D., Computer Science2003 - 2008Vrije Universiteit Brussel

Insights & ideas

The through-line

Across everything he says, one argument keeps resurfacing: American healthcare's problem is not a shortage of technology but a refusal to use the data it already generates. "The healthcare industry basically still kind of lives in the 1980s when it comes to technology" [1], and the consequence is a system costing "about 5 trillion dollar a year, which is nearly 20% of the GDP of the American economy, and if we don't do anything about it, it will just lead to the bankruptcy of the system overall" [1]. The fix he proposes is unglamorous and structural: unlock clinical and claims data, combine them, govern them properly, and let that combination push the industry from fee-for-service toward value-based care [1].

Tempering that is a realism about pace. He is not promising a fast turnaround, because "it takes also a long time to revert an industry that has been going the wrong direction for almost 50 years" [1]. The change he expects over the next decade is incremental and, at first, uneven in who benefits.

On the 1980s problem

He is explicit that the binding constraint is not engineering. The number one challenge is that healthcare CIOs still operate with proprietary systems and paper, and that they have yet to take the cloud migration step their counterparts in financial services already took [1]. The absurdity of the status quo is best measured commercially: "Companies like Iron Mountain make billions of dollars in filling airplanes with patient records, 9 billion a year" [1]. Set that against ordinary consumer expectations and the gap becomes indefensible. "I do know if I pull up my phone here what did I order at Amazon for the last 10 years, and Google even knows where I was 5 years ago" [1], while a patient's own medical history remains scattered and inaccessible.

On combining clinical and claims data

Today each dataset is used narrowly and separately: clinical data mainly by pharma for A/B testing in trials, claims data mainly for insurance risk underwriting [1]. His central technical claim is that the value sits in the join. "Once you start combining data these two sources together you get in one plus one is more than two effects" [1], and those effects open entirely new categories of use, care management and population health among them [1]. This is the mechanism by which the shift from fee-for-service to value-based care actually becomes possible rather than merely aspirational [1].

On prevention as a data problem

He extends the same logic to chronic disease. The four main chronic diseases, cancers, metabolic disorders, cardiovascular and neurodegenerative diseases, share a common set of risk factors [1]. If the correlations between risk, behavior and inflammation can be understood from data, prevention itself becomes data-driven rather than guesswork [1]. It is the longest-horizon payoff he describes, and it depends entirely on the combining work being done first.

On regulation as the infrastructure

Where others might treat regulation as friction, he treats it as the foundation. Good data infrastructure starts with regulatory tailwind, and the analogy he reaches for is AT&T's telephone cables running out to connect even a single remote house in Texas [1]. Universal reach was mandated, then built. He is correspondingly wary when regulation is drafted loosely: the clause in Europe's AI Act banning applications that "influence behavior" troubles him precisely because it is vaguely defined and could end up excluding valuable healthcare opportunities [1].

On data as money in a bank

His preferred frame for patient data ownership is fiduciary. "If you give your money to a bank you trust the bank with their fiduciary responsibility to invest it as such and give you full transparency into their investments, and I think we have to have the same mindset with data" [1]. That means entrusting health data to institutions under an explicit duty, with full transparency and governance controls attached [1]. It also inverts a common privacy intuition: cloud platforms make downstream data use more observable and auditable than records sitting on a local doctor's laptop [1]. Centralisation, done with governance, is what makes accountability possible.

On AI, and what it is not yet good for

He draws a sharp line around where AI belongs in healthcare today. The focus should be specific, human-in-the-loop use cases, care management recommendations being the example he gives, rather than generative AI, which still hallucinates and is far from ready for this domain [1].

On the next ten years

His forecast is deliberately unromantic about sequencing. Better data will initially make insurance companies richer, through quality-based government reimbursements, before a rebalancing eventually gives patients more control over their own data [1]. The payer benefits first; patient sovereignty arrives second.

Takeaways

  • Treat the join between clinical and claims data as the actual product: separately they serve trials and underwriting, together they produce "one plus one is more than two" effects and enable care management and population health [1].
  • Do not wait for permission structures to appear on their own. Good data infrastructure starts with regulatory tailwind, on the AT&T model of mandated universal connection [1].
  • Push back on vague regulatory language early; the AI Act's ban on applications that "influence behavior" is undefined enough to sweep in legitimate healthcare work [1].
  • Sell cloud migration to healthcare CIOs on auditability, not just cost: downstream data use is more observable in a governed platform than on a doctor's laptop [1].
  • Scope AI to narrow, human-in-the-loop decisions such as care management recommendations, and keep generative AI out of the critical path while it still hallucinates [1].
  • Expect payers to capture the first decade of gains via quality-based reimbursements, with patient control over data arriving later [1].
  • Anchor the business case in the macro number: 5 trillion dollars a year, nearly 20% of GDP, on a trajectory toward "the bankruptcy of the system overall" [1].

This page shows public professional information only, each fact cited. Is this you? send a correction, or ask for removal within 24 hours, no questions asked.