
How Standardized Cell Labeling Could Fix Biotech's Data Problem
Flow cytometry can find a cell population in seconds, but naming it consistently across labs is a problem still unsolved. If you've ever tried to compare cytometry data across studies, CROs, or even two scientists in the same lab, you know the frustration: everyone gates their own way, and the same cell population ends up with three different definitions. That inconsistency is quietly limiting what AI can do with biotech data. Ryan Brinkman is VP Research Director of Flow Cytometry Bioinformatics at Dotmatics and Founding Director of SOULCAP, the Standard Ontology for Unambiguous Labeling in Cytometry and Phenotyping, built after years as an academic developing automated gating tools that kept running into the same naming problem. He's joined by Brian Wile, who is the General Manager of Flow Cytometry at KCAS Bio, a CRO that feels the cost of inconsistent labeling in client projects every day. You'll get a clear picture of how flow cytometry data moves from raw signal to labeled cell population, why that last step has resisted automation, and what a shared standard could unlock for machine learning models trained on this data. Ryan and Brian break down the gap between automated gating and consistent labeling, and why agreement, not just data volume, is what AI in biotech actually needs. This episode covers the mechanics of flow cytometry, an EVE Online citizen science project that trained a gating algorithm on hundreds of millions of human-labeled plots, and why cell population names like "Treg" or "natural killer cell" don't map to one agreed set of markers. Key Takeaways Gating automation solved half the problem: computers can now draw boundaries around cell populations, but nothing proves which name belongs on the result. A citizen science project turned hundreds of thousands of EVE Online players into an unlikely training set, collecting roughly half a billion labeled plots from people with zero biology background. The same cell type, like a regulatory T cell, gets defined by entirely different marker combinations depending on the lab or CRO, and those datasets can't simply be merged later. SOULCAP is running a Delphi-style consensus process among scientists worldwide to tie every cell population name to a specific, reproducible set of markers and experiments. Chapter Markers 00:00 Why naming cells is harder than naming genes 02:13 What flow cytometry measures, from cell to signal 06:21 The scale and complexity of high-dimensional cytometry data 09:20 How raw data becomes a gated cell population 12:02 Why gating has stayed manual for decades 16:19 Clusters of differentiation and the marker explosion 18:11 Training an algorithm with EVE Online players 25:31 Evaluating accuracy without a gold standard 31:34 Why solving gating doesn't solve labeling 36:32 What genomics got right that cytometry hasn't 38:09 Two definitions of a regulatory T cell, one label 44:28 The cost of inconsistent labeling for CROs and pharma 47:43 Introducing SoCAP and the Delphi consensus process 57:18 What automated labeling could make possible 65:38 How to get involved in SOULCAP Useful Links & Resources Ryan Brinkman on LinkedIn: https://www.linkedin.com/in/ryan-brinkman-9bb1103/ Brian Wile on LinkedIn: https://www.linkedin.com/in/brianwile/ KCAS Bio: https://www.kcasbio.com CorrDyn: https://www.corrdyn.com Connect With the Show Ross Katz on LinkedIn: https://www.linkedin.com/in/b-ross-katz/ CorrDyn LinkedIn: https://www.linkedin.com/company/corrdyn/ Where do you land on the naming problem? Have you had to merge cytometry datasets that turned out to use different definitions for the same cell type? Tell us about it in the comments. Visit corrdyn.com to learn how CorrDyn can help your organization extract value from data. #DataInBiotech #FlowCytometry #BiotechDataScience #Bioinformatics #LifeSciences
