What CDMOs Get Wrong About AI for Bioprocessing, According to Sapiens Health’s Seungik Cho

“How good is your AI model? That’s the wrong question. The right question is: what does your AI model actually help your company do?”

Seungik Cho, founder of Sapiens Health, has spent the past year pulling apart the assumption in biomanufacturing that a better-predicting AI model means a better digital twin. His research argues the opposite: that optimizing AI for prediction accuracy can cause these systems to fail in CDMO environments.

Seungik Cho is the founder of Sapiens Health and Director of Strategic Operations at Nucleate. He’s finishing his B.S. in Biological Physics at Rice University, where his research explores the intersection between computational biology and multimodal AI. He describes his work as building AI systems that are “trustworthy and deployable,” not just accurate on a benchmark. His recent publications include work on longitudinal CT lesion prediction (IEEE ISBI 2026) and gene network analysis (IEEE BHI 2025).

In the latest PharmaSource podcast episode, Seungik explains why he believes CDMOs are pointing their AI investment in the wrong direction and lays out an alternative model built around diagnosis rather than prediction, explainability rather than confidence scores, and institutional memory rather than one-off forecasts.

Prediction-First AI Doesn’t Work in CDMO Bioprocessing

Seungik’s argument starts with what he calls the identifiability problem. In a fed-batch process running 200 to 300 hours, the signals that actually determine final yield don’t show up at the start. As he put it:

“The signals that determine your final yield don’t come from the initial stage. They come from the sensor data between maybe fifty to sixty hours, and nobody can really check or identify because they can’t stay all day for two hundred hours. The key quality comes from offline measurements that arrive days later.”

Asking a model to forecast final yield from early online sensors, he argues, is asking it to do something that isn’t reliably possible. Seungik tested this using PenSim, an industrial-scale penicillin fermentation benchmark. The results were interesting:

“In the training stage, the R-squared test was 0.997. It was incredible. But for the test R-squared level, it was minus 0.04. It’s just worse than predicting the mean. The model learned the noise during the training process, and it completely fell apart on the real batch processes.”

His conclusion: in small-data, highly regulated settings like CDMO bioprocessing, “predicting early” may simply be the wrong objective to optimize for.

The Causal Fidelity Gap: When Accurate Isn’t Correct

Even models with strong short-term prediction accuracy can fail a more basic test, Seungik found, asking what happens if a process variable changes. He calls this the causal fidelity gap, and it’s more subtle than the identifiability problem, but just as damaging:

“We asked those questions to the model — what if we increase the feed rate from 25% to 29%? In the ideal scenario, it should give the response back that the yield goes up, because we increased the feed rate. But the model says yield goes down. Those are not biologically coherent signals.”

A model that produces biologically incoherent answers to simple counterfactual questions is not trustworthy enough to inform a real operating decision, however good its short-horizon accuracy looks on paper.

Why Black-Box Models Can’t Survive a GMP Audit

For Seungik, the deeper issue with prediction-first AI is regulatory. Biomanufacturing isn’t a task where a confidence score and an output value are enough:

“Black box AI, which just gives the results back in a confidence level and a value, is wrong, because we don’t know which batch they tried to get the data from and were really trained from.”

Every AI-flagged deviation has to be defensible to a human reviewer and, ultimately, to a regulator:

“It must be justified by the QA reviewer — understanding why this AI model tried to generate this response, documenting it, and signing off on it. So those are the real-world things that CDMOs have to deal with. A black-box network that outputs a risk score — the metrics alone will not support that. We need a more explainable approach.”

He points to three specific pressures that make CDMO bioprocessing different from other AI-adjacent industries: data scarcity (companies often don’t keep records complete enough to train on safely), long, partially observed time horizons, and a regulatory environment where every AI output needs a documented rationale.

The Diagnosis-First Alternative

Rather than asking an AI model to predict outcomes, Seungik proposes reframing it as a detector of abnormal behavior:

“The reframing is: instead of treating your AI model as an oracle that tells you what the outcome will be, we can treat this AI model as a sensor for departures from normal behavior. We’re not asking, what will this yield be? We’re asking, is this batch behaving the way normal batches behave? And if not, what’s exactly different?”

He described the resulting architecture as three layers. The first is soft sensors, interpretable models like gradient-boosted regressors running on process variables and Raman spectroscopy data, producing real-time estimates that are explainable, versionable, and auditable by QA. The second is failure signatures: defining how a nominal batch behaves and tracking deviations from that pattern over time, such as oxygen uptake rate departures that signal a metabolic failure before it shows up in yield. The third is a knowledge graph that stores those signatures alongside recipes and outcomes, so that the pattern becomes searchable rather than living only in a senior engineer’s memory.

“So it’s more like a tractable problem in a small-data, regulated setting, which is suitable for CDMO environments. And it produces outputs that are actually actionable.”

Process Memory: Turning Institutional Knowledge Into a Queryable System

The knowledge-graph layer is where Seungik sees the most value; what he calls process memory. Today, that knowledge typically lives in the heads of long-tenured staff:

“Most CDMOs have institutional knowledge where senior managers have some know-how that they accumulated over the years. But often, when new people come in as a QA reviewer, there’s a gap between what seniors know and what new people know.”

His pitch is to make that memory queryable, the way an engineer would query a search engine:

“The failure signature has appeared in three prior campaigns. In campaign fourteen and twenty-three, the action was to reduce the feed rate by fifteen percent and hold the dissolved oxygen, and both batches recovered to acceptable yield. In campaign thirty-seven, there was no action taken, and the batch was out of spec.”

Crucially, Seungik is careful to explain that this is a support tool, not a replacement for human judgment: “It doesn’t eliminate the need for human judgment, and it won’t interfere with humans’ judgment. What it does is dramatically reduce the burden of investigation when QA reviewers have to dig for three hours to find previous documents.”

The Questions Every CDMO Should Ask About Explainable AI

Asked what a CDMO evaluating AI vendors should actually ask, Seungik offered three questions he believes cut through vendor pitches:

“First: if your model flags an anomaly or spots a deviation, what exactly can you tell me about why? … If the answer is just, ‘I predicted this using a regression model,’ that’s not good enough for GMP regulatory settings.”

“Second: how does your system behave in scenarios it hasn’t seen before? Can I give you a process perturbation that your training data doesn’t contain and get a biologically coherent answer back?”

“Third: what does your audit look like? If I need to produce a deviation report for a regulator, what does the system actually give me?”

His summary takeaway for the industry: “Prediction without explainability is not actionable. A dashboard without process memory doesn’t really make your organization smarter.” The useful version of AI, in his view, is one that helps a company “detect deviations earlier than I could before, in a way my process engineers can act on, and my QA team can document.”


EARLY BIRD TICKETS PRICES RISE SEP 18
Days
Hours
Minutes
Seconds