Reporting from the frontiers of health and medicine

AI patient profiling challenges the meaning of failed clinical trials

AI patient profiling challenges the meaning of failed clinical trials GenoMethods.org © genomethods.org
AI patient profiling challenges the meaning of failed clinical trials © genomethods.org
A failed Phase 3 trial does not always mean a drug is ineffective. Luca Pani of NetraMark argues that AI-driven patient selection could transform how clinical evidence is generated and interpreted.

When a Phase 3 clinical trial fails, the usual response is to write off the drug. But Luca Pani, Chief Innovation and Regulatory Officer at NetraMark, argues that the real problem often lies in how patients are selected. He believes the "average patient" is a statistical invention, and that drug development should focus on identifying who actually benefits—not just whether a drug works for the broadest group.

Pani brings deep regulatory experience to this argument. As a former Director General of the Italian Medicines Agency and a member of the European Medicines Agency Management Board, he has reviewed hundreds of evidence packages. At NetraMark, he is focused on a recurring problem: trials designed so broadly that they miss the very signals they are meant to detect.

In 2026, the FDA opened a public docket on statistical considerations for rare diseases and launched the RISE Workshop, highlighting a regulatory push for more nuanced trial designs and population selection.
Federal Register

Pani’s perspective shifted during the review of a major depressive disorder program. The Phase 2 trial looked promising, but the Phase 3 failed badly—Cohen’s D at 0.082, p-value 0.558. When NetraAI later analyzed the Phase 2 data, it found a clear responder profile. Applying this profile to the Phase 3 data changed the effect size to 0.346 and the p-value to 0.053. The trial still failed, but the analysis produced a testable hypothesis for future studies—a key distinction for both regulators and sponsors.

Why do companies keep running such broad Phase 3 trials? Pani is direct: commercial goals, scientific uncertainty, operational challenges, and regulatory caution all play a part. Companies want the largest possible label, but a diluted trial population can erase any sign of benefit. Biomarkers may not be validated or may only predict prognosis, not response. Narrowing eligibility makes recruitment difficult, so the default is to increase sample size instead of refining patient selection. Regulators want evidence that is neither too broad nor too narrow, and the complexity of human biology means single biomarkers rarely give a full answer.

Recent regulatory guidance reflects this tension. The FDA and other agencies have long encouraged stratification and enrichment strategies to select patients more likely to respond, rather than designing for an "average patient." According to the FDA CERSI Collaborative Research Projects, enrichment approaches and patient-focused endpoints remain a regulatory priority, with new clarifications released in 2026.


Industry experts in 2026 have directly linked the need for representative patient enrollment to the regulatory framework of each agency, reflecting increased scrutiny of trial design and endpoint alignment with regulatory requirements.
RAPS (Regulatory Affairs Professionals Society)

NetraMark’s approach is to use AI as a disciplined tool for building credible, interpretable responder profiles—not as a black box. For Pani, credibility is not just a p-value, but a full evidence package. The process starts with a clear decision: is the analysis for internal hypothesis generation, trial exclusion, or regulatory submission? Data quality is critical; machine learning cannot fix poor data. The model must separate true treatment benefit from placebo response, pass robustness checks, and be practical for real-world screening. Most importantly, the model should recognize when it cannot classify a patient, and accept that some cases are genuinely indeterminate.

AI-driven stratification is already changing how companies decide whether to move forward after Phase 2. Pani expects the old yes/no decision to be replaced by a range of options: broad or enriched populations, hierarchical testing, validation studies, or regulatory consultation. The FDA and EMA are moving toward human-focused, risk-based AI standards, and systematic analysis of treatment-effect differences is becoming standard in trial planning. The habit of "just increase the sample size" is fading. For example, recent data from a major depressive disorder Phase 3 program shows that pivotal trials often enroll about 450 patients to account for population differences, as reported in Xenon's press coverage.

There is also an ethical side. If a failed trial later reveals a subgroup that benefited, the industry faces a dilemma: you cannot promise individual patients they would have responded, but you cannot ignore the data either. The responsibility is to analyze carefully, communicate clearly, and avoid letting AI-driven exclusion rules become rigid. As Pani says, “We do not have sufficient knowledge to absolutely exclude any fellow human from the possibility of responding to a treatment, just because a machine says no.”

Clinical development is not about abandoning randomized trials or statistics—it is about making them smarter. The real opportunity is in patient selection. The future belongs to those who can spot differences early, turn them into clear hypotheses, and test them in new studies. The "average patient" may be a statistical idea, but the impact on real people is concrete. The next step for the industry will come from asking sharper questions about who truly benefits—not just from running bigger trials.

Elena MacLeod Clinical biotechnology and CAR-T editor GenoMethods.org
Biotechnology Newsroom

Elena MacLeod

Elena MacLeod is Clinical Biotechnology Editor at GenoMethods, covering CAR-T, engineered cell therapies, gene therapy, clinical trials, cancer immunology and regulatory developments. Her evidence-first reporting focuses on trial design, patient populations, safety, efficacy, response durability and the limitations that determine how early clinical results should be interpreted.