Margaret Gamalo, Pfizer
Editor’s Note: This originally appeared in the Biopharmaceutical Report and is reposted with permission. View the original article for references.

Picture the trial lifecycle, where many of you already live, but now with AI quietly managing much of the plumbing:
Before breakfast, a simulation engine sweeps through thousands of design variants—sample sizes, accrual curves, interim looks, stopping rules, and surfacing trade-offs we once uncovered only through repeated team meetings. By lunch, an AI drafting assistant proposes eligibility criteria with the pragmatism of a seasoned clinician. It flags contradictions you would rather catch now than at site initiation, and they highlight fairness or feasibility risks before they balloon into a screen-failure bonanza.
Three months later, an AI monitoring agent detects potential anomalies in the data, forwards them to clinicians and the trial manager for discussion, and updates Bayesian posteriors in near real time for safety and early blinded efficacy. It then nudges you when pre-specified rules are close and logs both the decision and reasoning behind it. A year or two later, a reproducible pipeline runs exactly as pre-specified, checking results against a synthetic twin for accuracy. And when you finally open the draft clinical study report, the tables, listings, figures, and narrative read like one coherent story rather than 12 appendices colliding at the printer.
If AI can take care of all that, what does it leave for you? Keep that question in mind as we step back and consider the broader context in which we work. In the end, my goal is to let you reflect on your own role in this evolving landscape.
AI has not changed who’s at the table—sponsors, regulators, payers, healthcare professionals, and patients—but it has tightened the clock and deepened the interdependence of every move. Drug development has always been cross-functional and buffeted by crosswinds; that has not changed. What has changed is the cadence. Discovery cycles compress. Data volumes explode. And decisions cascade faster across chemistry, manufacturing, preclinical, clinical, biostatistics, safety, market access, and controls (which define and prove the drug’s composition, and quality). The challenge has never been just speed; it has always been trust—at scale. Over the past two decades, we’ve navigated a persistent arc of headwinds:
- Waves of innovation that stretched capital and teams
- Shifting definitions of “value” across a broader set of stakeholders and debates over who defines it
- Regulatory complexity multiplied across regions
- Supply chain shocks that turned timing into a moving target with digitization, new data, and AI risks from patient privacy to model governance
These forces aren’t new, but the speed with which we meet them is. AI broadens the pipeline and accelerates discovery, amplifying both opportunity and interdependence—but constraints remain. Faster is not automatically better; it just means that we collide with the same limits sooner. Our mission does not change: Deliver breakthroughs people can trust at costs health systems can sustain. To do that, we pair acceleration with rigor—clear evidentiary standards, sound statistics, privacy- and quality-by-design principles, and transparent benefit-risk communication. Moving faster should also mean moving safer.
The momentum of AI in discovery is real. Analyst estimates vary, but many project a substantial share—possibly around 30%—of new drug programs discovered in these next few years will be AI-enabled in some way. For 2024–2029, growth projections for global AI-in-drug-discovery range around 25–30% annually. It will be fueled by cost/time pressures, broader AI adoption, exploding life-science data and compute, pharma–AI partnerships, looming patent cliffs, generative-AI–enabled design, and demand for personalized medicine.
Open any life-science feed (STAT, Endpoints, Pink Sheet, even LinkedIn) and you’ll see AI’s fingerprints across the stack—big tech-big bio tie-ups, foundation models moving from structure prediction to de novo design, and university-industry consortia accelerating target/chemistry workflows. However, as AI compresses discovery timelines, development must adapt in step.
There’s no stop sign, but the playbook must evolve. The future is modular, risk-tiered investigational new drugs; predictive and in-silico toxicology with auditable error bounds; manufacturing process acceleration and bridging; and adaptive designs that unify dose escalation, cohort expansion, and early proof-of-concept—especially outside oncology. Speed will no longer be exceptional; it will be expected. That means the infrastructure around it—pre-clinical, clinical, statistical, regulatory, and operational—must mature in parallel.
In early development, modular investigational new drugs could open first-in-human studies with core pharmacokinetics and short-term toxicology, layering long-term studies and special populations as data mature. US sponsors often face slower Phase I entry because first-in-human authorization can default to a one-size-fits-all process, while some regions allow faster starts for clearly lower-risk programs. A balanced fix is a formal risk-tiered first-in-human pathway that links data-package and protocol safeguards (e.g., MABEL-based starts, sentinel/staggered dosing, exposure caps, real-time stopping rules), so low-risk assets move faster, while high-risk first-in-class agents remain under enhanced protection. This is in line with the commentary by Scott Gottlieb. Predictive or in-silico toxicology can complement animal studies, provided their models are transparently validated and bounded by measurable error. Risk-tier frameworks may also emerge, where lower-risk or well-characterized modalities qualify for streamlined investigational new drugs, while first-in-class or high-uncertainty compounds maintain full preclinical requirements. As more candidates reach first-in-human adaptive designs that merge dose escalation, cohort expansion and early proof-of-concept will become essential—particularly in chronic, non-oncology settings. Statistics provide guardrails that keep this speed trustworthy by defining operating characteristics, quantifying uncertainty, and preventing repeats of “rush-to-clinic” failures like TGN1412, a novel superagonist anti-CD28 monoclonal antibody. Acceleration is only progress if it remains safe, auditable, and scientifically sound.
Payers and health technology assessment organizations have also moved upstream. In the EU, the Joint Clinical Assessment forces early alignment on PICO (Population, Intervention, Comparator, Outcomes) and comparators. PICO forces a decision-relevant question that matches real clinical practice, and the comparator is essential to estimate relative effectiveness and cost-effectiveness rather than absolute performance. Without an appropriate, justified comparator, health technology assessment organization results risk bias, poor transferability, and conclusions not actionable for payers or guideline bodies. In the US, the Medicare Transitional Coverage for Emerging Technologies pathway enables earlier, conditional coverage tied to post-market evidence plans. Meanwhile, in the UK, the NICE Early Value Assessment offers provisional adoption with explicit evidence commitments. For statisticians, that means designing for access early. Payer-relevant estimands must sit alongside primary endpoints and measures—such as time-to-next treatment, hospital-free days, and treatment-free intervals. Resource use must be built in, not bolted on. When access is conditional, real-world evidence programs—such as registries, burden, and utilization studies—must be established early as target trial emulations with pre-analysis protocols and transportability controls.
Patients are changing, too. Many now arrive as informed consumers: AI and natural language processing help prescreen eligibility; “blue button” tools surface nearby trials; and patient portals reveal time and travel costs, along with remote visit options. In this emerging marketplace of trial choice, ranking must be fair, explainable, and resistant to gaming. And the trial burden must be explicitly modeled, or feasibility projections will fail. The statistical toolkit expands accordingly: uplift modeling to estimate incremental recruitment benefit; constrained bandits to allocate patients fairly under burden and capacity limits; conjoint analysis to quantify real-world trade-offs; heterogeneous-treatment-effect modeling to identify who truly benefits; and target-trial emulation to ensure resulting claims remain grounded in clinical reality.
Together, these strands create a new equilibrium. Sponsors win through global pipeline partnerships and randomized evidence packaged with AI-informed post-market loops that continuously earn trust. Regulators converge on flexible, risk-tiered investigational new drugs, keep randomized trials as anchors of truth, and use AI with randomized controlled trial–real-world evidence embedding to extend generalizability (ideally harmonized through International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use guidance). Payers press for conditional reimbursement paired with AI-enabled real-world monitoring, while patients increasingly act as informed consumers. In this ecosystem, systems biostatistics become the connective tissue of evidence: aligning estimands with regulatory and payer decisions, architecting adaptive designs and simulations, and synthesizing randomized controlled trial and real-world evidence under explicit assumptions and sensitivity analyses. We don’t eliminate bias; we expose, mitigate, and quantify it so that every choice about benefit, risk, and access remains fair, transparent, and auditable.
As the scientific and regulatory landscape becomes more complex—with integrated data streams, divergent global frameworks, and accelerating decision cycles—the role of statisticians becomes more vital than ever. Our discipline anchors evidence amid volatility and complexity. First, signal versus noise: The convergence of omics, clinical, electronic health record, and claims data generates a torrent of patterns, and statisticians discern truth from coincidence. Second, regulatory credibility: If a model is not interpretable, validated, and auditable, it is not deployable. Third, integration complexity: Without causal structure, multimodal data degenerates into a decorative quilt of bias. We establish the weights and guardrails that preserve inferential integrity. Fourth, decision risk: As fragmentation increases, so does the cost of error; we quantify trade-offs so leadership can decide with clarity and confidence. Fifth, ethics and fairness: When an algorithm systematically underserves a subgroup, it’s not a technical flaw, it’s an ethical failure and the responsibility to detect and correct it lies with us.
Statistical stewardship requires embedding statistical principles within AI systems rather than treating AI as an opaque instrument. Causal inference must reside within predictive pipelines; bias correction must occur where it alters actions; and transportability must be made explicit rather than assumed.
Statistical stewardship requires embedding statistical principles within AI systems rather than treating AI as an opaque instrument. Causal inference must reside within predictive pipelines; bias correction must occur where it alters actions; and transportability must be made explicit rather than assumed. We construct operating-characteristic frameworks that stress test trial designs and portfolios against population shifts, supply disruptions, enrollment volatility, and patient nonadherence. We translate the models into evidence that satisfies International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use, US Food and Drug Administration, European Medicines Agency, National Medical Products Administration, and regional regulatory expectations. And we insist on explainability that can be interrogated and replayed—what drove the decision, what alternatives were considered, and how conclusions evolve when assumptions change. The distinction between tooling and stewardship lies in judgment—in knowing when the right answer is a better model and when it is a better question.
Used well, large language models are accelerators, not autopilots. In a systems-biostatistics workflow, they manage the infrastructure—the drafts, retrieval, and code scaffolds—so human time is spent on judgment and decision-making rather than operational assembly. They can outline derivations and simulations; propose eligibility criteria that we re-rank for coverage and fairness; and generate draft protocol sections, statistical analysis plan shells, and data monitoring committee charters anchored in precedent. They can perform structured and quality-focused reviews of analysis plans for alignment between specification and data. They can also scaffold real-world evidence studies with confounder libraries, directed acyclic graphs, and sensitivity panels tailored to payer questions. Retrieval-augmented generation keeps analyses grounded in precedent rather than speculation.
Acceleration, however, requires brakes. If a language model can influence anything that touches a patient, it must be governed like a medical device: documented, monitored, versioned, and equipped to abstain when confidence is low.
Acceleration, however, requires brakes. If a language model can influence anything that touches a patient, it must be governed like a medical device: documented, monitored, versioned, and equipped to abstain when confidence is low. Hallucinations and overconfidence should be limited by automated fact checking against verified knowledge sources, calibrated uncertainty, and conformal prediction (a statistical framework that provides a way to make reliable predictions with a guaranteed level of confidence). Guardrail erosion in extended interactions requires conversation-state monitoring and adversarial stress testing. Bias and harmful content require structured audits, counterfactual testing, and fairness-constrained training or reranking to preserve subgroup equity. Large language models can automate plumbing, but never judgment.
These all point to who we must become. Think of a “T”-skilled statistician. The horizontal bar of the “T” represents breadth, the ability to think throughout the system from molecule to market; trial-to-access, seeing the full chessboard of regulators, payers, investigators, supply chains; and most importantly, patients. The vertical stem represents depth: subject matter expertise in disease biology; endpoints; trial operations; health technology assessment; and the complex mathematical foundations that connect them.
We are the ones who recognize when a data collection plan invites missingness, when an endpoint lacks sensitivity, or when a method rests on unverifiable assumptions. Then, we propose what will work instead. The boldface “T” is a reminder to act boldly, but with guardrails: Use large language models for drafting and scaffolding, yet insist on calibration, provenance, and pre-specified rules for trustworthiness. Automate the plumbing; never automate the judgment.
In a systems-biostatistics model, our role is not diminished by AI; in fact, it is strengthened to enable sound decisions in an accelerated world. We design evidence that speaks to both regulators and HTA bodies. We weave randomized and real-world evidence into coherent narratives. We stress test portfolios against disruption and fragmentation. We keep fairness, interpretability, and safety visible in every adaptive step. So, what remains for statisticians in an AI-accelerated world? Everything that matters in this new board game. Think broadly. Integrate deeply. Act boldly—with guardrails. Do that, and we will not simply keep up with AI. We will ensure it delivers what truly counts: better, faster, more trustworthy outcomes for patients.

Leave a Reply