• Skip to main content
  • Skip to secondary menu
  • Skip to primary sidebar
  • Skip to footer
  • Homepage
  • About Us
  • Advertising
  • Submission Instructions
  • Editorial Calendar
Amstat News

Amstat News

The Membership Magazine of the American Statistical Association

  • Printed Issues
  • Practical Significance Podcast
  • Additional Features
  • Columns
  • Member News
  • Departments
You are here: Home / Additional Features / The Evolving Role of Statisticians in the Pharmaceutical Industry: Leveraging Advanced Statistical Analytics and Artificial Intelligence 

The Evolving Role of Statisticians in the Pharmaceutical Industry: Leveraging Advanced Statistical Analytics and Artificial Intelligence 

July 1, 2026 Leave a Comment

A stethoscope sitting on top of a printed graph.
Haoda Fu, Amgen, and H. Amy Xia, Amgen 

Editor’s Note: This article originally appeared in the Biopharmaceutical Report and is reposted with permission. View the original article for references. 

Haoda Fu's headshot. He has short black hair and is wearing glasses and a dark suit jacket.
Haoda Fu

Pharmaceutical statisticians have come a long way over the past half century, evolving from backroom number-crunchers to essential contributors across the entire drug development spectrum. Once viewed primarily as support staff ensuring regulatory compliance, statisticians today are equal partners in research and development teams. They influence decisions from early drug discovery and clinical development to manufacturing and commercialization. This expanded role has been driven by multiple converging forces. Advances in computing and new data sources (e.g., genomics, real-world clinical data) have enabled innovative statistical methodologies. Meanwhile, the rise of AI and machine learning offers powerful tools to extract insights that were previously difficult to analyze in electronic information. 

Amy Xia's headshot. She has short brown hair and is wearing a white jacket.
Amy Xia

At the same time, the pharmaceutical industry’s external environment has grown more challenging. With increasing costs, fewer new therapies are approved each year and stakeholders demand greater transparency and evidence of value. Statisticians have responded by embracing new analytic techniques and stepping into leadership and collaboration roles that were virtually unheard of decades ago. Ultimately, we aim to demonstrate how statisticians not only adapt to an evolving landscape, but also increasingly lead innovation in pharmaceutical R&D and beyond. 

Historical Context of Statisticians’ Roles in the Pharmaceutical Industry 

The use of data and statistics to improve patient outcomes has been part of healthcare for thousands of years and remains crucial today. An early example of data-driven healthcare comes from the Book of Daniel in the Bible. In 500 BC, King Nebuchadnezzar of Babylon believed a diet of meat and wine would keep his people healthy. However, some young men chose to eat vegetables and drink water for 10 days. They appeared healthier, so the king allowed them to continue their diet. This was an early instance of using an experiment to make a health decision. In the 18th century, James Lind, a ship’s surgeon, conducted one of the first controlled clinical trials. He tested treatments for scurvy and found that oranges and lemons were effective. 

Modern biostatistics in drug development began in 1946 with the introduction of randomization and controlled trials. Randomization was first introduced in 1923, and Sir Austin Bradford Hill conducted the first randomized controlled trial in 1946, showing that streptomycin was effective for tuberculosis. This study demonstrated how randomization, control groups, and statistical testing could guide medical decisions. A significant change came with the 1962 amendments to the US Food, Drug, and Cosmetic Act, following the thalidomide tragedy. These amendments required the FDA to demand “substantial evidence” from controlled trials to prove a drug’s effectiveness, not just safety. This led drug companies to realize the necessity of statistically designed trials for approval, leading to a surge in hiring statisticians to meet FDA requirements. By the late ’60s and ’70s, statisticians were key members of clinical research teams, mainly designing trials, calculating sample sizes, and analyzing data for regulatory submissions. 

In the ’70s and ’80s, the role of statisticians in pharma grew with new regulatory initiatives. A key development was the FDA’s New Drug Application rewrite in the early ’80s, which required a formal statistical review for every new drug application and a statistician as a co-author of clinical trial reports. These changes solidified statisticians’ roles in the drug approval process. However, they were still seen as technical support, ensuring analyses were correct and compliant. As Frank Wesley Rockhold noted, even after NDA reforms, statisticians mainly executed analyses and calculated sample sizes, rather than shape study designs or development programs. Most focused on late-phase clinical trials and manufacturing quality assessments, with little involvement in early research phases or nonclinical areas. 

By the ’90s, several factors increased statisticians’ influence. Pharmaceutical R&D became more global and complex, with larger trials and more data. Regulatory agencies worldwide adopted harmonized standards for trial conduct and statistical practice. The International Conference on Harmonization issued guideline E9: Statistical Principles for Clinical Trials, emphasizing the importance of statistics in trial design, analysis, and interpretation. According to Rockhold, ICH E9 gave statisticians more “leverage and authority in drug development,” highlighting the need for a strong statistical foundation for credible evidence. Statisticians began contributing strategically, advising on clinical programs and study designs. The industry also recognized that information is the key output of R&D, boosting the demand for statistical thinking to maximize data value in discovery, preclinical studies, clinical trials, and post-market surveillance. 

Another milestone in the ’90s was the rise of powerful statistical software and personal computing that enabled advanced analyses and simulations. Statistical programming languages like SAS became essential tools for pharma statisticians. In the late ’90s and early ’00s, statisticians expanded into new areas: safety data mining for adverse event detection; support for epidemiological studies; and clinical pharmacology modeling (e.g., PK/PD analyses for dose selection).  

In the ’00s and ’10s, the statistician’s role expanded significantly. The FDA’s 2004 Critical Path Initiative aimed to modernize medical product development science, advocating for innovative statistical approaches. The initiative highlighted challenges like biomarker validation, enrichment trial designs, missing data handling, multiplicity issues, and model-based evidence. These all required sophisticated statistical input.  

In the following decades, regulators released guidance documents on adaptive trial designs, noninferiority trials, multiple endpoints, and real-world evidence, expanding statisticians’ toolkit and responsibilities in clinical development. By the ’10s, statisticians were regarded as “absolutely critical for efficient and effective drug development,” serving as key contributors or consultants in all R&D areas. The role evolved from a support role to a strategic, interdisciplinary one that positioned statisticians to tackle 21st-century challenges, including the big data revolution and AI integration in pharmaceutical research. 

The following key catalysts have accelerated this evolution of statisticians’ responsibilities and skill sets:  

  • Advances in computing and software that exponentially widened analytic possibilities 
  • Development of innovative statistical methodologies, coupled with regulatory encouragement that fostered adoption 
  • Emergence of new data types and large datasets that demanded novel analytical approaches  

Advances in Statistical Computing and Hardware 

The practice of statistics in pharma has changed markedly in tandem with advances in computational power.  

Early pharmaceutical statisticians worked in an era of limited computing power, often performing calculations by hand or with basic mechanical aids. The mid-20th century saw the introduction of mainframe computers, but computational resources remained scarce and specialized. This inherently constrained the complexity of analyses that statisticians could practically undertake. Over time, however, revolutions in computing hardware and the advent of statistical software radically transformed the toolkit of the pharmaceutical statistician. By the late 20th century, improvements in processing speed and data storage (following Moore’s Law) enabled routine execution of intensive methods that were previously impractical. In parallel, the development of high-level statistical programming languages and software packages—the Statistical Analysis System in the ’70s and the open-source S language (and later R) in the ’90s—provided user-friendly platforms to implement complex analyses. Widespread adoption of these tools meant statisticians could manage larger datasets and apply more sophisticated models with relative ease.  

One direct outcome of these advances was the rise of simulation-based analysis and design. With greater computing resources, statisticians began to use Monte Carlo simulations to evaluate trial properties and optimize study designs before any patients were enrolled. By the ’00s, it became routine to simulate thousands of trial iterations to assess a design’s probability of making correct decisions or to model various what-if scenarios. Such computationally intensive work was not feasible in earlier decades.  

The increasing availability of fast computing also facilitated resampling and other modern methods. Techniques like the bootstrap (for estimating confidence intervals) and Markov chain Monte Carlo (for Bayesian analysis) gained traction in clinical research once computers were up to the task. The net effect was an expansion in statisticians’ capabilities: Rather than being limited to relatively simple trial designs and analyses, they could now explore a much richer design space and fit more complex models to data.  

Indeed, contemporary statisticians often write extensive code in SAS, R, or Python to manipulate datasets, implement custom analyses, and even create interactive dashboards for data visualization—a blending of traditional statistical skills with what we now call data science. 

Better computing didn’t just change how fast statisticians work—it expanded what they work on. Previously, statisticians’ contributions might have begun only after data collection (analyzing final trial results). However, modern computing power allows them to influence studies from the planning and design stage onward, running simulations to inform optimal sample sizes, endpoint definitions, and decision criteria. For instance, clinical trial simulation became an established practice for complex trial planning by the ’10s, allowing statisticians to quantify the trade-offs of various design choices under myriad scenarios. As datasets grew from tens of patients in the ’60s to tens of thousands of patients (or millions of observations) in the 21st century, the statistician’s role expanded to include ensuring data integrity, traceability, and reproducibility through efficient programming and validation. The history of success with SAS as a de facto industry standard is one testament to how important computing environments became to pharma statistics. More recently, open-source tools (R and Python in particular) have gained acceptance, further empowering statisticians to use cutting-edge techniques and share reproducible code. The dramatic improvements in hardware and statistical software from 1950 to present have been fundamental catalysts for transforming the statistician’s role from a manual calculator of p-values to a computational strategist capable of exploring vast design and analytic possibilities. 

The next wave of computing advances will continue to shape the statistician’s role. For example, current trial simulations primarily focus on addressing scientific questions like family-wise type I error control, power, or the posterior probability of trial success. These simulations often come before running a clinical trial. As computing power grows, we expect statisticians will increasingly leverage real-time data during ongoing clinical trials to run simulations that address not only scientific questions, but also operational questions like the impact of opening additional sites to speed up enrollment. Addressing these questions can lead to optimizing clinical trial operations. We think statisticians collaborating with cross-functional teams and designing and implementing real-time simulations will be key to the next generation of clinical trials. 

Growth of Advanced Statistical Methodologies and Regulatory Encouragement 

As computing capabilities grew, so did the development of novel statistical methodologies for clinical trials. From the ’80s onward, statisticians began proposing innovative trial designs and analysis methods that could make drug development more efficient and informative. Two prominent examples are adaptive trial designs and increasing use of Bayesian statistical methods. 

These innovative designs represented a break from the fixed, one-size-fits-all designs that dominated clinical research since standardizing randomized controlled trials in the post-war era. At the same time, the industry has recognized the increasing cost of drug development and the need to improve efficiency. However, the uptake of these innovations in industry was initially slow, until regulatory bodies, especially the FDA, actively encouraged adoption. Regulatory guidance has been a crucial catalyst in legitimizing and accelerating the use of advanced methods by pharmaceutical statisticians.  

Adaptive designs allow pre-planned modifications to certain aspects of clinical trials (sample size, randomization ratios, treatment arms) based on interim analysis of accumulating data. The conceptual appeal of adaptive trials is clear: They can make clinical research more flexible, efficient, and potentially help find effective treatments faster or use fewer patients.  

For example, rather than stick to a static design, an adaptive trial might start with multiple dose groups and use interim results to seamlessly drop ineffective doses or reallocate more patients to promising treatments. By utilizing ongoing results, adaptive designs can ethically benefit patients (more patients get the better treatments) and scientifically improve the chance of trial success or reduce needed resources. These advantages were recognized in the statistical literature by the ’90s, but there was early hesitation in the conservative regulatory environment to accept trials that departed from the traditional fixed protocol. This began to change in the 2000s and 2010s. A milestone was the FDA’s 2010 draft guidance on adaptive design, followed by comprehensive FDA Guidance in 2019 that explicitly outlined principles for adaptive trials. This guidance not only provided the industry with a clear roadmap on how to plan and analyze adaptive trials, but also it sent a strong signal that regulators welcome well-justified adaptive approaches.  

Currently, an ICH E20 guidance on adaptive design for clinical trials is underway to delineate the principles of adaptive designs and regulatory considerations. Statisticians were central to this shift: They had to develop new statistical methods to ensure, for instance, that making mid-course modifications would not inflate the family-wise type I error (false positive rate). They also engaged in extensive simulations, as recommended by FDA, to demonstrate operating characteristics of adaptive designs before implementation. As a result of these efforts, adaptive designs are now increasingly common in clinical trials across therapeutic areas (from oncology to cardiology), and pharmaceutical statisticians have expanded responsibilities in designing interim analyses, setting adaptation rules, and liaising with data monitoring committees. Adaptive methods have moved from an experimental idea to a mainstream tool, catalyzed by regulatory acceptance. 

Bayesian methods have similarly grown in prominence. The Bayesian framework for data analysis offers an intuitive and flexible approach, where evidence is accumulated sequentially and prior knowledge can be formally incorporated into current trial analysis. For decades, classical (frequentist) statistics dominated drug trials, but Bayesian statistics began gaining traction for problems where traditional methods were less efficient. This included trials in rare diseases, early-phase studies requiring use of prior data, and other applications in safety signal detection and meta-experimental design and analysis. Bayesian analyses can produce direct probability statements about treatment effects (e.g., the probability a drug is better than control), which are appealing to decision-makers. They also allow for more continuous learning from data rather than an all-or-nothing hypothesis test. Bayesian methods like probability of study success evaluation have been used for internal decision-making. 

However, adopting Bayesian approaches in regulated clinical trials required convincing both scientists and regulators of their validity and robustness. A key turning point was in medical devices. In 2010, the FDA’s Center for Devices and Radiological Health released guidance for the use of Bayesian statistics in medical device trials. This document explicitly acknowledged that Bayesian methods, when properly applied, could reduce required sample sizes or study durations by incorporating prior evidence. They also offered benefits in trial design flexibility. Notably, by formally addressing the question, “Why are Bayesian methods more commonly used now?” and similar questions, FDA guidance clarified misconceptions and provided best practices for sponsors.  

This endorsement catalyzed a surge of interest in Bayesian designs not only for devices, but also eventually in drug trials. In drug development, Bayesian methods have seen increased use in exploratory Phase II trials, adaptive dose-finding (e.g., Bayesian dose-escalation methods in oncology), and confirmatory trials with regulatory acceptance (especially in rare disease settings, where leveraging external or prior trial data is invaluable).  

For example, the pivotal Pfizer/BioNTech mRNA COVID-19 vaccine study (BNT162b2) employed a design and analysis framework described as Bayesian. Notably, in autoimmune disease development, Amgen’s program for systemic lupus erythematosus entered the FDA’s complex innovative trial design pilot program, proposing that endpoint will be evaluated using a Bayesian hierarchical model with non-informative priors. By 2024, a Lancet review advocated for broader use in clinical research, noting that Bayesian statistics offered a flexible and informative approach that facilitated both design and interpretation of trials. The authors emphasized that owing to its different conception of probability, the Bayesian paradigm can incorporate evidence in ways that enrich inference and decision-making.  

FDA leadership also highlighted Bayesian and adaptive designs as promising innovations for modernizing clinical trials, in discussing the 21st Century Cures Act, which encouraged exploration of novel trial designs and analytical methods for speeding therapy approvals. Recently, the FDA’s Center for Drug Evaluation and Research launched the Bayesian Statistical Analysis Demonstration Project to foster the use of Bayesian methods in “simple” phase-III drug trials (e.g., non-adaptive or sequential designs). The program allows sponsors to use Bayesian analyses as either primary or supplemental analysis. It also offers regulatory interaction and methodological support.  

In practice, statisticians’ roles have expanded to include mastering these advanced methodologies, educating project teams and regulators about them, and developing technical justifications needed for their use. Where a ’70s-era statistician’s toolkit might not have extended beyond t-tests and chi-squares, a statistician today might design a complex adaptive Bayesian trial with multiple interim looks and dynamic randomization.  

Another example of methodological innovation is the emergence of master protocols (platform trials, basket trials, umbrella trials), which allow evaluation of multiple therapies and/or multiple diseases within a single trial infrastructure. These designs, which became especially prominent in the ’10s (notably in oncology), require sophisticated statistical coordination, such as sharing control groups, dropping or adding treatment arms on the fly, and possibly using Bayesian borrowing of information across sub-studies.  

Statisticians were instrumental in conceiving these designs, but their broad adoption was again facilitated by regulators. In 2017, Janet Woodcock and Lisa LaVange from the FDA authored a New England Journal of Medicine review explaining the value of master protocols and providing a regulatory perspective on how to conduct them rigorously. They illustrated that such designs can accelerate drug development by studying multiple hypotheses under a common protocol. But they also cautioned that statistical complexities must be managed. Following this, the FDA issued a formal guidance on master protocol trials, further cementing regulatory encouragement.  

The net effect of these trends is that statisticians are now far more deeply involved in trial design strategy. They’re not just answering “How do we analyze the data?” They’re also answering “What is the optimal way to design this study to begin with?” As a 2010 industry review, “Statisticians in the Pharmaceutical Industry: The 21st Century” put it, statisticians in pharma have evolved into “full and equal partners with clinical and regulatory scientists” in trial planning and drug development strategy.  

This cultural shift means statisticians today often co-lead discussions on a program’s evidence generation plans. They ensure innovative designs like adaptive and Bayesian trials are used appropriately and transparently, satisfying scientific rigor and regulatory standards.  

In summary, the growth of advanced methodologies—and the feedback loop of regulatory guidance and endorsement—has been a key catalyst in expanding statisticians’ responsibilities. It pushed them into new roles: methodological innovators, architects of novel trial designs, and front-line communicators who articulate the benefits and limitations of designs to regulators and clinical teams. 

New Data Types and Large Datasets 

The modern pharmaceutical landscape is awash with data sources that scarcely existed a few decades ago. In early times (’50s–’80s), clinical trial results captured on paper case report forms were the primary data source for statisticians. With limited computation tools, analysis was often restricted to basic descriptive statistics and simple statistical methods, such as t-tests, ANOVA, and chi-squared tests. Later, longitudinal data from electronic CRFs and databases became more common, allowing for richer analyses of treatment effects over time, methods like mixed-effects models, and more complex statistical techniques.  

Nowadays, companies contend with real-world data from healthcare databases, genomic, other “omic” data from advanced laboratory technologies and wearable sensors and patient devices. The advent of these new data types—often high-volume, high-velocity, and high-variety—changed the statistician’s job. Statisticians have had to develop and adopt new methodologies, expand their expertise into realms traditionally outside classical biostatistics, and collaborate closely with experts in fields like bioinformatics and machine learning. The rise of large, complex data sets has broadened the statistician’s role from trial-centric analysis to a more holistic “clinical data science” role. 

One key area is real-world data and real-world evidence. Real-world data refers to patient health status and/or delivery of healthcare routinely collected from a variety of sources. It could include electronic health records, claims and billing data, data from product and disease registries, patient-generated data, and data gathered from sources like mobile devices. 

Historically, real-world data was not heavily used in regulatory decisions due to concerns about bias and quality. However, efforts in the US have increased use of real-world evidence for regulatory and clinical insights. The 21st Century Cures Act of 2016 required the FDA to explore real-world evidence for drug approvals. By 2021–2022, the FDA had issued guidance on using real-world evidence and approved some drugs based on real-world studies.  

Statisticians play a crucial role in analyzing these datasets and dealing with issues such as bias, confounding and missing data, and data quality. They must also explain to regulators how observational data approximates randomized trial evidence. This expansion means statisticians now work in outcomes research, safety surveillance, and policy. The FDA’s 2021 guidance on using electronic health records for regulatory decisions further expands their responsibilities, and since 2021, FDA has published a series of guidance documents related to real-world data and evidence in data, design, conduct, and regulations. 

Advances like genome sequencing have introduced large-scale data to drug development. These data are complex and require sophisticated modeling. Statisticians have been key in developing tools to analyze this data, contributing to bioinformatics. They design experiments, pre-process data, and develop algorithms to identify important genes or biomarkers. As precision medicine grows, statisticians will help identify patient subgroups with genomic markers that predict drug response. They also work on companion diagnostics, linking biomarkers to treatment outcomes. Their role has expanded from asking, ”Does the drug work on average?” to ”For whom does the drug work?” This role requires collaboration with lab scientists and skills in multivariate modeling and machine learning. Statisticians also ensure data validity in new analytical domains. 

Digital health data is another new area. Devices like smartphones and wearables collect real-time patient data, creating “digital endpoints” in trials. These endpoints offer a more comprehensive view of patient health. Statisticians validate and analyze these endpoints, addressing challenges like data volume and missing data. They work with clinicians to ensure digital measures correlate with clinical benefits. The FDA has shown interest in digital health technologies, issuing guidance on digital tools in trials. The COVID-19 pandemic increased the acceptance of digital endpoints. Statisticians now work with data scientists to refine algorithms and design trials with remote data capture. 

Finally, underpinning all these new data domains is the rise of AI and machine learning in drug development. Pharmaceutical companies are increasingly using ML models for tasks ranging from drug discovery (e.g., predicting molecule-target interactions) to patient/site selection and outcome prediction in clinical trials. Statisticians often lead this work, collaborating closely or leading on the validation of such models.  

Notably, regulatory agencies have begun to acknowledge AI/ML in submissions. By 2025, the FDA reported seeing over 500 product submissions (across drugs and biologics) that incorporated AI/ML approaches, spanning discovery, trial optimization, and post-market safety analysis. This marks a significant new responsibility for statisticians: evaluating and perhaps even developing predictive algorithms to ensure they meet standards of evidence and lack undue bias.  

The FDA has encouraged sponsors to employ cutting-edge analytical tools, such as using ML on real-world data to detect safety signals or interpret complex endpoints. As such, statisticians now collaborate on cross-functional teams and contribute expertise in the validation of applying principled cross-validation, setting up prospective validation studies for algorithms, and quantifying uncertainty in model predictions.  

The data science revolution has not obviated the need for statisticians—it’s expanded their purview. This breadth is a direct consequence of the influx of novel data types that require novel analytic thinking and methods. 

Emerging AI Technologies in Pharmaceutical Research 

Perhaps the most transformative catalyst in recent years has been the rise of artificial intelligence and machine learning in pharmaceutical research. AI technologies are reshaping how data are generated, analyzed, and how trials are conducted. 

In the past, pharmaceutical data was mostly just numbers in tables. Now, AI has expanded what we consider “data” to include things like molecular sequences, medical images, text, and audio. Machine learning models can now learn from unstructured information that couldn’t be analyzed before. For example, large language models are trained on diverse text sources like Wikipedia, which shows how text can be turned into valuable scientific data. Similarly, AI in biotechnology has used protein databases to create new functional proteins in a computer. This means protein databases once used for manual searches are now used for AI-driven protein design, allowing us to create enzymes with specific functions from scratch. These examples show how the concept of “data” is growing and how new data types are driving unexpected advances in pharmaceuticals. 

Statisticians play an important role in understanding this flood of digital data. AI gives statisticians the chance to work with complex datasets that were too difficult to analyze before. There’s a clear path for turning raw data into useful insights:  

  • Digitalization: Turning paper records into digital form 
  • Datafication: Organizing these digital records so they can be analyzed 
  • Knowledgefication: Finding patterns and insights from the data and to answer various what-if questions 
  • Intelligencefication: Using AI to recommend optimal decisions based on that knowledge 

We see this happening in pharmaceutical research and development. Big companies have digitized years of clinical trial protocols, patient records, and regulatory documents. Once these documents are digitized, statisticians can start analyzing them, linking trial criteria to outcomes and study designs to success rates. This allows them to ask important questions like, ”What makes some trials succeed while others fail?” or ”How do certain criteria affect patient enrollment and outcomes?” Recent studies using real-world patient data have shown that many traditional trial restrictions don’t significantly affect outcomes, and relaxing these restrictions could increase the number of eligible patients without harming results. This is an example of “knowledgefication”—turning large, unstructured data into insights that can improve trial design. 

The last step of “intelligencefication” is about to happen. This could mean AI helping in design trials. For example, after learning from many past trials, an AI agent might suggest the best inclusion and exclusion criteria for enrolling patients to best differentiate treatment efficacy and safety. It can also recommend the best schedule of activities to maximize trial success. We can imagine AI tools that combine information from regulatory documents, scientific publications, conference abstracts, and early experiments to predict what concerns regulators might have about a new drug. Statisticians, with their skills in data analysis and experiment design, will be crucial in checking and using these AI recommendations. By leading the digitalization and analysis of diverse data—and by carefully evaluating AI’s suggestions—statisticians help ensure the pharmaceutical industry’s new data leads to reliable knowledge and smart actions.  

In short, the growth of data in pharmaceutical research further increases the influence of statisticians in organizing data and making decisions, which reinforces their role as key players in AI-driven research. 

Go Beyond Traditional Statistical Methods 

Modern AI doesn’t just improve traditional statistics; it often surpasses them, opening new scientific areas. A great example is AlphaFold2 by DeepMind, which revolutionized how we predict protein structures. Before, predicting a protein’s 3D shape from an amino acid sequence required expert-crafted features and significant domain knowledge. AlphaFold2 changed this by using deep learning to directly predict structures from sequences, skipping the extensive human interventions. It achieved high accuracy—even without similar known structures—and matched experimental results for many targets. This breakthrough, published in Nature in 2021, showed that AI can learn complex biological patterns from data without needing detailed chemistry or physics rules.  

Crucially, this shift opens a new lane for quantitative scientists with strong mathematical and programming skills. Complex biological and chemical problems that used to be the domain of specialized computational biologists are increasingly being tackled with general-purpose, data-driven methods. Statisticians, given their training in rigorous modeling and algorithm development, are well positioned to contribute to this emerging domain, often referred to as “digital biology.”  

The nature of work in digital biology (such as de novo protein design or ligand generation) often involves advanced mathematics and computations that go beyond classical computational chemistry training. For instance, modern generative models for molecular structures exploit concepts from differential geometry and Lie groups to enforce physical symmetries (e.g., rotational or translational invariances of molecules) in the learning process. Geometric deep learning frameworks have been developed to handle data on non-Euclidean domains like protein surfaces and molecular graphs, encoding invariances under rotations/reflections by design.  

Implementing and extending these models requires fluency in linear algebra, group representations, and high-performance computing—skill sets much more aligned with statisticians or applied mathematicians than traditional wet-lab scientists or computational biologists/chemists. In effect, drug discovery is becoming more of an algorithmic science. This creates ripe opportunities for statisticians to lead methodological innovation in protein engineering, small-molecule drug design, and other areas where sophisticated modeling (rather than domain-specific intuition alone) drives breakthroughs. 

Beyond AlphaFold, numerous other examples illustrate how algorithmic approaches are reshaping pharmaceutical R&D. For statisticians, each of these advances signals a domain where their expertise can be applied in novel ways such as designing modeling strategy, ensuring rigorous validation, and quantifying uncertainty in predictions. Notably, many such AI-driven discovery techniques emphasize prediction and optimization (e.g., finding a molecule that maximizes a predicted efficacy score) rather than classical inferential statistics.  

This highlights an important cultural shift that renowned statistician Leo Breiman presciently discussed in his “two cultures” essay. Breiman argued that much of traditional academic statistics focused on data models and inference under an assumed “true” model, whereas a different culture—exemplified by machine learning—focused on algorithmic models aimed at predictive accuracy. He urged statisticians to embrace this algorithmic approach for complex problems, where the goal is often prediction or discovery, not estimating a pre-specified parameter.  

The current wave of AI in pharma is a testament to Breiman’s point: Many breakthroughs (like protein folding or de novo molecule generation) are essentially large-scale prediction problems, where flexible algorithms trump analytical formulas. Statisticians who adapt to this mindset—valuing predictive performance and computational experimentation alongside traditional inference—can substantially broaden their impact. 

The rise of AI methods is pushing the boundaries of what quantitative scientists do in pharmaceutical research. Statisticians equipped with strong coding abilities and mathematical depth are in an excellent position to drive these innovations. They can develop new algorithms, evaluate AI models, and ensure methods are applied soundly. By venturing beyond the confines of traditional statistical methodology, while still upholding standards of rigor and clarity, statisticians can become key players in cutting-edge domains like AI-driven drug discovery, precision medicine, and digital health. Their contributions will complement those of domain specialists, blending data-centric problem-solving with scientific insight to accelerate pharmaceutical progress. 

Broadening Responsibilities and Impact Across the Value Chain 

The role of statisticians in the pharmaceutical industry has evolved dramatically in recent years, from a narrow focus on clinical trials to a broad involvement across the entire drug discovery and development lifecycle. Historically, a pharmaceutical statistician’s influence was largely confined to phase II/III clinical development: designing trials, analyzing efficacy and safety data, and supporting regulatory submissions. Today, statisticians are increasingly embedded in cross-functional teams from early discovery and preclinical research to manufacturing, quality control, and post-marketing surveillance and health economics. This expansion is driven by the growing recognition that the statistician’s core skill set—quantitative reasoning, experimental design, data interpretation, and uncertainty quantification—is invaluable at every stage where data are generated and decisions are made.  

Gregory G. Enas and John S. Andersen presciently noted that statisticians are uniquely trained to improve decision-making “from the very early stages of drug discovery until patients, payers, and regulators are satisfied,” essentially advocating for statisticians to become key contributors in all phases of the enterprise. Two decades later, this vision is being realized. Statisticians now collaborate with chemists and biologists in discovery research; optimize processes with engineers in chemistry, manufacturing, and controls groups; and partner with physicians and epidemiologists to assess real-world outcomes post-approval. The modern pharmaceutical statistician often serves as a quantitative strategist, not only ensuring analyses are sound, but also guiding what data to collect, how to collect it efficiently, and how to interpret it to drive business and scientific decisions. 

Besides drug discoveries, the role of statisticians is rapidly evolving as AI becomes integral in fields beyond traditional statistics. In clinical development, statisticians leverage AI to enhance trial design and execution. Natural language processing algorithms analyze study protocols and electronic health records. Image technology and digital biomarkers help in enrollment—particularly in complex oncology trials by quickly finding eligible patients across extensive health networks. AI is also being used to simulate or augment control arms in trials through the creation of “digital twins”—virtual patient avatars generated from historical data. This innovative approach helped make trials more efficient and ethically palatable. Statisticians play a crucial role in validating these AI models to ensure they accurately represent patient outcomes and maintain scientific and regulatory rigor. 

In the manufacturing and supply chain sectors, AI is driving the transition to “Pharma 4.0,” a new paradigm of smart, data-driven production. Statisticians now collaborate with engineers to implement AI-based process monitoring and control systems. Machine learning models analyze process development data to identify optimal parameters and scaling conditions, accelerating development. AI-driven advanced process control systems can make real-time adjustments during production and ensure critical quality attributes remain within target ranges. The FDA has acknowledged the potential of AI in drug manufacturing, highlighting its ability to reduce development time and waste through improved process design. Statisticians are essential in deploying these advancements, from designing experiments to training AI models, validating their performance, and integrating statistical process control. 

In the realm of commercialization, AI and advanced analytics empower statisticians to drive better business decisions. AI algorithms inform demand forecasting and inventory optimization, analyzing historical sales and external data to accurately predict drug demand. This helps reduce stockouts and oversupply, optimizing the pharma supply chain. In marketing, AI tools help segment healthcare providers and patients, tailoring outreach to those most likely to benefit. Predictive analytics guide field sales strategies by integrating data on prescribing habits and patient demographics, enhancing targeting precision. Statisticians collaborate with AI to develop pricing strategies, using machine learning models to analyze market data and recommend optimal pricing. These advancements demonstrate that data-driven decision-making is becoming the norm in pharma, with statisticians translating AI-driven analytics into actionable business insights. 

Statisticians will also influence future AI data collection strategies. Traditionally, data collection in trials was solely focused on regulatory approval of the molecule. In the future, we envision companies will deliberately collect data not just to advance the current product, but also to improve the next generation of AI models that assist in drug design and development. This might entail, for instance, designing clinical studies that also create high-quality datasets for machine learning (such as rich biomarker panels or digital sensor data), recognizing that these datasets could inform many programs beyond the original trial. Statisticians will be key in planning such dual-purpose studies, balancing immediate needs with the long-term value. Techniques like adaptive sampling and active learning—where collected data is dynamically guided by algorithmic learning needs—could become part of trial design considerations. By advising on how to gather the most informative data for both human decision-making and machine learning, statisticians ensure that pharmaceutical data resources continuously feed the cycle of innovation. 

AI is expanding what scientists can do in pharmaceutical research. Solving today’s tough problems often requires complex models and heavy computation. Statisticians with strong coding and math skills are well-positioned to lead these innovations. They create new algorithms, evaluate AI models, and ensure these methods are used correctly. By moving beyond traditional statistics, statisticians can play key roles in AI-driven drug discovery, precision medicine, digital health, manufacturing, and commercialization. Their work will complement that of domain experts, combining data-driven problem-solving with scientific insight to advance pharmaceutical research. 

Conclusion 

The role of statisticians in the pharmaceutical industry has undergone a remarkable evolution, expanding in scope, influence, and importance over the past 50 years. From the early days of following the 1962 FDA reforms—when a few statisticians helped ensure new drugs had statistically sound evidence of efficacy—to the present day where statisticians are at the forefront of AI-driven drug development. The transformation is profound. We have seen how historical milestones, such as regulatory changes (’80s NDA guidelines, ICH E9) and technological advances (computing revolution, big data) set the stage for statisticians to move from the periphery to the core of decision-making in pharma. 

Driving this evolution are key factors like statistical computing, the proliferation of new data types demanding novel analytical methods, and the willingness of industry and regulators to embrace innovative statistical designs that can make drug development more efficient. The recent surge of AI has further catalyzed a paradigm shift, positioning statisticians as vital contributors to data science initiatives that span discovery through post-market use. This breadth of impact across the value chain—from molecule to market—exemplifies how far the influence of statisticians has grown beyond traditional boundaries. 

Crucially, statisticians have also grown in their leadership and collaborative roles. They are increasingly recognized as strategic partners who bring a data-driven lens to interdisciplinary teams. Whether it’s guiding a cross-functional team through designing an adaptive platform trial, negotiating the use of a novel surrogate endpoint with regulators, or explaining to commercial colleagues how an observational study supports a product’s value proposition, statisticians influence critical decisions at every step. Statisticians also serve as the bridge between their company and regulators on complex methodological issues. This role helps smoothen the adoption of things like complex innovative trial designs and real-world evidence considerations. 

The future trajectory points to statisticians continuing to be agents of innovation in pharmaceutical R&D. With the ongoing integration of AI, the growth of personalized medicine, and increasing reliance on real-world data, there will be even greater demand for statisticians who can blend quantitative rigor with creativity and strategic thinking. We anticipate statisticians playing leading roles in quantitative decision-making at every step of pharmaceutical research. This will require concerted effort in training and professional development. Academia and industry must work together to equip statisticians with a modern skill set that includes advanced modeling, machine learning, statistical computing (high-performance parallel computing), algorithms, mathematical optimization, and software engineering basics. The curriculum adjustments and competency development recommended in this paper are intended to future-proof the profession. 

The evolving role of statisticians in pharma is a success story. One of how a profession can adapt and expand to meet new challenges. Statisticians have leveraged advanced analytics and AI to augment and elevate their traditional work—not replace it. This drives better decisions and outcomes. They’ve become frontline leaders, ensuring evidence and data quality remain the bedrock of pharmaceutical innovation. The fruits of this evolution are evident: more efficient trials, more robust evidence of drug benefits and risks, and ultimately, a more informed approach to bringing therapies to the patients. Our evolving role will be characterized by leadership, innovation, and an unwavering commitment to using data for the betterment of public health.

Filed Under: Additional Features Tagged With: AI, Amy Xia, Biopharmaceutical Report, data sets, emerging technology, Haoda Fu, history, innovation, pharma, pharmaceutical, statistical analytics

Reader Interactions

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Primary Sidebar

Search

More to See

JSM 2027 Chicago, Illinois Logo

Shape the Future: Invited Session Proposals Sought for JSM 2027

August 3, 2026 By Meg Ruyle

Laptop open next to the words "online community"

How I Built a Portuguese-Language Analytics Community Online

August 3, 2026 By Meg Ruyle

New Member Spotlight: Collin Nill

August 3, 2026 By Meg Ruyle

colorful books

Members Offer Advice for Writing as a Team

August 3, 2026 By Meg Ruyle

ASA HOME

American Statistical Association

Communications from the Executive Director

ASA Leader Hub

ASA Career Connect

ADVERTISERS

STATA
SIAM

Archives

Categories

Footer

Editorial Staff

Managing Editor
Megan Murphy

Graphic Designers / Production Coordinators
Olivia Brown
Meg Ruyle

Communications Strategist
Val Nirala

Advertising Manager
Christina Bonner

Contributing Staff Members

Kim Gilliam

American Statistical Association
277 South Washington Street, Suite 370
Alexandria, VA 22314-3646
Phone: (703) 302-1857

 

Copyright © 2026 · Magazine Pro on Genesis Framework · WordPress · Log in