• Skip to main content
  • Skip to secondary menu
  • Skip to primary sidebar
  • Skip to footer
  • Homepage
  • About Us
  • Advertising
  • Submission Instructions
  • Editorial Calendar
Amstat News

Amstat News

The Membership Magazine of the American Statistical Association

  • Printed Issues
  • Practical Significance Podcast
  • Additional Features
  • Columns
  • Member News
  • Departments
You are here: Home / Departments / A Statistician's View / As AI Does More Analytics, What Should Statisticians Do More Of?  

As AI Does More Analytics, What Should Statisticians Do More Of?  

October 1, 2026 Leave a Comment

Zhiwei Zhu, Purdue University 

A few years ago, teaching statistical analysis meant spending considerable classroom time on execution: writing code, visualizing data, fitting alternative models, diagnosing model fit, and comparing accuracy measures. Today, a student can ask an AI system to perform much of that work in minutes. The same transformation is occurring in practice. An analyst equipped with AI can generate code, fit competing models, produce visualizations, and summarize results at a speed that would have been difficult to imagine not long ago.  

For statisticians, this raises an uncomfortable but useful question: As AI does more analytics, what should statisticians do more of?  

My own answer has evolved while working on two seemingly different projects, one examining the changing relationship among statistics, data science, and AI—as presented in my Harvard Data Science Review paper, “Data Science at a Fork in the Age of AI: Recognizing Divergent Missions in Research and Education”—and another reconsidering how time-series forecasting (and more broadly analytics) should be taught and practiced—as developed in my book Forecast by Design. They have led me toward the same conclusion: AI is expanding the territory of analytics while simultaneously reducing the cost of many analytical tasks. Both changes increase, rather than diminish, the importance of statistical judgment. 

The Map Is Getting Bigger—Much Bigger  

Much of modern statistics developed within a particular data world: structured observations and variables, limited samples and unknown populations, and relatively parsimonious models and parameters. Even when statisticians encountered text or images, we typically made them analytically manageable by converting them into features and structured representations.  

AI is breaking through those constraints.  

Large language and multimodal models increasingly work directly with language, imagery, audio, video, code, and other forms of human and machine expressions—what I call expression-based data. Developments in spatial intelligence and world models point farther outward, toward data grounded in perception, spatial relationships, movement, environments, and their evolution over time—what I refer to here as world-grounded data. What counts as computationally usable data is becoming substantially broader than the relational records around which much conventional analytics developed.  

A geographic analogy can be useful for thinking about this change. Structured data are like land. Over many decades, we built roads, maps, and institutions for navigating that land: relational databases, SQL, sampling theory, statistical models, metadata standards, and validation procedures. 

Expression-based data—the raw material increasingly powering generative AI—resembles an ocean: vast, multidimensional, fluid, and still unevenly charted. Large language models and other generative systems may be thought of as some of our first powerful watercraft. World-grounded data may push us farther still toward systems capable not only of representing information but of perceiving and modeling how environments evolve.  

The analogy matters because discovering the ocean does not make land obsolete. Most organizations will continue to make consequential decisions using sales records, clinical measurements, financial transactions, experiments, surveys, and operational time series. The statistical infrastructure built for such data remains indispensable.  

But discovering an ocean should make us reluctant to define the entire world as land.  

This expansion involves more than data type. I see at least two other departures worth watching. Classical statistics developed largely around sample-based inference: using limited observations to reason carefully about a larger, unknown population. Contemporary AI models increasingly begin instead with enormous training corpora that approximate populations at unprecedented scale. And where statistics has traditionally valued parsimony, interpretable structure, and explicit uncertainty, contemporary AI often derives capability from expansive representations containing billions of parameters.  

These are not arguments against statistics. They are reasons to ask which statistical principles travel well into this new territory, which require extension, and where entirely new analytical infrastructure may be needed.  

Meanwhile, something interesting is happening back on land. There is another side to the AI story that may be more immediately relevant to many statisticians. Even when we remain on familiar structured-data territory, AI is changing the economics of analytical work.  

Consider forecasting. The traditional workflow devotes substantial attention to selecting and fitting models. Should we use exponential smoothing or ARIMA? How should seasonality be represented? Which specification minimizes RMSE or MAE? These remain legitimate statistical questions. Statisticians need to understand the methods and help stakeholders understand their implications.  

But as AI makes many technical tasks easier and faster, organizations have greater reason to focus on what happens before and after forecasting: how method selection aligns with purpose, whether historical accuracy translates into future reliability, and when implemented forecast should no longer be trusted. 

For example, for a hospital anticipating capacity needs, underforecasting a surge may be much more costly than overforecasting it. In inventory management, the relative costs of stockout and excess inventory matter. A senior executive seeking directional understanding may not require the same level of precision as an automated replenishment system requires. These differences cannot be adequately captured by conventional forecast accuracy measures such as RMSE or MAE alone. 

That is why I argue that forecasting should be treated as a design problem rather than simply an analysis task. The central question shifts from “Which model forecasts best?” to “How should forecasting be designed to support a specific decision?” 

AI makes that distinction more consequential. When fitting another model becomes inexpensive, generating more models is not necessarily where another hour of a statistician’s time creates the greatest value.  

Accuracy Is Not Reliability  

One distinction from forecasting has become particularly important to me: accuracy and reliability are not the same question.  

We usually validate a forecasting model by looking backward. We hold out historical observations, generate forecasts, calculate errors, and compare models. That tells us something important: how the model would have performed on data we already possess.  

But a forecast is consumed in the future.  

After implementation, the environment evolves. Relationships drift. A pandemic may occur. Consumer behavior changes. A competitor enters. Forecast residuals (or errors) that were once random may begin moving systematically in one direction.  

This suggests another important role for statistics: forward validation. 

Instead of treating residuals merely as historical errors to minimize during model development, we can treat them as signals after deployment. Is their mean moving away from zero? Is their variance increasing? Are underforecasts becoming systematically more common? At what point should we maintain, refit, or rethink the forecasting system? Investigations to such questions have long been desirable and attended, but limited and difficult to act on.  

Forecast by Design places forward validation at the center of forecast governance: models do not simply get validated and deployed; their post-deployment reliability must be watched as time moves forward.  

This principle extends well beyond forecasting. A generative or agentic AI system that performs well on a benchmark today is not necessarily trustworthy tomorrow, in another context, with another user, or after repeated interaction. Statisticians know a great deal about validation. AI gives us both the need and the capacity to extend validation across context and over time.  

From Building Models to Designing Analytical Systems  

This leads me to a somewhat different answer to the question posed in the title. The future value of statisticians may lie less in competing with AI over who can produce a regression, ARIMA model, visualization, or block of Python code faster. It may lie in extending statistical reasoning across the entire life cycle of a decision system.  

Before analysis, that means asking what the observations actually represent, what is absent, whether the available data corresponds to the population or decision of interest, and whether the problem has been framed correctly.  

During analyzing data, it means understanding structure rather than merely searching across algorithms. In time series, for example, trend and seasonality may be visible components to model explicitly, while autocorrelation represents a less visible form of temporal memory. Machine learning offers still another lens. AI makes it easier to try all of them, but easier model production does not answer the harder question: Which representation is appropriate for the purpose at hand?  

After analysis, statistical responsibility extends to forward validation, monitoring, uncertainty communication, thresholds for action, and rules for human intervention.  

The endpoint is no longer simply an analytical output. It is a decision system.  

This is where I see an interesting convergence between an old statistical tradition and a new AI opportunity. Statistics has always taught us not to confuse computation with evidence. AI makes computation abundant. That may also make disciplined reasoning about evidence both more valuable and more distinctly a human responsibility. 

What Does This Mean for Statistics Education? 

The implications may be greatest in the classroom. It is tempting to conclude that because AI can write code and fit standard models, students need less technical foundation. I reach almost the opposite conclusion. 

Students still need mathematics, statistics, computing, and problem-solving foundations—not primarily so they can outperform AI at routine execution, but so they can recognize when an AI-generated analysis is inappropriate, unnecessarily complex, poorly validated, or disconnected from the decision.  

At the same time, simply adding an “AI module” to an existing statistics or data science course seems insufficient. If AI changes both what counts as data and what machines can perform analytically, we need to reconsider the division of labor between machine execution and human judgment.  

In my own analytics courses, I increasingly require students to use AI as both a learning partner and a thinking partner. As a learning partner, AI can help students explain a concept, generate code, compare methods, and extend their learning beyond the classroom. As a thinking partner, however, students must take responsibility for questioning assumptions, challenging reasoning, identifying weaknesses, and making judgements about what deserves trust.  

This partnership takes effort to build, and it allows classroom attention to shift toward questions organizations increasingly care about: Why are we analyzing? What data structure matters? What evidence would make us trust the results? What happens when analytical performance changes after deployment? What decision will actually be different because this analysis exists? 

The educational progression I now find useful is therefore not simply learn a method → apply a method but rather understand the data → develop analysis → validate the models → monitor reliability → design the decision.  

AI can participate in every stage, but responsibility for connecting those stages remains human. 

Statistics Does Not Need to Choose Between Land and Ocean  

The rise of AI is sometimes framed as a competition among statistics, data science, and computer science. I am increasingly unconvinced that competition is the most useful framing.  

There are really two opportunities before us. One points toward the ocean. As AI moves toward expression-based and world-grounded data, statistical and data-science thinking can contribute to unresolved questions of provenance, representation, validation, transparency, sustainability, and governance. The analytical infrastructure for this new data territory remains immature. 

The other opportunity remains firmly on land. Structured data and conventional analytics are not disappearing. But AI gives statisticians an opportunity to spend relatively less effort on routine analytical execution and more on the difficult work surrounding it: understanding what data means, determining what deserves trust, monitoring what happens after deployment, and connecting uncertain evidence to consequential decisions. 

“Data Science at a Fork in the Age of AI: Recognizing Divergent Missions in Research and Education” identifies the first development as a possible fork in the mission of data science and its relationship to statistics. Forecast by Design approaches the second from a more practical direction, using forecasting to illustrate how analytics can be taught in MBA and graduate analytics programs, and practiced in organizations, treating AI as both a learning and thinking partner.  

The two paths lead me to the same conclusion: As AI does more analytics, it opens new frontiers for statisticians, as well.  

We might need to do more around the analytics—more questioning before a model is built, more judgment about whether its output deserves trust, more validation after it enters the world, and more design connecting evidence to action. Perhaps that is one way statistics can enter the age of AI without losing what made it valuable in the first place. 

Filed Under: A Statistician's View, Departments Tagged With: A statistician's view, AI, analytics, decision-making, design problems, forecasting, forward validation, statistical judgment

Reader Interactions

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Primary Sidebar

Search

More to See

2024–2025 Degree Data: Awarding of Statistics, Biostatistics Degrees Steady Amid Staggering Growth for Data Science, Analytics Degrees 

October 1, 2026 By Olivia Brown

Member Showcase: Robert Oster on the Power of Collaboration, Leadership, and Giving Back

October 1, 2026 By Olivia Brown

Who’s Missing from the Data? The Case for Disability Inclusion 

October 1, 2026 By Olivia Brown

Meet New Member Naren Prakash: Aspiring Statistician with a Love for Statistical Research 

October 1, 2026 By Olivia Brown

Data Science Certification

ASA HOME

American Statistical Association

Communications from the Executive Director

ASA Leader Hub

ASA Career Connect

STAFF LIST

Kim Gilliam
Naomi Friedman
Amanda Malloy

Archives

Categories

Footer

Editorial Staff

Managing Editor
Megan Murphy

Graphic Designers / Production Coordinators
Olivia Brown
Meg Ruyle

Communications Strategist
Val Nirala

Advertising Manager
Christina Bonner

Contributing Staff Members
Kim Gilliam

American Statistical Association
277 South Washington Street, Suite 370
Alexandria, VA 22314-3646
Phone: (703) 302-1857

 

Copyright © 2026 · Magazine Pro on Genesis Framework · WordPress · Log in