• Skip to main content
  • Skip to secondary menu
  • Skip to primary sidebar
  • Skip to footer
  • Homepage
  • About Us
  • Advertising
  • Submission Instructions
  • Editorial Calendar
Amstat News

Amstat News

The Membership Magazine of the American Statistical Association

  • Printed Issues
  • Practical Significance Podcast
  • Additional Features
  • Columns
  • Member News
  • Departments
You are here: Home / Featured Stories / Unlocking Statistical Insights: NSF’s Push for Mathematical Foundations of Digital Twins

Unlocking Statistical Insights: NSF’s Push for Mathematical Foundations of Digital Twins

June 3, 2024 2 Comments

Short dark hair, slight smile
To strengthen the connection between the statistical community and National Science Foundation, we continue the series introduced in the May 2023 issue that poses questions to NSF program officers and awardees. If you have questions or comments for the program officers, send them to ASA Director of Science Policy Steve Pierson at pierson@amstat.org.

 
This month’s program officer is Yulia Gel from the Division of Mathematical Sciences in the NSF Directorate for Mathematical and Physical Sciences. The awardee responses are from Qing Mai of Florida State University.

Program Director

Yulia Gel is on leave from The University of Texas at Dallas. This is her third year as a rotator program director of the statistics program.

Digital twins research and NSF funding on digital twins—what is there for statisticians?

A digital twin is an emerging technology that allows for a dynamic virtual representation of various real-world physical systems, objects, and processes. One of the essential constituents of digital twins is the bidirectional connection with their physical counterparts. That is, digital twins can learn to synchronize and communicate with their physical counterparts by continuously exchanging information, using, for example, tools of artificial intelligence, or AI.

In particular, a recent report from the National Academies of Sciences, Engineering, and Medicine titled “Foundational Research Gaps and Future Directions for Digital Twins” defines a digital twin as “a set of virtual information constructs that mimics the structure, context, and behavior of a natural, engineered, or social system (or system-of-systems), is dynamically updated with data from its physical twin, has a predictive capability, and informs decisions that realize value.”

To enhance adoption of the digital twins in real-world applications—from health care to civil engineering—we first need to ensure digital twins do indeed reliably represent their physical counterparts in a broad range of what-if scenarios.

For example, how can we construct the confidence intervals of digital twin outputs while accounting for uncertainties such as modeling, data, and process? How can we relate those confidence intervals to decision-making? Is it possible to develop verification, validation, and uncertainty quantification techniques adapting to the physical twin evolving over time? How can we detect and mitigate biases that may potentially be introduced in the digital twin models and associated AI-based tools? How robust are the digital twin models to missing data, outliers, irregularly spaced records, and fusion of observations at various scales?

Many, if not all, these questions cannot be addressed without proper statistical inference and sound statistical methodology, with the machinery ranging from causal inference to robust statistics to experimental design to extreme value analysis.

Recognizing these needs, NSF has released several new solicitations on mathematical and statistical foundations of digital twins. The Foundations for Digital Twins as Catalyzers of Biomedical Technological Innovation NSF 24-561 is a joint tri-agency initiative among the NSF, US Food and Drug Administration, and National Institutes of Health. It focuses on mathematical and statistical principles behind digital twins and synthetic data used in the evaluation of medical devices and the relevance of the developed models in addressing current and emerging challenges affecting the development and assessment of biomedical technologies. The deadline for this solicitation is June 21.

In turn, the Mathematical Foundations of Digital Twins NSF 24-559 is a joint program with the Air Force Office of Scientific Research and supports foundational mathematical and statistical research on digital twins in applied sciences, without focusing on a particular application. The deadline for this solicitation is June 20.

It is hard to overstate the role of statistical sciences in digital twins research, and it is an exciting opportunity for statisticians to lead interdisciplinary projects.

Awardee

Short dark hair, slight smileQing Mai is a professor in the department of statistics at Florida State University. She earned her PhD from the University of Minnesota in 2013, and her research interests include high-dimensional data analysis, tensor data analysis, and machine learning. She has been the principal investigator or co-principal investigator for two NSF grants from the Division of Computing and Communication Foundations in the Directorate for Computer and Information Science and Engineering. She has also served on multiple NSF panels.

Mai and her Florida State colleague Xin Zhang received $479,000 from the Communications and Information Foundations program for their proposal, “Cluster Analysis for Highly Correlated, Heavy-Tailed, and Higher-Order Data.”

Tell us more about your project, including its motivation and goal.

Rapid advances in modern science and technology are resulting in the generation of data sets of unprecedented size and complexity. A common source of complexity in data sets is the presence of subpopulations. For example, a disease may have several subtypes, and customers may be attracted to different features of the same product. Cluster analysis is a popular tool to identify subpopulations, which affords a refined investigation on each of them.

This project aims to generate innovative clustering methods to understand the heterogeneity common in modern data sets. We answer the challenges of heavy tails, high correlations, and the tensor structure.

Such data sets frequently arise in contemporary scientific studies, but severely comprise classical clustering methods.

We propose a family of probabilistic models to accommodate such data sets, under which we develop model-based clustering methods.

In addition to the allocation of subjects, the methods in this research further find the defining features of each subpopulation. We pay special attention to efficient computation and rigorous theoretical study.

The research team will apply these methods to various real-world problems with the potential to affect multiple fields that rely on such data sets. Open source and user-friendly software will also be provided. Moreover, this project will be integrated with educational and outreach activities, including new courses, interdisciplinary training, and mentoring of underrepresented student groups in the mathematical and statistical sciences.

Describe your approach to the Division of Computing and Communication Foundations.

A critical task of applying outside the Division of Mathematical Sciences is to establish relevance to that specific entity. Some of my approaches toward this goal are the following:

      1. For each potentially interesting solicitation, I read abstracts of recent awards to determine whether my research overlaps with the funded projects, and, if so, which of my projects are most suitable.
      2. The guidelines of each solicitation are different, and I carefully follow them.
      3. In writing my proposal, I emphasize the unique contributions of statisticians to the problem of interest, as well as pay special attention to the computational aspects and scientific applications.

What advice do you have for others applying for NSF funding?

In recent years, the connections are significantly strengthened between statistics and many other areas such as computer science and engineering. Consequently, statistical research is increasingly valued by researchers from such fields. I suggest statisticians actively explore funding opportunities outside the Division of Mathematical Sciences, which increases the impact and visibility of statistics, as well.

Filed Under: Featured Stories, NSF Corner Tagged With: Digital Twins NSF 24-559, Division of Mathematical Sciences, model-based clustering methods, NSF Directorate for Mathematical and Physical Sciences, Qing Mai, subpopulations, Yulia Gel

Reader Interactions

Comments

  1. Mukesh Ram says

    August 8, 2024 at 8:01 am

    Great insights! Your blog offers crucial lessons on managing failed IT projects. I’ve also explored related themes in an ultimate guide on IT staff augmentation on my site, which might provide additional context to project recovery and team management. Would love to get your thoughts!

    https://medium.com/@mukesh.ram/lessons-you-can-learn-from-a-failed-it-project-96b6370bbcb7?postPublishedType=initial

    Reply
  2. Stan Young says

    January 24, 2025 at 2:22 pm

    An analysis method, Local Control Analysis, was proposed in 2013/2014. In this method, the data set is first clustered, and then a simple analysis proceeds within each cluster (e.g., a difference between two treatments, or non-parametric correlation between an outcome and a predictor). The test statistics from each cluster are then examined via recursive partitioning to see if they change systematically across clusters.

    Improved clustering would strengthen Local Control Analysis.

    Obenchain RL, Young SS. (2013) Advancing statistical thinking in health care research. Journal of Statistical Theory and Practice 7, 456-469.

    Lopiano KK, Obenchain RL, Young SS. (2014) Fair Treatment Comparisons in Observational Research. Statistical Analysis and Data Mining 7, 376–384.

    Reply

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Primary Sidebar

Search

More to See

JSM 2027 Chicago, Illinois Logo

Shape the Future: Invited Session Proposals Sought for JSM 2027

August 3, 2026 By Meg Ruyle

Laptop open next to the words "online community"

How I Built a Portuguese-Language Analytics Community Online

August 3, 2026 By Meg Ruyle

New Member Spotlight: Collin Nill

August 3, 2026 By Meg Ruyle

colorful books

Members Offer Advice for Writing as a Team

August 3, 2026 By Meg Ruyle

ASA HOME

American Statistical Association

Communications from the Executive Director

ASA Leader Hub

ASA Career Connect

ADVERTISERS

STATA
SIAM

Archives

Categories

Footer

Editorial Staff

Managing Editor
Megan Murphy

Graphic Designers / Production Coordinators
Olivia Brown
Meg Ruyle

Communications Strategist
Val Nirala

Advertising Manager
Christina Bonner

Contributing Staff Members

Kim Gilliam

American Statistical Association
277 South Washington Street, Suite 370
Alexandria, VA 22314-3646
Phone: (703) 302-1857

 

Copyright © 2026 · Magazine Pro on Genesis Framework · WordPress · Log in