This week, I am in Salt Lake City for the ASA’s annual Symposium on Data Science and Statistics. It’s such an amazing group of speakers and experts, gathered to share emerging technology driven by statistical science to make a beneficial impact on life in so many ways.
Biostatistics is always an important research area for ASA members, making SDSS a leading event for the latest developments in intelligent systems for health, medical, pharma, and related applications. So much is happening right now in this area with so much innovation and potential for good that I identified the development of intelligent systems as the top biostatistics opportunity for 2025.
One hot topic at SDSS this year was the new wave of AI. Open-source LLMs have arrived—open-source agents are next! This architecture welds powerful AI tools into a single operation. It starts with a query engine. The query accesses one or more databases, which can be huge. Ordinary LLMs receive a prompt directly, but these RAG (retrieval-augmented generation) tools send the prompt to a search engine and forward the search output to the LLM for processing. Agents are LLMs—often but not always prefaced by a RAG—that add a natural language processing tool to optimize the prompts and often another at the end of the process to make the LLM output easier to use.
With strong predictive power and the ability to leverage large databases, these agentic RAG tools offer tremendous opportunity to develop intelligent systems for biostatistics. Both new and existing biostatistics projects can add RAGs to better leverage large medical survey and public health databases. A recently published open-access article in Nature, titled “Retrieval-Augmented Generation for Generative Artificial Intelligence in Health Care,” analyzes the potential for using these tools to improve equity in health care research, mitigate bias, and address disparities in access to medical care.
Development of intelligent systems in biostatistics, including RAGs, often leverage large publicly available medical databases. The Global Burden of Diseases, Injuries, and Risk Factors Study project supports comparison of risk factors across countries around the world. A medical data science team led by Mohamad-Hani Temsah leveraged GBD data using ChatGPT-4 to develop health care plans optimized for individual patients.
Food and nutrition data from the US Department of Agriculture can be accessed using a RAG to evaluate diets of anonymized individuals in a survey to identify systemic nutritional weaknesses in a population for remedial action.
The World Health Organization continues to report COVID-19 cases and mortality by country by week, going back to the beginning of the pandemic.
These are just a few examples of the many data resources now available. The WHO data collection homepage is a wonderful place to look for ideas and resources for your next D4G project.
Data preservation efforts from statistical scientists in every area are having an enormous impact on supporting science, including Data for Good. A team led by Lucky Tran at Columbia University downloaded every publicly available resource from the US Centers for Disease Control and Prevention and copied them to the Internet Archive, a non-profit free public digital library that has become the new home for a lot of federal data at risk of going dark. Jonathan Gilmour of the Harvard School of Public Health has downloaded information about National Science Foundation awards going back to 1960 and added them to the Harvard Dataverse research archive. It’s a tremendous wealth of information on high-impact projects and researchers.
Also, looking at NSF grants over time (you all know I’m a time series statistician, right?) maps the trajectory of scientific innovation over decades. This allows us to discern future directions, overlooked areas of research, and a step-by-step “how-to” for advancing science for the public good.
Emerging methods, resources, and architectures for intelligent systems are being applied to an ever-increasing number of use cases in biostatistics—including age-old concerns in health and well-being—in a multitude of ways to create novel solutions for the benefit of all through Data for Good.
Getting Involved
In opportunities this month, SDSS will be in Milwaukee, Wisconsin, next year, so start making your plans now. You can visit the 2025 symposium website and check out the program for ideas, resources, and connections that support your next Data for Good project.
Also, the ASA is looking for volunteers to serve on committees that further the mission of the association in multiple ways. There are nearly 80 committees working in every area of interest. With so many groups, volunteers are needed regularly. Don’t be shy if you are early in your career! Students and young professionals bring important perspectives that are very much needed. You can nominate yourself or recommend someone else by going to the nomination website.


Leave a Reply