One of my Data for Good heroes is David Riedman, founder of the K–12 School Shooting Database. Back in 2017, the Center for Homeland Defense and Security began the project to keep track of school shootings and other firearm incidents at schools in the United States. When funding ran out in 2022, Riedman kept this important database going on his own. Today, the project continues its mission as a 501(c)(3) not-for-profit organization.
Unfortunately, not all data-at-risk stories have happy endings like this one. Every year, important data resources needed by Data for Good researchers are in danger of going dark or, worse still, losing data quality relative to earlier data. Shifting political agendas can lead to funding changes, resulting in the reduction or even elimination of staff supporting critical data. Currently, with so many significant policy changes underway, I have identified the resiliency and hardening of important data sets as a top Data for Good priority in 2025.
Strengthening the data infrastructure supporting D4G initiatives takes many forms. Data from past successful projects needs to be preserved and made available for future use. Where data has been collected using paper forms and documents, digital copies are needed.
Riedman’s work with the K–12 School Shooting Database provides an example of one type of data source perennially at risk: projects funded by federal, state, and local governments that cease once their goal has been met. In these cases, published papers remain but the data often goes dark, lost to future researchers.
COVID pandemic–era data is particularly at risk, even though preserving it is needed to address future pandemics. The outstanding effort by biostatisticians that produced an explosion of research during the pandemic must be followed up by preserving the data—and metadata!—for future use.
Data produced and curated by government agencies is also at risk due to changes in funding priorities and staff changes that can result. This has attracted media attention. For example, Darya Minovi of the Union of Concerned Scientists has written about multiple efforts to preserve data as a new administration takes over. Common concerns in media reports include data on the environment, climate change, and justice.
Hardening data infrastructure and building resiliency begins with identifying the data sets most needed, either in our own area or widespread D4G applications. For example, I have been maintaining a zipped archive of key demographic tables from the American Community Survey since the late twenty-teens. I have never needed to access the archive but it’s there to support longitudinal demographic analysis if it is ever needed.
Coordinating data resiliency efforts with colleagues in our own research areas is important to avoid duplication of data sources while others go unarchived. Maintaining a copy in a space with backup support is essential. While small files can be archived in cloud spaces, this can become expensive for larger files. In this case, data tables can be divided across multiple colleagues or organizations.
Building data resiliency through identifying important data sets and archiving copies of the critical ones is an activity well-suited to the D4G community, as it mitigates the risk of potential deletion, damage, or discontinuation due to lack of support. It also provides an opportunity to connect with new collaborators and expand research teams with common interests in data for good!
Getting Involved
In opportunities this month, February is a great time to start planning activities for Earth Day. Start by connecting with local organizations working in your area of environmental advocacy to host a hackathon. Get a sponsor to provide a site, then work with your partners to develop a data set to be shared with the hackathon teams, along with a specific task for them to accomplish. Finally, publicize your event—and be sure to include it on the ASA Community!


Leave a Reply