• Skip to main content
  • Skip to secondary menu
  • Skip to primary sidebar
  • Skip to footer
  • Homepage
  • About Us
  • Advertising
  • Submission Instructions
  • Editorial Calendar
Amstat News

Amstat News

The Membership Magazine of the American Statistical Association

  • Printed Issues
  • Practical Significance Podcast
  • Additional Features
  • Columns
  • Member News
  • Departments
You are here: Home / Additional Features / DATA: The Engine That Drives DataFest

DATA: The Engine That Drives DataFest

September 2, 2025 Leave a Comment

Robert Gould and Mine Çetinkaya-Rundel

You can’t have a DataFest without data. But how does the ASA DataFest find and assemble the data sets thousands of students scrutinize to squeeze out every drop of information?

Rob Gould and Mine Çetinkaya-Rundel have been selecting the data for DataFest from the beginning. Gould founded DataFest at the University of California at Los Angeles in 2011, when Çetinkaya-Rundel was in her final year as a PhD student. They have been planning DataFests ever since. 

One driving force behind founding DataFest was to provide undergraduate students with a real-life, data-driven problem. Data sets with large numbers of rows and columns, while common in the workplace or well-funded research labs, are rarely found in the classroom. As a result, students are likely to work with well-explored data that holds little mystery or room for discovery. 

In 2011, Gould and Çetinkaya-Rundel’s goal was to provide students with a challenge that had never been tackled and a client who was truly interested in what the students had to say. “We thought organizations looking for assistance with their data would be eager to have a group of talented undergraduates look hard at their problems,” wrote Gould and Çetinkaya-Rundel. “We imagined these ‘data donors’ would be mostly nonprofits and civic/government organizations, since these might not have statisticians or data scientists on staff or might have interesting problems they don’t have the time or resources to tackle.” Instead, Gould and Çetinkaya-Rundel found many organizations are eager to take part.

The first data set was provided by the Los Angeles Police Department. At the time, data-driven policing was relatively new, and the LAPD was eager to use data to reduce crime and increase public accountability. The LAPD had been working with UCLA researchers for a few years after their new publicly available crime map erroneously suggested the most crime-ridden location in Los Angeles was the headquarters of the Los Angeles Times. The LAPD was eager to recruit the best young data scientists into the new realm of predictive policing, so they were a natural donor candidate for the first DataFest.

The data required a lot of preparation before it was released to the students, and the entire UCLA Department of Statistics pitched in. For example, the location data used a unique origin point and proprietary mapping algorithm and had to be back-engineered to be mapped using standard latitude and longitude. In addition, Gould and Çetinkaya-Rundel decided to supplement the data set with the locations of transitional living facilities.

“Little has changed since then,” wrote Gould and Çetinkaya-Rundel. “The data often requires heavy preparation, and we often need to provide additional context to help the students have a productive weekend.” One thing that has changed, however, is businesses are more aware of the value of data, so it has become a challenge to get data from prominent corporations.

In the early years, Gould and Çetinkaya-Rundel were able to entice major tech companies into donating data: Edmunds.com; Ticketmaster; e-Harmony; and Expedia. But as companies grew wary of supplying data, students expressed interest in seeing data from beyond the corporate world. In response, data has come from the UCLA Department of Psychology (CourseKata Project), Yale School of Medicine (Play2Prevent Lab), Canadian National Women’s Rugby Team, American Bar Association, and Rocky Mountain Poison Control Center.

Gould and Çetinkaya-Rundel find the data sets mostly through word of mouth, though some come from cold calls. As the years go by, more DataFest alumni are in positions to donate data, so Gould and Çetinkaya-Rundel ask former participants, presenters at the Joint Statistical Meetings, presenters at their local departmental seminars, and other faculty. “Ideally, we have about three promising leads by September,” wrote Gould and Çetinkaya-Rundel. “We engage in conversations with these leads about what we’re looking for and whether we will get their institution’s permission to use the data. We ask for sample data and use this to try to flesh out a challenge and assess whether the data will meet our criteria.”

Usually, one or two leads drop out as they realize they can’t get institutional buy-in or don’t have the time to assemble the data. Sometimes, they can donate data the following year, which helps.

Once Gould and Çetinkaya-Rundel have a data donor, they get the full data set and begin cleaning and preparing a codebook. The ASA supports a graduate student to help them with this, and this person usually does the bulk of the testing. Testing includes the usual tasks of data cleaning but also includes simulating how students might approach the data and looking for potential roadblocks they might meet.

Gould and Çetinkaya-Rundel make decisions about which aspects they should clean based on their assessments of how much time it might take the students to fix the data. This is an iterative process, with many queries back to the data donor, and leads to several versions of increasingly refined data.

Preparing a codebook is also a major endeavor, as organizations often have little documentation that can be shared publicly. This might be because they rely on institutional knowledge and practices, the documentation has proprietary information, or the assembled data is entirely new to the organization. 

During this process, Gould and Çetinkaya-Rundel write and rewrite the challenge, which is the second-most important aspect of  the competition. “We strive to provide a challenge that gives all students a path forward, regardless of their level of data sophistication,” they wrote. “A good challenge provides sufficient leeway for students to exercise creativity and ‘choose their own adventure.’ At the same time, the challenge needs to be precise enough that students are clear about what the data donor is looking for.”

The most important component of DataFest is the data, and the primary requirement for the data is that it be rich. For Gould and Çetinkaya-Rundel, “richness” is measured mostly by the number of distinct variables and the possibility of enhancement through external data or feature engineering with the available data. The context is key; it must be accessible to students so they do not waste valuable time learning a new subject but instead have sufficient understanding to think of interesting questions. There is also a novelty aspect; the data should not be publicly available, although parts of it can be. The data donor must also be willing to record a video to put a face with the data set so it is clear real people are listening and curious about what the students might find.

A constraint is that the data be large, but not too large. Gould and Çetinkaya-Rundel limit the data set and documentation to under 3 GB (2GB is ideal), which is about the limit of what can be distributed simultaneously to many students. 

Additionally, Gould and Çetinkaya-Rundel cannot ask students to sign nondisclosure agreements, and this makes it difficult for data donors from private institutions. Students are asked to read a statement that informs them that by taking part in DataFest, they cannot use the data for any purpose other than DataFest without permission from the data donor. 

Finding data for DataFest has been one of the more rewarding aspects of Gould and Çetinkaya-Rundel’s careers, although there have been years of great stress. They hope to enlarge their data-finding committee. If anyone is interested, send an email to Gould. And if you know of anyone with data to share, pass on Gould’s email address.

Photo of male, big smile, gray hair

Rob Gould

Gould is a teaching professor in the University of California, Los Angeles, Department of Statistics and Data Science. The founder of ASA DataFest in 2011, he is a Fellow of the American Statistical Association and recipient of the ASA Founder’s Award. He is also a co-author of the introductory statistics textbook Exploring the World Through Data.

    Female, blonde shoulder length hair, big smile

    Mine Çetinkaya-Rundel

    Çetinkaya-Rundel is professor of the practice and the director of undergraduate studies in the department of statistical science and the director of first-year experience at Duke University. She is a Fellow of the ASA, recipient of the Waller Education Award, and co-author of R for Data Science and OpenIntro.

      Filed Under: Additional Features, DATAFEST, Special Features Tagged With: cleaning data, CourseKata Project, data sets, datasets, LAPD, Play2Prevent lab, sourcing data, students, UCLA

      Reader Interactions

      Leave a Reply Cancel reply

      Your email address will not be published. Required fields are marked *

      Primary Sidebar

      Search

      More to See

      Jennifer L. Green: The Collaborative Life of a Statistics Professor and Teacher Mentor

      September 1, 2026 By Megan Murphy

      Meetings, Manuscripts, and Meows: Spend a Day with Charlotte Walsh

      September 1, 2026 By Megan Murphy

      Students’ Statistical Thinking When Using Generative AI

      September 1, 2026 By Megan Murphy

      What Students Taught Me About Teaching Statistics

      September 1, 2026 By Megan Murphy

      STATAcorp. Efficiency matters. Stata is easy to use, so you spend less time learning software and more time focusing on your research
      Data Science Certification

      ASA HOME

      American Statistical Association

      Communications from the Executive Director

      ASA Leader Hub

      ASA Career Connect

      ADVERTISERS

      STATA
      SIAM

      Archives

      Categories

      Footer

      Editorial Staff

      Managing Editor
      Megan Murphy

      Graphic Designers / Production Coordinators
      Olivia Brown
      Meg Ruyle

      Communications Strategist
      Val Nirala

      Advertising Manager
      Christina Bonner

      Contributing Staff Members
      Kim Gilliam

      American Statistical Association
      277 South Washington Street, Suite 370
      Alexandria, VA 22314-3646
      Phone: (703) 302-1857

       

      Copyright © 2026 · Magazine Pro on Genesis Framework · WordPress · Log in