• Skip to main content
  • Skip to secondary menu
  • Skip to primary sidebar
  • Skip to footer
  • Homepage
  • About Us
  • Advertising
  • Submission Instructions
  • Editorial Calendar
Amstat News

Amstat News

The Membership Magazine of the American Statistical Association

  • Printed Issues
  • Practical Significance Podcast
  • Additional Features
  • Columns
  • Member News
  • Departments
You are here: Home / Additional Features / Up the Creek with Many Possible Paddles: How Students Cope with Being Awash in Data at DataFest 

Up the Creek with Many Possible Paddles: How Students Cope with Being Awash in Data at DataFest 

September 2, 2025 Leave a Comment

Jessica Karch, Jennifer Noll, James K. L. Hammerman, and Traci Higgins

As big data becomes ubiquitous across multiple sectors, it is increasingly important to create opportunities for students to work with large, authentic, complex data—referred to as LAC data here—during their undergraduate training. Working with LAC data is challenging, as students need to extract meaning from the data and work and think like data scientists. “In a data science situation—especially as a beginner—you often don’t know what to do. You have way too much data, and the data you have is confusing. Even though you’re not literally in a boat, you are awash in data,” write Tim Erickson and Ernest Chen in the introduction to “Introducing Data Science with Data Moves and CODAP.” 

To cope with being awash in LAC data, students need to manage complexity and navigate many possibilities. This can lead to feeling overwhelmed, especially as students move from formal classroom contexts to navigating real-world data that has not been cleaned, pre-structured, or connected to a chapter teaching a particular technique.  

DataFest presents a unique opportunity for students to engage with LAC data. In many ways, the DataFest challenge mimics real data science work: The data is messy and complex, with many observations and variables. Also, the task is not prescribed but instead must be constructed around a meaningful problem that can be addressed by the data and that yields solutions or insights the client can act upon. 

These authentic data investigations involve a complex workflow involving the following six actions, as Hollylynne Lee and her collaborators describe in their 2022 Statistics Education Research Journal article, “Investigating Data Like a Data Scientist: Key Practices and Processes”: 

  1. Framing the problem in relation to the real-world phenomena and broader context and pose investigative question(s) 
  2. Considering and/or gather data, which may involve examining the data at hand and considering data-gathering methods, the measures, the type of data and how it is structured, sample size, and what questions it can address and whether additional data is needed 
  3. Processing the data by organizing, structuring, cleaning, and possibly transforming the data or computing new metrics 
  4. Exploring and visualizing to identify patterns and relationships 
  5. Considering models 
  6. Communicating and proposing action to stakeholders 

This workflow is typically iterative and nonlinear. For example, a data scientist may realize the current structure of the data does not allow them to answer the problem as formulated and engage in a further round of considering the data, identifying different variables of interest, processing the data to arrive at new metrics, or merging in additional data—which in turn leads to a reformulation of the question.  

Over the course of two years at six sites, the research team for our Improving Undergraduate STEM Education: Directorate for STEM Education-funded study interviewed participants from 28 teams approximately one week after DataFest about their approach to the DataFest challenge. We used Lee and collaborators’ framework for the data investigation workflow and Erickson’s concept of authentic data science as producing an experience of feeling “awash” to explore how teams navigated the complexity of working with LAC data.  

Our qualitative analysis of the interview data revealed teams often expressed a sense of being awash in the data early in their investigation as they worked on developing a framing. In traditional, hypothesis-driven statistics, problems are well defined and the data collection is designed to flow from the investigative question(s). In the classroom, the data may be given, but it fits within the context of learning new techniques and students are often given the problem they are expected to solve.  

At DataFest, teams interact with very large, messy, complex data sets and must consider how to plot a path through that data to some sort of actionable insight or data product the data donors or judges recognize as meaningful and relevant. To narrow down options and give the investigation focus, teams sought a way to frame the problem and generate investigative questions that could anchor their work. We found the following three common anchoring strategies that teams used to find this framing: 

  1. Lean in on contextual or domain knowledge to generate questions that feel meaningful or relevant 
  2. Formulate a problem space that allow for showcasing of certain technical skills or techniques 
  3. Adopt a data-driven approach using simple exploration of the data to identify where the data seem to clump or patterns that were interesting or unexpected 
Strategy 1. Leveraging contextual knowledge

Students used salient domain knowledge to help frame their investigation. In our first year of data collection in 2023, students were asked to make sense of data from the American Bar Association’s online tool that connected clients with lawyers for pro bono legal advice. One team had domain knowledge about incarceration in the United States from a previous course. They used this domain knowledge as an initial way to frame the data, interrogating it for data about incarcerated people. Although this strategy did not pan out (the legal issues were civil, not criminal, so there was minimal data about incarceration), this outside knowledge initially provided some traction, giving their work direction. This helped them overcome the feeling of being overwhelmed by the size and complexity of the data sets.  

When this framing was discovered to be at odds with the data available, the team shifted momentarily to a data-driven approach, exploring question categories to see which accounted for the bulk of the data, but then again leaned on contextual knowledge as they reformulated their framing of the problem to focus on understanding characteristics of clients posting under this category. 

Strategy 2. Showcasing and/or centering technical skills

Teams narrowed in on a way to frame the problem by taking stock of what computational or statistical tools they were familiar with and considering how applying those tools could lead to a particular framing of the problem space. For example, when one team first explored the data, they immediately began to consider what tools they could use to further explore the text data without doing a lot of time intensive processing. Topic modeling fit these criteria, so the team applied this technique in an exploratory way to gain a better understanding of the data. This analysis produced an unexpected finding. Their questions flowed from this finding and led to a productive framing when they began to discover that certain lawyers had an outsized impact on the data—they drew on their knowledge of “whales” from mobile gaming (Strategy 1) to conceptualize this trend and create a productive framing around investigating these lawyers’ answers.  

Strategy 3. Leveraging the data themselves 

Teams often immersed themselves in the data through exploratory data analysis. In this strategy, students used a variety of techniques to familiarize themselves with the data and see what emerged in a very open way. For example, in Year 2, which focused on data from an online statistics textbook, (see “CourseKata’s Experience as a DataFest Donor”), one team described feeling overwhelmed by the number of variables in the data set. To ground themselves, the team read the descriptions of the variables in the data definition document and began making graphs through a process of trial and error. The team had a vague sense of what their task was—give feedback to CourseKata—and visualizing the relationships between variables helped them see what stood out and what trends merited further investigation and framing. 

There are multiple entry points for teams of students to engage with LAC data during authentic data science experiences. When students feel “awash” without a way to frame the investigation, they can bootstrap using knowledge that speaks to the contextual domain, that draws on technical tools to process and work with the data, or that employs exploratory data analysis to seek patterns that demand further explanation. 

Just as there’s no one best way to approach LAC data, DataFest helps students learn that being a data scientist means being flexible and able to employ multiple strategies and sources of knowledge when navigating the full data investigation process. DataFest provides opportunities to learn how to integrate knowledge from computational, statistical, and domain knowledge sources to deepen their data inquiry process.

Editor’s Note: This material is based on work supported by the National Science Foundation under Grant No. DUE 2216023. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the authors and do not necessarily reflect the views of the NSF.

Big smile, long dark hair

Jessica Karch

Karch is a senior researcher at TERC, which uses qualitative and mixed methods to study science learning and learning environments at the undergraduate and graduate levels with a focus on equity.

    Female, Big smile, blond hair braided in pig-tails.

    Jennifer Noll

    Noll is principal investigator at TERC. Her background focuses on K-12 and undergraduate statistics and data science education through innovative curricula, technology, and teacher professional development. 

      Male, short gray hair, slight smile

      James K. L. Hammerman

      Hammerman co-directs the STEM Education Evaluation Center at TERC. For more than 20 years, he has worked as an evaluator, designer, teacher educator, and adviser for innovative statistics and data science education projects, engaging formal and informal learners of all ages to investigate and make sense of data.

        Female, shoulder length hair, big smile, dimples

        Traci Higgins

        Higgins is a senior researcher in STEM education at TERC. She has more than 20 years’ experience conducting research and developing educational materials, processes, and models to support STEM learning and teaching both in and out of school, focusing on K-8 mathematics, data science K-12+, the social sciences, and interdisciplinary thinking.

          Filed Under: Additional Features, DATAFEST, Special Features Tagged With: Awash in Data, data science, DataFest teams, datascientist, Investing, LAC data, Statistics Education Research Journal

          Reader Interactions

          Leave a Reply Cancel reply

          Your email address will not be published. Required fields are marked *

          Primary Sidebar

          Search

          More to See

          JSM 2027 Chicago, Illinois Logo

          Shape the Future: Invited Session Proposals Sought for JSM 2027

          August 3, 2026 By Meg Ruyle

          Laptop open next to the words "online community"

          How I Built a Portuguese-Language Analytics Community Online

          August 3, 2026 By Meg Ruyle

          New Member Spotlight: Collin Nill

          August 3, 2026 By Meg Ruyle

          colorful books

          Members Offer Advice for Writing as a Team

          August 3, 2026 By Meg Ruyle

          ASA HOME

          American Statistical Association

          Communications from the Executive Director

          ASA Leader Hub

          ASA Career Connect

          ADVERTISERS

          STATA
          SIAM

          Archives

          Categories

          Footer

          Editorial Staff

          Managing Editor
          Megan Murphy

          Graphic Designers / Production Coordinators
          Olivia Brown
          Meg Ruyle

          Communications Strategist
          Val Nirala

          Advertising Manager
          Christina Bonner

          Contributing Staff Members

          Kim Gilliam

          American Statistical Association
          277 South Washington Street, Suite 370
          Alexandria, VA 22314-3646
          Phone: (703) 302-1857

           

          Copyright © 2026 · Magazine Pro on Genesis Framework · WordPress · Log in