
The southern California region played host to two ASA DataFests this spring. The first was at the University of California at Los Angeles (UCLA) April 5–7 and included more than 350 students from UCLA, the University of California at Riverside, University of Southern California, Pomona College, and Cal Poly San Luis Obispo (orange circles on the map). The second took place May 3–6 at Chapman University, which hosted 128 students from Chapman, Cal Poly Pomona, California State at Fullerton, California State at Long Beach, Orange Coast College, University of California at Irvine, University of California at San Diego, and University of California at Santa Barbara (blue circles on the map).
This year’s data challenge came from Ming-Chang Tsai, a researcher at the Canadian Sports Institute. The data consisted of GPS and accelerometer records from every game played in the previous season by the Canadian National Women’s Rugby 7 team, as well as data on daily training and medical reports. The challenge issued was to comment on the role of fatigue on the team. The students took a variety of approaches. Many struggled with methods for using self-reported measures of such items as exertion, sleep quality, and mood with more objective medical measurements. Others chose to focus on particular players and examine the variability in their daily exertion and relationship with game play.
One of the main prizes in the Chapman University DataFest went to a team from a two-year college, which is a first. Team Memory Leak (Phuoc Do, Hector Elias, Jacob Leenerts, Ryan Millett, and Naomi Valentin) from Orange Coast College took the Best Use of External Data. The Best Insight award at Chapman went to team git rekt (Natanael Alpay, Christina Berardi, Dylan Davis, Noah Ferrel, and Jennifer Prosinski) from Chapman University, and the Best Data Visualization award went to team Joint Distribution (Brett Galkowski, Eisah Jones, Jason Kahn, Satyam Tandon, and Shivan Vipani) from UC Irvine.
Honorable mention for Best Insight at Chapman went to team Resonant Skunks (Nic Cordova, Max Haggard, Christopher Moore, and William Simmons) from Chapman University, and honorable mention for Best Data Visualization was awarded to team Deep Data (Ivy Gu, Brandon Newberg, Raymond Nguyen, Junlin Wang, and Hanwen Ye) from UC Irvine.
The Chapman judges chose to honor two additional teams for their outstanding work. Team Top Quantile (Anastasia Franio, Joseph Gadbois, Kristy Le, and Earl Zedd) from Cal State Long Beach were recognized for the Best Use of Statistical Software, and team unDATAble (Estaban Escobar, Ryan Flynn, Abram Garcia, Christine Hoogendyk, and Kevin Tsoi)—with students from both Chapman and Cal Poly Pomona—were recognized for Best Data Forensics.
The UCLA event was its largest yet, with more than 70 teams competing. Judging took place in two rounds. In the first, teams were randomly assigned to one of five rooms, and a panel of judges chose three teams from each room to promote to the final round. Fifteen teams then presented their findings a second time to a new panel of judges, who chose the finalists.

The prize for Best Insight went to the Wild Cardinals from UCLA (Marina Hu, Sarah Truax, Hunter Carlisle, Jonathan Chang, Devyn Fisher). Surprisingly, the members of Wild Cardinals had never met before DataFest, and yet functioned as a well-oiled machine. The coveted Best Visualization Award at UCLA went to team Above Average from Cal Poly San Luis Obispo (Nicole Hill, James Kao, Evan Shui, Jenna Landy, and Markelle Kelly). Judges awarded the Best Use of External Data prize to Pomona College’s Flock of SQLs (Madelyn Andersen, Amy Watt, Connor Ford, Adam Rees, and Ethan Ashby). Many thought that if a prize had been given for best team name, they would also have won that.
The UCLA judges awarded three Judges Choice awards to teams with exemplary work that didn’t quite fit into the main award categories. These prizes went to team Naive Baes (Ambrish Parekh, Paul Boulos, Yifun Zhu, Qufei Wang, and Chelsea Lee) from USC, The Wranglers (Alexander Lao, Kyle Maxwell, Maxwell Blau, Charlie Liou, and Hanson Egbert) from Cal Poly San Luis Obispo, and team 99th Percentile (Jiayu Lyu, Shiyu Ji, Yanghui Wang, Xinran Qian, and He (Iris) Yang) from UCLA.
The success of ASA DataFest depends greatly on its visiting “mentors,” who offer advice and wisdom and sometimes help students get out of ruts or debug code.
The UCLA event was organized by Linda Zanontian and the UCLA Stats Club DataFest committee, led by students Margaret Koulikova and Pashmeen Kaur. The Chapman event was organized by Michael Fahy and Madeline Bauer.
Supporters of ASA DataFest at UCLA included the UCLA Department of Statistics, UCLA Division of Physical Sciences, UCLA IS Associates, true[X], Beyond Yoga, Wolfram Language, Stickeryou.com, Subway, Brothers All Natural, Sticker Mule, and Red Bull. Supporters of the ASA DataFest at Chapman included QAD, Data Tree by First American, the Orange County/Long Beach ASA Chapter, Rogue Cloud, Chapman University Schmid College, Chapman University MLAT Lab, Concise, R Consortium, RStudio, Missing Variables, Monster Energy, and Essentia Water.
ASA DataFest is a national event held at multiple institutions throughout the United States, Canada, and Germany. This year, 41 sites hosted a DataFest, with an estimated 3,000 students participating. To learn more or get involved, visit the DataFest website.


Leave a Reply