
With a PhD in statistical astrophysics, David Corliss works in analytics architecture at Ford Motor Company while continuing astrophysics research on the side. He serves on the steering committee for the Conference on Statistical Practice and is president-elect of the Detroit Chapter. He is the founder of Peace-Work, a volunteer cooperative of statisticians and data scientists providing analytic support for charitable groups and applying statistical methods to issue-driven advocacy in poverty, education, and social justice.

When disaster strikes, great minds turn to more and more … hackathons! Data for Good researchers perform vital roles in mining social media, mapping, geospatial analysis, resource optimization, and other analytic functions, partnering with organizations that put boots on the ground to provide critically important information, logistical support, and guidance for hands-on relief efforts. Crowdsourcing can play an important role in bringing together many resources in a short time in response to a crisis. As a result, many organizations now sponsor hackathons and data dives and create applications to collect crowdsourced information from people in affected areas.
One of several recent events was a hackathon at RPI in September. With many more than 100 people participating, the Call for Code Hackathon: Natural Disaster Preparedness and Relief Code Challenge was organized as a competition, with a cash award for the winner and an opportunity to work with IBM to implement their project. In addition to an opportunity to work on this important project to help people in need, students and the sponsoring company also benefitted from the opportunity to meet and get to know each other.
Many government organizations are involved in crowdsourcing events. One example is the Houston Hackathon, an annual event focusing on local challenges and issues. Now in its sixth year, the event on May 18 hosted by Sketch City brought together more than 400 people for 24 hours of Data for Good. This event has served as a catalyst for the Houston analytic community to focus efforts in time of need. Following Hurricane Harvey, which struck Texas last August, Houston Mayor Sylvester Turner described efforts: “While it was still raining, Sketch City and others in the civic technology community were creating solutions to aid our rescue and relief efforts. … These products saved lives, put resources in the hands of those in need, and are now being used to assist other communities after disasters.”
One government agency playing an important role in organizing crowdsourcing efforts is the Federal Emergency Management Agency (FEMA). We usually think of FEMA as the government resource for direct, onsite support for disaster response. In addition to all the efforts on the ground, a huge logistical effort is needed behind the scenes—and that means data and analysis.
FEMA’s Disaster Crowdsourcing Exchange
Late last year, FEMA organized a Data for Good disaster response event at their Washington offices and online. As one of the online participants, I was able to see first-hand how statistical volunteers worked together to crowdsource data vital to the recovery efforts. Scores of participants were guided by a core group of professionals to “ground truth” images: assigning the correct identity objects appearing in the digital image. Efforts focused on buildings, roads, and other landscape features. Many of these were severely damaged by the hurricanes, making old maps obsolete. Even areas with less damage have required new development due to incomplete features in earlier maps.
As an online participant for this combined virtual and in-person event, I found the infrastructure generally worked well. There were occasional connectivity issues and sometimes speakers at the in-person meeting forgot the online people. After registration and a general introduction to the project, the crowdsourcing event began with training on the mapping software used to capture the data. It took about an hour to get familiar with the tool and begin to capture useful data.
While many people worked on the mapping effort, other opportunities to contribute included designing new projects, testing software, and developing the strategy for crowdsourcing projects in the future. Volunteers selected a number of opportunities to help, dividing them into project teams. Team leaders met with their volunteers to give specific guidance for individual tasks. After this briefing, the volunteers began working with the data.
Here is how their process for crowdsourcing ground-truth data works. First, images are taken of affected areas. Individual images are pieced together in a mosaic to produce a complete picture. This digital image is then converted into an editable map, which can be annotated with descriptions of objects appearing in the image. The entire map is broken into sectors, and a high-level map is created with thumbnails of the map sectors. Volunteers, who had been trained on the software, clicked on the thumbnail of a map sector to bring up a view of the editable map. The volunteers moused across a feature visible on the map to add a polygon corresponding to the object. The shape recorded in the map is annotated using drop-down menus to describe the designated object. The software has different drop-down lists for each type of object, with items on the list giving more detail about the object (e.g., a dirt road, minor paved road, or highway). Users can save their data to the editable map and come back later to complete a sector.
The sectors I saw took perhaps a half hour to complete. This gives volunteers the flexibility to contribute larger or smaller blocks of time, as their schedule and commitments permit.
All this improves the ability to plan and execute disaster response. New data on roads enables disaster response to reach the most remote areas. Ground-truthing buildings enables responders to see where homes and business have been damaged or destroyed and identify what resources appear to have survived the disaster.
While this event responded to a disaster, crowdsourcing data collection to support Data for Good projects is found in many areas, including public health, crime, and homelessness. New crowdsourcing data-sharing applications like GatherIQ from SAS enable a widening community of volunteers to share data supporting analytics for good causes. Google your area of interest and there’s a fair chance someone is crowdsourcing data today. If not, start your own!
Do you know of a Data for Good crowdsourcing data project? Let me know! I would love to hear more about it. Maybe your project can be featured here in Stats4Good!


Leave a Reply