The 2024 Joint Statistical Meetings in Portland, Oregon, featured more than 20 sessions and panels about statistical data privacy and confidentiality. Following are a few highlights.
Multiple invited panel and paper sessions covering the federal statistical system, a new national data infrastructure, and national data governance highlighted the need for privacy-preserving technologies, tiered access regimes for data sharing, and a common approach to balancing the privacy and confidentiality of respondents against the needs for data sharing for public policymaking. Some speakers discussed specific research areas using statistical disclosure limitation to help share data to inform agricultural policies and access federal earnings data. In the case of accessing federal earning data, a discussion focused on the numerous ways the Internal Revenue Service’s Statistics of Income program—one of the 13 federal statistical agencies—is exploring methods such as synthetic data, differential privacy, and secure remote queries to help researchers and policymakers gain better insights into individual and establishment earnings data.
A common session topic was the creation of synthetic data sets, such as in the context of federal statistics and health data. Panelists argued that among an arsenal of tools, synthetic data is one valuable tool for sharing and disseminating data. While it may not be the right tool for every situation, it provides another possible solution.
For example, data stewards can use synthetic data in a tiered access system in unison with other data-release methods. A session about generating synthetic data focused on use cases for when synthetic data is a feasible solution, such as allowing researchers to prepare code for their statistical analyses prior to accessing the confidential data. For this use, the synthetic data does not need to provide statistically valid results but could be a “mock data set” (i.e., a data set that has the same schema as the confidential data set). In other use cases, synthetic data will need to provide statistically valid results.
However, it is important to emphasize that synthetic data will never allow for the same richness of analyses as confidential data because the complexity of the statistical relationships in the synthetic data is limited by the sophistication of the synthetic data generation model. In synthetic data, one can only detect those relationships, which are incorporated within the synthetic data model.
The tradeoff between disclosure risk and statistical utility is a key aspect of statistical data privacy, and numerous sessions highlighted this issue. One session focused solely on defining data utility. Although definitions vary in different contexts, the session speakers emphasized defining what utility of privacy-protected data means and preserving the usefulness of the data for research or policymaking.
Another session considered the tradeoff between risk and utility, with a focus on policy and proper implementation. Even just concretely defining what is meant by risk and utility is usually intricate and context dependent.
Each panelist provided their perspective on risk and utility, which differed in both subtle and not-so-subtle ways. One speaker advocated for differential privacy to numerically quantify risk, although they admitted differential privacy can be difficult to interpret and it is hard to achieve both satisfactory utility of the released data and meaningful privacy as measured by differential privacy.
Another speaker used the p-percent rule (i.e., whether the second-largest contributor can use the aggregate total to determine the largest contributor’s value to within p-percent) to measure disclosure risk in some of their data releases. Risk and utility are context-dependent due to the wide variety of types of data that need to be protected (e.g., survey data, social science data, agricultural data, etc.) and the variety of formats of published data (e.g., tabular data, microdata, regression coefficients, etc.), so this makes the discussion complicated.
Ultimately, the speakers all highlighted difficulties in communicating what risk and utility mean and the need to be able to better articulate these concepts to a diverse set of stakeholders who have varying degrees of privacy expertise.
James Bailie, PhD student, Harvard University and members of the ASA Privacy and Confidentiality Committee contributed to this article.

Leave a Reply