Sections

Commentary

Sharing more, protecting more: Three lessons from the Safe Data Technologies project

Shutterstock / MMD Creative

On June 4, the Department of Commerce issued a disclosure avoidance order that has reignited the heated discussion on the tradeoffs among data privacy, access, and accuracy. The order has focused on the choice of suppressing and coarsening the data instead of using “noise infusion,” which includes data swapping and synthetic data to protect official statistics and data. The challenges of balancing these competing needs are not unique to the U.S. Census Bureau and Bureau of Economic Analysis. Across the federal statistical system, agencies must provide timely evidence while maintaining strong privacy protections in an environment where data are fragmented across agencies, programs, and legal authorities. This situation raises the question of what other privacy-enhancing technologies and techniques (PETs) are available to balance privacy, access, and accuracy without withholding more data or reducing the information to levels that are not useful.

The Safe Data Technologies (SDT) project, a collaboration involving the Statistics of Income Division (SOI) at the IRS, the Urban Institute, and other partners, offers a case study for how new PETs can move from theoretical concepts into operational tools and provide evidence on appropriate PETs to improve the tradeoff. Although the PETs developed for SDT focused on expanding access to confidential tax data through synthetic data and privacy-preserving validation servers, SDT provides lessons learned that extend well beyond administrative tax data and synthetic data applications. In other words, SDT may help inform ongoing debates about the role of synthetic data under the order and across a wider range of federal statistical and data access contexts.

This blog presents three lessons that other agencies across the federal statistical system and potentially others, such as state statistical systems, may learn.

1. Innovation from moving theory to practice requires different types of partnerships

One reason SDT has advanced beyond a research prototype is that it was supported through multiple funding streams and partnership models, which has been both valuable and challenging. The opportunity lies in drawing on different types of support at different stages of development, where the challenge is maintaining a consistent long-term vision while balancing the priorities and expectations of the multiple funders.

In our case, those partnerships aligned with different stages of the project. Early work was supported by Arnold Ventures and the Alfred P. Sloan Foundation, whose investments enabled exploratory research and prototype development. Sloan Foundation support has also advanced related efforts through the Re-Engineering Statistics using Economic Transactions and Economic Indicators Initiative projects, creating further opportunities to identify lessons learned and build complementary capabilities. Those initial efforts generated the evidence needed to secure funding from the National Science Foundation’s National Center for Science and Engineering Statistics (NSF/NCSES) and, ultimately, direct investment from the SOI as funding became available.

These funding sources served different purposes. Philanthropic funding created space for experimentation. Research grants supported technical advancement and evaluation. Federal agency funding had operational requirements, implementation, and long-term sustainability.

For federal statistical agencies and developers alike, the lesson is that innovation should be viewed as a multi-stage process that depends on strong, sustained partnerships. Different types of partnerships become valuable at different points, from initial proof of concept to eventual deployment. In other words, the project’s success depends on the collaboration of multiple partners. SOI, as an official statistical agency, along with academic and nonprofit researchers, technology developers, and funders, each contributed unique expertise, resources, and perspectives that helped advance the project from research to operational use while maintaining a shared long-term vision.

2. Design for portability and plan for adaptation

Many of the challenges faced by federal statistical agencies are similar. Agencies must protect confidential information, provide access to approved users, comply with legal requirements, and operate under resource constraints, including limited staff time and computing capacity.

These shared challenges create opportunities to develop PETs that can reduce administrative burden and accelerate data access and analysis. However, the SDT project demonstrates that successful adoption depends on much more than whether a technology works in theory.

As the project matured, the SDT team found after engaging with others outside of SOI that implementation requirements varied substantially across government entities, from local agencies to federal statistical organizations. Infrastructure, governance processes, hardware, security requirements, and staff expertise all influence whether a system can be successfully deployed and maintained.

This challenge becomes even more apparent when considering how technologies developed for one agency can be adapted for another. A system designed around administrative tax data and SOI workflows may require significant modifications before it can be used by the Census Bureau, Bureau of Economic Analysis, Bureau of Labor Statistics, or state data systems.

As the core SDT team considers expanding the work beyond federal tax data, a major focus has been identifying which components can be standardized or packaged for broader use and which aspects must remain tailored to local environments. To support this effort, the team is in the process of engaging state and local partners to assess technical, governance, legal, and operational barriers to adoption; provide feedback on implementation requirements and governance frameworks; and participate in project reviews and testing activities.

The broader lesson is that developers should engage a diverse set of stakeholders and use cases early in the process. Building PETs that work in a single environment is not enough. Adoption of PETs requires understanding where agencies share common needs, where they differ, and how a solution can be adapted across those differences without losing its core functionality.

3. Building agency capacity matters as much as building technology

New technologies and techniques often receive the most attention, but one of the most important lessons from the SDT project is that organizational capacity matters as much or more than the technology and techniques themselves.

For instance, implementing new PETs introduces new questions for agency leaders and data stewards. How should privacy risks be evaluated? How much accuracy is sufficient for a given use case? What governance processes are needed before a new system can be trusted in a production environment?

The SDT project has invested heavily in helping answer these questions while engaging and training SOI staff. Through these engagements, the project created documentation, open-source tools, several internal and external dissemination materials, stakeholder engagement, technical guidance, and advisory activities that help agencies understand both the capabilities and limitations of the technology. We have carried this approach forward as we adapt the synthetic data and validation server framework for state agencies, where we focus on governance requirements, implementation guidance, operational planning, workforce readiness, and hands-on technical assistance. Through training sessions, co-design activities, and ongoing support, our goal is to help agency staff develop the knowledge and confidence needed to operate and govern these systems without our support.

The lesson is that successful adoption requires more than acquiring a new technology. Agencies need the tools, training, governance frameworks, and internal expertise necessary to make informed decisions about implementation and long-term use once a contractor or vendor is no longer around. This is especially salient today, as agencies are adjusting their disclosure policies. New PETs or other technologies alone will not solve organizational challenges. Lasting impact comes from building the capacity within agencies to evaluate, adopt, and manage new approaches effectively.

Future of PETs in government

The future of federal statistical agencies still requires navigating fragmented data systems, evolving privacy regulations, constrained resources, and growing demand for timely evidence. SDT project illustrates that moving from innovation to impact requires three things: multiple strong partnerships that support different stages of development both technical and funding support, technologies and techniques designed with real-world implementation in mind, and investments in the people and governance structures responsible for operational decisions.

The importance of these lessons learned extends beyond SDT. For example, NSF/NCSES’ SEDSyn-23 project brings together NCSES, Urban Institute, RTI International, RAND, and Vassar College to develop and evaluate synthetic versions of the Survey of Earned Doctorates. Similar to SDT, this project demonstrates how collaboration among statistical agencies, researchers, and technology experts, combined with a focus on real-world implementation and investments in people and governance structures, is essential for advancing PETs from research to practice.

For agencies considering adopting PETs, these lessons demonstrate that long-term success depends on building the partnerships, expertise, and governance frameworks needed to sustain innovation long after a pilot project ends.

Author

The Brookings Institution is committed to quality, independence, and impact.
We are supported by a diverse array of funders. In line with our values and policies, each Brookings publication represents the sole views of its author(s).