Written evidence submitted by the OpenMined Foundation (SMH0046)
OpenMined appreciates the opportunity to provide evidence to the inquiry into the links between algorithms used by social media and search engines to rank content, generative AI, and the spread of harmful or false content online. With society increasingly exposed to algorithmically-driven online services and AI-generated content, we support the Committee's efforts to gather evidence and support research into the potential negative impacts of these tools and services.
Whilst our response speaks directly to some of the questions put forward in the call for evidence, the main point we wish to communicate is that many of the questions put forward simply cannot be answered in an evidence-based way due to access barriers that restrict meaningful external research and oversight of the algorithmic systems that drive content recommendations on online platforms. We are hopeful that new requirements under the Online Safety Act 2023 can help to alleviate these challenges, but the effectiveness of the new regulatory regime will depend on how the Government and Ofcom choose to operationalise the Act’s provisions. Here, we provide details of projects OpenMined has conducted to facilitate researcher access with X, Microsoft’s LinkedIn, Dailymotion, Reddit, and the governments of France, New Zealand, and the United States. By leveraging open-source data infrastructure powered by privacy-enhancing technologies (PETs), we demonstrate how a new paradigm of access – that facilitates access to answers, not datasets – can overcome these barriers and thus enable third-party researchers, Government, and regulators to answer important questions about the impact of algorithms and AI systems on society.
OpenMined is a global non-profit that specialises in the problem of facilitating access to internal, secure data and AI systems by external parties — often using PETs. We began in 2017 as an informal, online meeting point for researchers from different fields relevant to this problem. Since then, our free online courses have garnered 20,000+ student registrations, our freely available, open-source software has been contributed to by over 500 open-source contributors, and numerous academics, startups, industry, and policy professionals have launched or pivoted their careers during their time in our community. Broadly speaking, we are well-positioned to represent the perspective of the PETs community — and its relevance to third-party auditing of proprietary AI systems.
Our work has been featured in the White House’s 2023 report on Privacy-Preserving Data Sharing and Analytics[1], the Royal Society’s 2023 report on PETs, and the United Nations’ 2023 PET guide. We have worked with the private sector, including organisations likely to be classified as ‘categorised services’ under the Online Safety Act, building and deploying privacy technology with: NVIDIA, Google (1, 2, 3), Meta (1, 2, 3, 4, 5), Microsoft (1), Twitter/X (1, 2), Dailymotion, and Reddit (1, 2). Our work with the public sector includes co-launching and chairing the United Nations PET Lab, deploying on the United Nations Global Platform, and testing and improving our software in collaboration with UN PET Lab Members, which include national statistics offices in the United Kingdom, Netherlands, Italy, and the United States. We have also entered a technical partnership with the UK AI Safety Institute to help develop and deploy digital infrastructure hosted by the Institute physically in the UK to facilitate external access to data relevant to AI safety research. Taken together, we have worked extensively on this problem with various members of the public and private sectors, both in the UK and abroad.
As a non-profit, we are supported by a diverse pool of funders, including Open Philanthropy, the Center for Security and Emerging Technology within Georgetown University's Walsh School of Foreign Service, the Alfred P. Sloan Foundation, the Office of the Prime Minister of New Zealand, and additional funders. As a result of this generous support, everything we produce we give away for free, and — amidst a sea of startups — we are one of the only well-funded non-profits focusing on the end-to-end problem of external access to private data. We believe we are uniquely positioned to offer neutral, objective, and informed comments on the technical conditions and procedures necessary to enable robust external oversight of platform systems and data.
In early 2022, OpenMined partnered with Twitter (now X) to advance algorithmic transparency. Over the following year, we worked closely with Twitter’s Machine Learning Ethics, Transparency, and Accountability (META) team on the challenge of facilitating external research on their production recommender algorithm. The project focused on measuring political bias during the 2020 U.S. Presidential Election. While we had partnered with technology companies on privacy technology many times before, facilitating the study of large-scale production systems running on private user data came with unique privacy, security, and intellectual property challenges. Finding solutions required work across our own engineers and researchers in partnership with Twitter’s product, engineering, research, legal, and policy teams. Even with buy-in from senior stakeholders — who launched the project voluntarily — external access to internal data and algorithms is not a trivial task.
By the end of 2022, the project with X was successfully expanded into a new scope: the Christchurch Call Initiative on Algorithmic Outcomes (CCIAO), which was announced in a joint press conference between New Zealand Prime Minister, at the time, Jacinda Ardern and French President Emmanuel Macron. Phase 1 of this initiative started with New Zealand, the United States, Twitter, and Microsoft, with OpenMined joining later as a technical partner. France and Dailymotion joined soon after to complete the full list of participants for Phase 1 of the initiative. In Phase 1, four researchers performed quantitative analysis on impression data related to the production recommender systems at two online platforms and completed six research projects to better understand the relationship between the algorithms and the content recommended to users.[2] The access modality chosen for these research projects was remote data science - performing analysis using data a researcher cannot see - powered by a novel digital infrastructure, PySyft. This access modality was chosen to protect data providers' legitimate interests, including their trade secrets, confidentiality, and the security of their services.
Remote data science powered by PySyft was again chosen earlier this year when OpenMined partnered with Reddit to develop an external researcher access program and increase the scale of researchers collaborating with Reddit and each other. The program has currently reached a few dozen researchers who are actively engaging with Reddit data around a range of important and impactful topics, including analysing and improving the quality of political and cross-partisan discussion, developing models for detecting and managing AI-generated content, and exploring strategies and tooling to support healthier, more inclusive online communication, to name a few.
Like the project with Twitter, our recent projects with LinkedIn, Dailymotion, and Reddit have been interdisciplinary efforts across research, engineering, policy, legal, and product disciplines. In line with our non-profit mission, all the infrastructure we build is freely available under open-source licenses and deployed inside a data provider’s internal software system, allowing a data provider full control and oversight over their private data without any dependency on external third parties - including OpenMined. Participation by independent researchers also happens voluntarily and for free — this is our charitable mission, which is supported by our donors. Based on these experiences, we share some of what we’ve learned below.
As widely reported in the press,[3] various social media, news, and messaging platforms were mediums through which misinformation was broadcast that likely contributed to inciting the Summer riots, and through which organisers were able to coordinate their actions. However, in the current information environment it is very difficult to make more concrete conclusions about the role of social media and their algorithmically-driven recommender systems. The main sources to draw such conclusions lie within social media platforms, which are not incentivised to voluntarily admit their complicity in spreading harmful misinformation. External parties with incentive to uncover the truth lack sufficient and appropriate access to the main sources of information about these systems and their usage.[4] For example, it is exceedingly difficult to answer critical research questions such as:
● What was the scale of misinformation on platforms in the lead up to the riots?
● How did harmful posts propagate across different platforms?
● What impact did platforms’ recommender systems have on propagating or amplifying misinformation or other harmful content related to the riots?
● Were there posts that were created by generative AI tools that contributed to the spread of misinformation? What was the scale of this?
● Were there statistically significant, causal links between the design of content recommender systems on social media platforms and the riots that followed?
The critical problem preventing these questions from being answered is that external researchers, regulators, and government agencies typically do not have access to data on how proprietary AI systems are being used in the real world, due to legitimate barriers associated with user privacy, security, and intellectual property. This limits important research to internal teams, who may keep their findings private, and prohibits research into cross-platform phenomena or research that can establish causation by linking platform data to outcomes data. The result is that government, the regulator, and society at large has a poor understanding of the societal impacts of these widely used systems, relying primarily on anecdotal evidence or small-scale studies. This lack of a robust evidence base means there is a significant risk that resulting policy interventions – no matter how well-intentioned – are ineffective in practice.
The introduction of the Online Safety Act in 2023 includes provisions that may enable Ofcom to have more rigorous oversight of proprietary algorithms, particularly when they are suspected of contributing to harmful outcomes. Specifically, Ofcom is provided with information gathering powers that include the ability to “observe the carrying out of empirical tests” run by a service provider[5], and to appoint a skilled person – defined as “an external third party with relevant expertise” – to carry out an independent assessment of a service provider’s compliance with specific duties.[6] These new powers mark a significant step forward in facilitating greater accountability. However, they have only recently been under consultation, so it is too early to draw conclusions about their effectiveness.[7] OpenMined has previously made recommendations for how these provisions can be operationalised most effectively.[8]
Whilst these information gathering powers are very welcome, they still constrain who can access platform data (i.e. Ofcom and Ofcom-appointed skilled-persons) and under what circumstances. We are therefore encouraged to see an amendment to the Online Safety Act in the Data Use and Access Bill focused on creating a regime for broader researcher access. As with the information gathering powers, the effectiveness of this regime will depend on how it is operationalised. There is a similar regime under the European Union’s Digital Services Act which came into force in 2022 – whilst the DSA has mandated that large platforms provide APIs for external access, early evidence[9] suggests there has been limited usage due to a lack of awareness of the existence of these APIs, lengthy application procedures, and a high burden of liability placed on researchers using the APIs, amongst other challenges.
Our contention is that these challenges can largely – if not entirely – be overcome by building services that provide researchers the ability to run research queries against platform data remotely, retrieving the results of their queries rather than having direct access to the underlying data. In short, we advocate for infrastructure that can create a paradigm shift of facilitating access to answers, not datasets. Such infrastructure was leveraged by the CCIAO project to successfully empower third-party researchers to conduct audits on proprietary production recommender systems at LinkedIn and Dailymotion. We provide further details on how this infrastructure works below.
Modern PETs allow for a fundamental change in how external research is framed. Current external access conversations narrowly discuss external researchers obtaining “access to data” - assuming that researchers need direct read access or a copy of raw datasets in order to perform their analysis. But data is merely a means to an end. What researchers really want is access to answers. They aren’t after 1 billion tweets. They want to know whether the tweet ranking is biased based on race or gender. They aren’t after ten million video uploads — they want to know whether TikTok’s video feed is driving mental health issues in teenagers. They don’t want a database; they want the final histogram or table of metrics in their final research paper, which will surface an important insight about society. Everything else is a means to that end.
Several PETs focus on facilitating the creation of (verified) answers without seeing the underlying data. The PETs industry hasn’t coalesced on a term for this yet — data spaces, federated learning, trusted research environments, trusted execution environments, data clean rooms, secure enclaves, secure multi-party computation, remote data science, and other terms all encompass this ideal — but we call this structured transparency.
In short, structured transparency is about a new approach: researchers access answers instead of data. We have found that this approach has enormous implications for running external research programs. First and foremost, it does not infringe upon the rights and interests of the data provider, including the protection of their confidential information, in particular, their trade secrets and the security of their services. When an external researcher sees raw data from a service provider, there’s almost no way to guarantee that they won’t upload it to the dark web, sell it on the side, or use it for use cases other than what they promised. In a structured transparency framework, this is called “the copy problem.” If a data provider shares a copy of a dataset with a researcher, the data provider can no longer control how the researcher will use the dataset. Once a data provider makes a copy of a dataset and gives that copy away, they lose technical control over the information and have to trust that the researcher or any other recipient of the copy will not misuse the information.
Social institutions attempt to prevent people from misusing a piece of information; the United States Government passed HIPAA to protect medical information and enforces that law through various regulations. The European Union has GDPR, and the UK has UK GDPR. California has the CCPA. But these are difficult to enforce, as once information is copied, there's no guarantee a data provider, an oversight authority, or anyone else can find out where the information went, what it was used for, or do anything about it. A researcher can sign every legal agreement under the sun, but in most cases — if they obtain raw data — preventing misuse is broadly unenforceable.
Our ability to copy information is nearly impossible to stop without an incredible reduction in individual liberty. The consequences are often associated with and initially felt by the data providers, who typically make the initial trade-off when deciding whether to share information, weighing the benefits of sharing information with the risks of misuse. However, the long-term consequences are intimately related to people, as the information often relates to personal details about their lives, and the misuse of a copy of their data translates to harm they experience in the digital world, physical world, or both.
If a researcher obtains access to data, data providers' concerns can be both legitimate and significant. However, if the only thing a researcher ever acquires is a verified answer to a specific question they propose, then this concern is mitigated. This concern is also mitigated by reducing the amount of copies of data available - less attack surface, less potential for misuse, and greater personal privacy. Furthermore, if the researcher never acquires a copy of the data, their legal liability can be significantly reduced. This can incentivise more researchers to participate in these research programmes, and can massively simplify the onboarding processes.
The CCIAO project demonstrates the application of structured transparency in practice. Researchers were able to conduct quantitative research into proprietary production systems leveraging infrastructure that overcame the copy problem through the application of two PETs: remote execution and differential privacy.
At a high level, this setup works as follows: a service provider loads relevant data assets (e.g., user logs, impression data, etc.) into a high-side server deployed on their local infrastructure inside their firewall. They then deploy a low-side server that contains mock assets—assets that directly imitate the structure of the real assets but contain fake, non-sensitive information. The low-side server is made accessible to authorised external researchers so that they can prepare, test, and iterate on their audit/evaluation code using mock assets downloaded to their local machine. This step ensures that researchers specify their code with appropriate precision to get the appropriate result. We refer to the high-side and low-side servers together as comprising the service provider’s Datasite.
Once content with their code, the researcher shares it with the service provider. The service provider confirms the audit/evaluation goals are as specified in the code and can then approve the project to be executed against the private assets on the high-side server and return the result to the oversight organisation. The researcher now has the results of an audit/evaluation run against the private assets, crucially without ever directly seeing the assets. The application of PETs such as differential privacy streamline the approval process, by providing mathematical guarantees that private information cannot leak from the dataset. More detailed technical information can be found in the CCIAO Phase 1 report.
Finally, we note that generative AI tools (e.g. ChatGPT, Gemini, etc.) introduce new challenges for meaningful external oversight. Whilst research into the impacts of traditional social media has been hamstrung by access challenges, there has still been crucial research carried out by civil society, academia, and the OSINT community by leveraging social media content available on the open web. With generative AI, similar research may not be possible given the prevailing interface through which individuals interact with these systems is a private chatbot interface. Understanding the societal impacts of such tools requires access to user logs from these chatbot interfaces, which poses major challenges around privacy and surveillance. We therefore strongly encourage the Committee to consider oversight implications not just for traditional social media, but also for emerging generative AI systems, and to explore the role of privacy-preserving mechanisms for facilitating research into their impact on society.
In closing, we emphasise that an effective external access regime is fundamental for understanding the societal impacts of social media platforms. Without this, we cannot robustly understand the role that these platforms – and the algorithmic systems that support them – may have played in events such as the Summer riots.
We propose that privacy-preserving infrastructure based on novel PETs offers a promising path forward for implementing data access provisions in a way that better serves both researchers and data providers. By shifting from a paradigm of "access to data" to "access to answers," such a regime could significantly reduce security risks, administrative burden, and infrastructure costs while increasing research efficiency and scale. Making this paradigm shift would ensure meaningful research access while protecting platforms' legitimate interests and user privacy, in a more technically sophisticated and sustainable way. In doing so, we can facilitate a nuanced, continuously developing understanding of the impacts of algorithmic systems, which can inform policy interventions that more effectively curb harms and foster greater accountability.
18 December 2024
[1] Search “PySyft”
[2] To learn more about the outcomes of CCIAO Phase 1 see Chen et al. (31 October 2024). AI Transparency
in practice: What was learnt from third‑party audit of recommender systems at LinkedIn and Dailymotion. https://www.christchurchcall.org/content/files/2024/11/Christchurch-Call-AI-Transparency-in-Practice-Report-October-2024.pdf
[3] See e.g. Violent Southport protests reveal organising tactics of the far-right, BBC News
[4] This is in response to the questions:
● To what extent do the business models of social media companies, search engines and others encourage the spread of harmful content, and contribute to wider social harms?
● How do social media companies and search engines use algorithms to rank content, how does this reflect their business models, and how does it play into the spread of misinformation, disinformation and harmful content?
● What role did social media algorithms play in the riots that took place in the UK in summer 2024?
[5] See S. 100(3) of the Act, and p.66 of its Explanatory Notes
[6] See S. 104 of the Act, and p.67 of its Explanatory Notes
[7] This is in response to the question: How effective is the UK's regulatory and legislative framework on tackling these issues?
[8] See OpenMined’s response to Ofcom’s information gathering consultation here
[9] See Report on EDMO Workshop on Platform Data Access for Researchers
[10] This is related to the question: What role do generative artificial intelligence (AI) and large language models (LLMs) play in the creation and spread of misinformation, disinformation and harmful content?