Science, Innovation and Technology Committee
Oral evidence: Innovation showcase, HC 523
Tuesday 11 March 2025
Ordered by the House of Commons to be published on 11 March 2025.
Members present: Chi Onwurah (Chair); Emily Darlington; Dr Allison Gardner; Steve Race; Dr Lauren Sullivan; Adam Thompson; Martin Wrigley.
Question 9
Witness
I: Dr Jennifer Fleming, Co-ordinator, Protein Data Bank in Europe, EMBL-EBI.
Witness: Dr Fleming.
Chair: Welcome to this morning’s session of the Science, Innovation and Technology Committee. First up, we have our innovation showcase. The Committee is seeking to understand how the UK supports innovators, and what more can be done. To inform our work, every week a different member of the Committee brings an innovator to share their story. This week it is Lauren’s innovator.
Q9 Dr Sullivan: Thank you very much, Chair. It is my great pleasure to introduce Jenny Fleming. We did our PhDs together in Dundee. Jenny has single-handedly made grams and grams of protein, which she crystallised beautifully, and was always my go-to person when, “Oh no, I’ve got a problem. How do I solve this?” She is absolutely amazing. She has worked in Europe, in Germany and now she is working at Protein DBe on AlphaFold. Given the real interest in AI and AlphaFold, and all these sorts of things, to solve our life science problems, I thought it would be brilliant to hear what Jenny has to say.
Dr Fleming: Good morning. Thank you for having me. I am Dr Jennifer Fleming. I am here today to share how the European Molecular Biology Laboratory, the EMBL, the European Bioinformatics Institute, the EBI, and the Protein Data Bank in Europe, the PDBe, are driving innovation in biomedical research. I will start by talking about Demis and the Nobel prize in chemistry last year, which was for AlphaFold. AlphaFold is an artificial intelligence system that was developed by DeepMind, a research tool from Google. Shortly after his win, Demis spoke about how the EBI helped in the Nobel prize.
AlphaFold revolutionised how we obtain protein structures by allowing researchers to predict highly accurate 3D structures of proteins. I have a model of a protein here. Proteins are macromolecules. They are essential building blocks of life. Understanding their 3D structure, which you see there in the model, is essential for uncovering how they function, how they are involved in diseases and how we can design drugs to treat those diseases.
Before AlphaFold, obtaining protein structures was one of my jobs. It was an incredibly difficult, expensive and time-consuming task. With AlphaFold we can predict the structures of proteins even when there is no experimental evidence available. This really revolutionised and answered the grand challenge of protein structure prediction. However, this breakthrough was only made possible due to well-curated databases which were used to train the AI model. I would like to highlight where this database comes from, and how we can support further discoveries like AlphaFold. To do that, I need to take a step back and introduce the EMBL.
The European Molecular Biology Laboratory is Europe’s only intergovernmental laboratory for life sciences research, supported by 29 member states, of which the UK is one. With sites across Europe, each focused on a different aspect of research, the mission is to advance biomedical research for providing the tools, data and resources that help scientists tackle important biological questions. One of the key areas of that is bioinformatics. This is where the European Bioinformatics Institute, hosted here in the UK, just up the road in Kingston, plays a central role. At the EBI we are dedicated to providing open access data, tools and services that empower life sciences across various fields. Our goal is to ensure that the data and knowledge we provide not only support the scientific discoveries but enable their real-world applications.
Let’s take a look at how this data process is managed through the data cycle. We have scientists around the world who generate data and make their discoveries. They then publish their findings and deposit that data in places like the EBI, just as books are made available through libraries. At the EBI, experts archive and make the data available to all scientists free of charge. Not only do we do that, but the experts add value to the data, as a single data point in itself is not very meaningful. It is only when data are collated, analysed and put into context that true meaning, understanding and insights can emerge. This is why, at the EMBL-EBI, we focus not just on gathering the raw data but on enhancing it and making it available and accessible in ways that actually drive impactful scientific discoveries.
We manage a vast array of biological data resources at the EBI, but today I will focus on one of the 40-plus data resources, the Protein Data Bank in Europe, as I am the co-ordinator there, and its role in driving innovations such as the Nobel prize-winning AlphaFold. The PDBe, which is within the EBI, plays a vital role within the Worldwide Protein Data Bank, a global collaboration dedicated to collecting, validating and distributing the 3D macromolecular structures that I was talking about. Specifically, the PDBe manages deposition and structural data from Europe. It creates it and makes experimentally solved structures accessible to researchers. We then add value to the data by providing user-friendly interfaces that allow them to explore the structures and enable them to better understand and answer their biological questions.
In recent years, the field of structural biology has undergone a remarkable transformation. We have moved from managing a limited number of experimentally determined structures to handling an ever-expanding collection of structural data. This growth has been driven not only by advances in experimental methods such as cryo-electron microscopy, but in the breakthroughs of predictive techniques like the aforementioned AlphaFold. This has significantly increased the amount of data that we have.
The sheer scale of the data now presents both challenges and opportunities. The challenge lies in making sense of such a vast amount of information and ensuring its accessibility to researchers across various fields. In harnessing this data we have the opportunity to accelerate scientific discoveries and innovation by providing open access to accurate, high-quality protein structures in the context of other biological data. An example is AlphaFold.
AlphaFold was revolutionary, but it required specialist knowledge and access to high-performance computing, which initially limited its impact. To ensure that the benefits of this breakthrough were accessible from the very beginning, Google DeepMind collaborated with the PDBe and developed the AlphaFold database, to bring the fruits of the revolution to a broader audience. It now provides over 200 million pre-computed structures for free, as well as the tools and training to investigate and understand those structures. That accelerates research and innovation in drug discovery, biotech and disease understanding, while reducing the environmental impact of high-energy computational resources.
Through our collaboration on the AlphaFold database, PDBe continues to lead the way in making protein data open, accessible and impactful. AlphaFold was only made possible due to the well-curated PDBe database, which was used to train the AI model. Now it returns to the community. The curation was only possible due to the long-term and sustained funding of the PDBe that allows it to collect, archive and enrich the data.
This leads me to how we keep the UK at the forefront of the AI revolution in life sciences and get to the next breakthrough. An important part of it is sustaining long-term innovation. Long-term, stable funding, particularly from the UK Government, is essential for maintaining open access biological data. It is governmental support that enables ambitious, long-term projects and fundamental research that do not have immediate commercial returns. It also allows the leveraging of additional funding sources that amplify the impact of the initial investment. As data volumes grow beyond 100 petabytes, scalable computable infrastructure becomes critical. Investing in global scale infrastructure is far more efficient than supporting numerous smaller disconnected efforts. Open science accelerates innovation by ensuring that high-quality, structured data is widely accessible for both academia and industry. Siloed fragmented databases slow progress, while a co-ordinated approach ensures that data remains reusable and scalable.
Finally, as we reflect on AlphaFold and the AI-driven revolution, we see there is an explosive demand for computational skills. To fully harness this transformation we need investment, not only in data and infrastructure but in people, ensuring that the next generation of scientists and engineers is equipped to lead the way. I thank the PDBe team, and thank you for your attention.
Chair: Thank you very much. That was fascinating. It is good to see the relationship between British investment and Google DeepMind unfolding the secrets of proteins.