Written evidence submitted by the John Innes Centre (BIG0061)
Executive Summary:
Scientific research is greatly benefitting from new technologies that give rise to ‘Big Data’.
Whilst capital investment is required to align the scientific computing infrastructure with the demands of increasing data quantities, we see the bottlenecks in exploiting ‘Big Data’ in the following:
The John Innes Centre
This report is being submitted on behalf of the John Innes Centre (JIC) with input from colleagues across the Norwich Research Park, in particular from the Genome Analysis Centre (TGAC).
The John Innes Centre is an independent, international centre of excellence in plant science and microbiology.
Our mission is:
• To generate knowledge of plants and microbes through innovative research
• To apply our knowledge of nature’s diversity to benefit agriculture, the environment, human health and well-being
• To train scientists for the future
• To engage with policy makers and the public
The research we do makes use of a wide range of disciplines in biological and chemical sciences, including microbiology, cell biology, biochemistry, chemistry, genetics, molecular biology, computational and mathematical biology.
We receive funding from the UK Biotechnology and Biological Sciences Research Council (BBSRC) for research that directly address BBSRC strategic objectives in food security, human health and industrial biotechnology.
‘Big Data’ at the John Innes Centre and the Norwich Research Park
Science is facing a number of ‘Big Data’ related challenges. Biological research in particular has benefitted from step-changing technological developments that are enabling data to be produced at unprecedented rates. In the area of predictive agriculture (and the fundamental plant biology underpinning it), we are starting to gain a systems level understanding of how plants interact with their environment through the emerging fields of crop phenomics, pathogenomics and new sequencing technologies.
Whilst capital investment is required to process, store, transmit, query and backup the data, from our experience we see the major challenges relating to ‘Big Data’ as a problem of addressing the shortage of skilled computational biologists and statisticians.
With the advent of “Big Data” there will be a foreseeable shortfall in digital professionals to meet the demand. One study estimated that 58,000 new jobs will be required by 2017 (CEBR “Unlocking the value of Big Data” http://www.sas.com/offices/europe/uk/downloads/data-equity-cebr.pdf), whereas other estimates see the skills gap to be even wider and suggest that by 2020 a further 300,000 digital professional jobs in the big data field will be required (SAS “Investment needed to meet UK demand for big data skills and analytics”
http://www.sas.com/en_gb/news/press-releases/2014/october/demand-big-data-skills-analytics.html). A similar trend can be expected for scientific research. Clearly, these roles need to be fostered by effective, timely, and relevant education and training to ensure that the next generation of digital professionals and data scientists can meet the market demand.
The goal of collecting data has to be to extract meaning from it. “Big Data” techniques are something of a hot topic in industrial settings, however the implementation of computational approaches for analysing, visualising and, most importantly, interpreting ‘Big Data’ in the biosciences are in their infancy.
To make the most of the investment made to collect the data, the data and the associated software for its analysis needs to be readily available, easy to use and intuitive. The data need to be described in standardised computer-accessible formats, through the addition of metadata via controlled vocabularies, i.e. ontologies. Furthermore the data need to be maintained, updated, and integrated with new data on timescales that go beyond the typical grant funding periods. The storage of data needs to provide long-term access and ideally at a low cost to funding agencies, universities, institutes and scientific publishers.
The development of heuristics to deal with ‘Big Data’ will not follow a one-size-fits-all approach and will need to be tailored for domain specific problems (discipline and application based). This requires computational and statistical expertise within various scientific disciplines. As such, a strong commitment to further training in computational, mathematical and statistical methodologies will be key for the UK to maintain its scientific lead in biological sciences.
The rate of data generation across the Norwich Bioscience Institutes has grown rapidly over recent years. Our site wide storage usage for big datasets is increasing by over 50% annually, with a comparable increase in processing and backup requirements. The scale of computational infrastructure required to support the acquisition, storage and processing of these data is immense. Managing this infrastructure requires highly skilled staff, which presents a serious challenge for recruitment and training.
Whilst researchers may have access to funding streams involving industrial partners, these are often restricted to small-to-medium enterprise (SMEs). If we are to truly realise the data-driven approaches for Big Data interpretation, we need to be able to involve big industrial players (Microsoft, Intel, Google, Apple, etc.), and this should be actively encouraged.
Selected Examples:
Plant pathology in the genomics era: With rapid advances in DNA sequencing technology over the past decade, genomics is now embedded within the toolbox of every 21st century plant pathologist. Scientists can now freely explore plant pathogens at the genomic level, searching for signatures that may convey their ability to cause disease. Beyond single genome analysis, in Norwich we have also developed a genomics-based approach for plant pathogen population surveillance that is dependent on the generation of high-resolution data to track pathogens that pose a significant threat to several major global crops that include wheat and maize. However, the generation of hundreds of large complex genomic datasets (each ~20Gb) from pathogen-infected plant samples poses a significant challenge regarding the development of efficient computational pipelines for their analysis. An additional struggle is to recruit scientists to such projects with the biological background to design meaningful hypothesis-driven experiments, whilst having the required numerical skill base to exploit these datasets to their maximum potential. For scientists to appreciate the panoramic view afforded by the generation of such sizeable and comprehensive datasets it is paramount that we break down the boundaries between research disciplines and utilize these data resources to their full potential in a coordinated manner. As we move forward, with the cost and speed of genome sequencing decreasing rapidly, in Norwich we are embracing a multidisciplinary approach to research. We are actively bringing together scientists from an array of disciplines to work side by side to leverage these technical advances to answer key questions in plant pathology, aimed at ultimately achieving global food security. There is a desperate need to attract specialist data scientists into this area and to tailor computational training specifically for bench-biologists to make this unprecedented wealth of data accessible and truly exploitable by experienced researchers but also non-experts.
Field Phenotyping: Agricultural research programmes are being transformed by the use of novel field phenotyping technologies, which are being exploited to characterize the behavior of crops under field conditions. Quantification of agricultural traits is a key step towards selecting favourable genotypes and improving crop performance. Some recent phenotyping technologies include UAV (unmanned aerial vehicle), remote sensors, robotics and non-invasive imaging devices, which can easily generate terabytes of multi-dimensional phenotypic data in just one field trial. Finding solutions for processing large phenotypic datasets to extract meaningful results has become one of the greatest challenges faced by crop scientists today. ‘Big Data’ produced by high-throughput and high-resolution examinations of crop growth and its interaction with the environment is well beyond comprehension of a single research group or scientists from one research domain. Therefore, a joint effort needs to be made by researchers from different disciplines, including bioimaging informatics, computational biology, mathematical modelling, computer science, crop genetics and plant physiology, to catalyze breakthrough discoveries in crop research in the UK. Key to progress in this upcoming area will be the ability to handle and process ‘Big Data’ and a current hurdle is the recruitment of sufficiently skilled staff.
‘Omics Data Infrastructure: The Genome Analysis Centre houses a large BBSRC-supported National Capability in genomics and scientific computing. This UK-focused provision of sequencing hardware, high-throughput laboratory robotics, and high-performance computing has allowed TGAC to not only serve the sequencing needs of an international community, but also forms the basis of fundamental scientific research into biological systems on the NRP. These large-scale data generative processes require efficient and effective data storage and analysis platforms. Through our relationship with industrial computational vendors, we are able to store and analyse huge quantities of data that underpin world-leading scientific collaborations.
However we are seeing an increase in amounts of available data and a reduction in the time it takes to generate it, far outpacing Moore’s Law. With the advent of the sub-$1000 human genome and a single sequencing experiment that can fill a computer hard drive, we are now faced with very different challenges that were seen even as recently as 5 years ago. Researchers now need to find, share, and interpret large complex datasets, integrate them with their own, and then present their findings for reuse and reproducibility. Enabling access to huge datasets for researchers working in a range of biological disciplines demands investment in both novel hardware and software layers. Only through sustained reliable funding of these e-Infrastructures will we be able to innovate, implement new techniques, and maintain effective globally-connected platforms to help biologists and bioinformaticians continue to answer the key questions of our time.
Recommendations:
September 2015