Dr Peter Oliver, Director of STFC Scientific Computing, shares a glimpse into the innovative work of the international centre of excellence for advanced computing expertise and digital research infrastructure
For more than 50 years, the Science and Technology Facilities Council (STFC) Scientific Computing centre has provided a foundation for research partnerships, building pioneering research communities, and collaborating on a global scale with researchers and industries.
STFC Scientific Computing has a multidisciplinary, collaborative approach by combining scientific domain expertise, software engineering, platform development, and systems infrastructure to address complex research and development challenges.
Scientific discovery is increasingly limited not by instruments alone, but by the ability to move, manage, analyse, and interpret data at scale. STFC Scientific Computing sits at that junction of linking world-class facilities, advanced computing, artificial intelligence (AI), trusted data infrastructure, and international research communities.
Its focus is on the data and compute challenges that arise from large scientific facilities, such as those at the STFC Rutherford Appleton Laboratory (RAL) – the ISIS Neutron and Muon Source (ISIS), the Central Laser Facility (CLF), Diamond Light Source (Diamond) – and international facilities such as the Large Hadron Collider (LHC) at CERN. A key feature of these experimental facilities is that continuous data output is increasing exponentially with an associated increase in the complexity of data analysis.
To learn more about the work of STFC Scientific Computing, The Innovation Platform spoke with its Director, Dr Peter Oliver.
How have the capabilities and ambitions of STFC Scientific Computing changed in recent years, in line with the rapidly accelerating high-performance computing landscape?
The large-scale, multidisciplinary facilities (MDFs) generate vast amounts of data, and researchers increasingly need to be able to analyse the results of their experiments almost in real time. Because of this, high-performance computing and data systems now need to support both traditional modelling and simulation, and newer AI-driven methods.
AI and machine learning are rapidly becoming central to scientific discovery, so we have developed our AI for Science capability to address specific challenges that these scientific facilities encounter.
For instance, our Ada Lovelace Centre (ALC) is bringing together experts from across our programmes and themes to develop autonomous laboratories and services for facilities like Diamond and CLF. AI is used to change how experiments are performed and this not only makes things easier for the researchers but also makes the science smarter and more reliable.
What are the key priorities of STFC Scientific Computing at present?
Our key focus is on the data – we have a fantastic capability for managing very large volumes of data and compute at scale to help the MDFs with their experiments.
As well as developing automated processes to take the raw data from a sample to processed data and analysis by scientists, we provide ‘Data Analysis as a Service’ and a computing platform known as Ada. This provides users with remote data analysis, bringing together software, data, and compute resources for a variety of workflows. Through this service, we support thousands of researchers from around the world who use our Ada platform to perform their experiment analyses from ISIS and CLF without needing specialist computing knowledge or training. At any given moment, we can expect to see between 50 and 100 users accessing their data on Ada for immediate analysis. Prior to this service and platform, users would have had to download their data to a portable storage disk and take it with them.
Through the ALC, we are ensuring that data from experiments is AI ready. We need huge amounts of data to train new AI models and the experiments being carried out turn the MDFs into data-generating factories. Therefore, we can create a framework for the data and get insights on how to create new processes for those areas where there are knowledge gaps – such as the causes of battery degradation, for example.
We are using AI to enhance scientific discovery – such as developing models for interatomic potentials for materials research. One such AI model is MACE-MP that can predict how atoms behave in different materials, and we are one of the lead partners in this project. MACE can accurately simulate the behaviour of atoms in solids, gases, liquids, and chemical reactions. And it is a model that works across many different types of chemistry and materials science, so you don’t need a separate model for each science area. It makes advanced scientific simulations faster, cheaper, and more accessible for researchers.
Through our Computational Science Centre for Research Communities (CoSeC), we are engaging with a diverse range of research communities across the UK and further afield. A vital element of CoSeC’s work is providing continuity and longevity for the development, support, and maintenance of research software infrastructure. Some of the Collaborative Computational Projects it supports have been providing research communities with key software, delivered to users as supported packages and platforms to enable their research outputs, for several decades. This ensures that the research software is protected and continues to be developed long after the initial grant has finished.
A very important priority for the department is developing the skills and capabilities of our staff, and especially those who are early in their careers. We have very active Apprentice and Graduate programmes, and many of the staff who join us through these are able to continue their careers as full-time members of our research and operations themes.
One of the ways we encourage staff to gain new skills and engage with other like-minded groups and individuals is through attendance at relevant technical conferences and workshops. Additionally, we organise STFC’s ‘Computing Insight UK’ – now the prime UK conference for HPC and related research, which is held in Manchester each December and brings together academic and industry researchers, early career scientists, and HPC infrastructure providers.
Can you explain more about the intricacies and logistics behind the department? For example, how do you receive and store such large amounts of data?
STFC Scientific Computing is extremely complex.
Our staff are based at two sites – the Daresbury Laboratory in Cheshire, and the Rutherford Appleton Laboratory in Oxfordshire – and they have a broad range of in-depth expertise which is embodied in our themes:
- Computational materials;
- Computational life sciences;
- Computational engineering;
- Computational mathematics;
- AI for science;
- Infrastructure;
- Data engineering;
- Software engineering;
- Open science;
- Cyber security; and
- Platforms and services.
These are all supported by our Business Operations group, and each theme can be called upon to provide expertise wherever it is required. Programmes within Scientific Computing use this matrixed theme approach to ensure they have the capability to deliver scientific and technical solutions.
The beauty in this model is that it’s flexible and can respond to new programmes, as well as developing new expertise within the themes. We currently have six key programmes:
- Ada Lovelace Centre (ALC), maximising the scientific impact from the large MDFs;
- Computational Science Centre for Research Communities (CoSeC), enabling collaborative computational research by developing software as an infrastructure. This programme is funded through the Engineering and Physical Sciences Research Council and UKRI Digital Research Infrastructure (DRI);
- Digital Research Infrastructure for STFC (IRIS), a cooperative community creating digital research infrastructure to support STFC science, facilities and beyond;
- Physical Sciences Data Infrastructure (PSDI), which is building and sharing data collections. Funded through UKRI-DRI;
- Data and Analytics Platform for National Infrastructure (DAFNI), providing a computing platform for research into decision-making for national infrastructure. Funded through UKRI-DRI; and
- Square Kilometre Array (SKA) Regional Centre. Funded through STFC PPAN.
STFC Scientific Computing manages over half an exabyte (500 petabytes) of scientific data, making it one of the largest scientific data infrastructures in the UK.
As well as the continuous streams of data coming from the MDFs, we also manage data from international science collaborations, including the Large Hadron Collider (LHC) at CERN, Earth -observation satellites, and space missions. On tape, we store over 400 petabytes (PB) of data with a further 190PB on disk.
We operate a high-performance science network with an aggregate internal switching capacity of more than 30 terabits per second, connecting instruments, storage platforms, and compute facilities. This allows scientists to analyse data where it resides, avoiding the impractical task of repeatedly moving ever larger datasets between locations.
The scale of data movement is equally remarkable. In a typical month, data entering the RAL site averages around 600 gigabits per second (Gbps), of which approximately 200 Gbps originates from the LHC at CERN.
Similar large-scale data challenges arise across many scientific domains. Scientific Computing supports the European Space Agency’s Gaia mission, which has revolutionised our understanding of the Milky Way by precisely measuring the positions, motions, and characteristics of more than one billion stars, with approximately 2PB of Gaia data hosted at RAL to support UK astronomy research.
STFC Scientific Computing’s DAFNI Programme combines large datasets, advanced modelling, and high-performance computing to help researchers, policymakers, and industry understand the resilience of transport, energy, water and communications systems in the face of challenges such as climate change, extreme weather, population growth, and the transition to net-zero emissions.
How important is international collaboration in your work?
International collaboration is central to STFC Scientific Computing’s work. The department provides critical data and computing infrastructure for global science programmes, including CERN’s Large Hadron Collider, the Square Kilometre Array, European Space Agency missions such as Gaia, and international Earth observation and climate research through JASMIN. It also works with partners across Europe, the US, and India on areas such as AI for Science, advanced materials, batteries, aerospace modelling, and research software.
JASMIN is a globally unique data analysis facility, located within the STFC data centre. Can you elaborate on the running and workings of this facility and what it enables?
JASMIN is the UK’s leading environmental data analysis platform and is jointly managed by STFC Scientific computing and the Centre for Environmental Data Analysis (CEDA), part of RAL Space. It is funded by the Natural Environment Research Council, which has recently invested £3m to renew aging infrastructure and replace virtualisation hardware and software. This will enable improvements, broaden the use of JASMIN, and integrate it with other DRI.
The facility was designed and built to meet the collaborative needs of the users, and it exemplifies the approach for scientists to analyse data where it resides.
It combines large-scale storage, compute, and networking to enable researchers to analyse enormous Earth observation and climate datasets. These data are central to research into climate change, helping scientists monitor changes in the atmosphere, oceans, ice sheets and land surfaces, and improve predictions of the impacts of global warming. For example, scientists are currently using JASMIN to understand the role of air quality improvements in increases in extreme heat.
JASMIN currently has:
• Over 2,500 users among over 500 tenant projects across environmental science domains;
• A batch computing cluster offering approximately 55,000 cores for a diverse set of user workloads;
• Approximately 90PB of high-performance disk storage as collaborative space for projects and for the CEDA Archive;
• Approximately 100PB of tape storage for long-term offline archive storage of data;
• An on-premise cloud enabling tenant projects to construct bespoke computing environments with cloud-native technologies; and
• A ‘data transfer zone’, optimising science data transfers with multiple 100Gbps connections to JASMIN’s high-performance storage and 400Gbps site connectivity outward to the outside internet.
What do you think or hope that the near future holds for STFC’s HPC offering?
For UK research capability, STFC is in a unique position. Via the UK National Laboratories, of which Scientific Computing is a part, STFC can play an important leadership role in the delivery of scientific compute and data needs of UKRI and the Department for Business, Innovation, Science, and Trade.
STFC Scientific Computing is helping to deliver greater impact by providing insights into some of the most pressing scientific and societal problems we face today; as well as orchestrating large datasets so they can be used efficiently and reliably; and connecting academic researchers, industries and data, across the UK and around the world.
By bringing together HPC, data and AI, STFC aims to maximise scientific, economic, and societal impact.
Please note, this article will also appear in the 27th edition of our quarterly publication.