Global Nr.

Task Nr.

Task

Description

Output

Outcome

Lead Contributors

WP 1

The Data-driven Forecasting Work Package (WP1) focuses on advancing meteorological forecasting using data-driven approaches with traditional numerical weather prediction systems. This WP aims to develop models specifically also for regional forecasting at kilometer-scale resolution. The project will create data-driven limited area models (LAMs), develop global stretched-grid systems, and investigate downscaling techniques to establish a leading forecasting system that provides accurate and reliable forecasts across different spatial and temporal scales. Key activities include ensuring dataset access and sharing, defining evaluation metrics, developing and testing different models, and collaborating on scaling efforts. By the end of the project, WP1 aims to evaluate the strengths and weaknesses of various model approaches and contribute to the advancement of data-driven forecasting technologies, ensuring the delivery of high-quality forecasts for diverse applications and sectors. We expect that multiple project participants will have real-time data-driven forecasts running in parallel with their current operational systems.

The tasks outlined in the following encapsulate a diverse array of activities essential for achieving the project objectives. All Beneficiaries contribute to all tasks of Work Package 1 with some acting as Lead Contributors, as specified for each task.

1

1.1

Evaluation Metrics

Define evaluation metrics and headline variables that will be used to assess the performance of data-driven forecasting models. This task is essential for establishing clear criteria for model evaluation. The task involves a thorough review of existing evaluation metrics used in meteorological forecasting and determining which metrics are most relevant for assessing the performance of data-driven models, especially at the km-scale and for evaluating performance for extremes.

Document outlining the chosen evaluation metrics and headline variables.
(Tentative: Common software for computing a base-set of evaluation metrics either against test dataset or observations.)

Clear criteria established for assessing the performance of data-driven forecasting models, enabling objective evaluation and comparison of different approaches.

Met Éireann, RMI, Météo-France, LVGMC, GeoSphere

2

1.2

Datasets for high-resolution regional analysis data.

Ensure anemoi-datasets can deal with the diversity of data we have in the community and evaluate and extend its design and functionality. In this task we generate anemoi-datasets for various regional and global datasets in order to test and evaluate anemoi-dataset as a common tool that generates datasets that are compatible for all models and architectures employed in the project.

Various plugins of Anemoi for various sources of regional datasets (e.g.
CERRA). Feedback concerning usability and functionality of anemoi-datasets. New release of anemoi-datasets capable of integrating the required datasets.

anemoi-datasets can generate ML-ready datasets for regional datasets of interest in the pilot project, allowing to more easily compare training and inference using different datasets.

KNMI, DMI, MeteoSwiss, AEMET, RMI,
MET Norway, Météo-France,
FMI, GeoSphere

3

1.3

Extension of Anemoi to AICON

Integration of the relevant functionality of the AICON model into the Anemoi

infrastructure such that we can easily evaluate and compare this approach with the other existing model architectures. In particular the task integrates the ICON data grid mesh into the hidden mesh encoder, as well as the nested hidden graph approach of AICON.

Version of Anemoi that can be used to train directly AICON on the ICON native grid and with the nested hidden graph of AICON.

Infrastructure of Anemoi can be used to train AICON on ICON global and LAM datasets.

DWD

4

1.4

Data-driven LAM Model

Build a data-driven limited area model(s) (LAM) using existing exemplar models as references. The aim is to create a model tailored for regional forecasting at kilometer-scale resolution. A special focus will be given to variables important for high-resolution LAM applications (e.g., precip, low cloudiness, wind gusts, …).
Development of LAM model(s) involves selecting appropriate architectures (such as graph neural networks, vision transformers or spectral neural operators) based on models like AIFS and Neural-LAM. The LAM models should integrate boundary conditions by combining a cutout global dataset with the embedded high resolution dataset over the region of interest. Investigation of the effect of boundary zone width and training on different datasets for LAM predictions.
With respect to infrastructure, we integrate important functionality of existing approaches like Neural-LAM into a common framework (Anemoi) such that we can easily compare various approaches. In particular, the task integrates the concept of hierarchical mesh into the graph of Anemoi. Also the LAM mode of Anemoi is consolidated and tested with more than one regional dataset. Sensitivity experiments will be performed to determine
the ideal length of training datasets balancing model accuracy and computational efficiency.

A system for training, validating and testing data-driven LAM models with
various model architectures and proper boundary conditions tested over at least one region.

Successful development of a tailored LAM model capable of producing accurate forecasts at kilometer-scale resolution, demonstrating the potential of data-driven approaches for regional forecasting.

DMI, MeteoSwiss, DWD, RMI, Météo-France, Met Éireann, GeoSphere

5

1.5

Stretched-grid Data-driven Model

Continue the design and implementation of a global stretched-grid system using suitable architectures, such as those used in the AIFS model based on Anemoi. The objective is to develop a model that seamlessly transitions from coarse resolution to high-resolution grids over specific domains of interest. The task requires exploring different approaches for constructing stretched grids and adapting existing models like AIFS to support this configuration.
Approaches like high resolution hidden mesh over the regional domain should be merged into the main framework for use by all institutions on various datasets. It involves creating training datasets for stretched grid models, model training, and testing over at least one region to assess its performance compared to other approaches. Sensitivity experiments will be performed to determine the ideal length of training datasets balancing model accuracy and computational efficiency.

A system for training, validating and testing stretched-grid data-driven models tested over at least two regions.

Creation of a global stretched-grid system using suitable architectures, with improved resolution over specific domains of interest, showcasing the effectiveness of data-driven models for transitioning between different grid resolutions.

MET Norway, KNMI, MeteoSwiss, AEMET,  FMI, , GeoSphere

6

1.6

Downscaling / Superresolution

Downscaling using generative AI is an alternative approach to auto-regressive model emulation, which has the potential to ease the generation of convection resolving ensembles (e.g. SEEDS from Google or CorrDiff from NVIDIA).
Here, the focus is on investigating and implementing downscaling techniques to translate coarse-scale data (such as global data-driven forecasts) to higher spatial resolution for better accuracy at the regional-to-local level, specifically for resolving convection. This task involves researching and implementing machine learning-based downscaling methods to improve the resolution of global forecasts. It includes evaluating the effectiveness of these methods through testing over at least one region and exploring the representation of extremes. The task will benefit from (but not duplicate) the work in Destination Earth project DE_371 as well as other on-going projects.

Downscaling method tested over at least one region, demonstrating improved
resolution.

Successful implementation of downscaling techniques to enhance the resolution of global forecasts for better accuracy at the regional level, contributing to the refinement of data-driven forecasting systems. Potential applications go beyond forecasting and can also be applied to reanalysis datasets.

Météo-France, MeteoSwiss

7

1.7

Intercomparison of Model Approaches

Compare the performance of the different model approaches in Tasks 1.4, 1.5, and 1.6 for at least one geographical region. The aim is to evaluate the relative strengths and weaknesses of each approach. The task involves conducting a comprehensive analysis comparing the performance of different model approaches based on predefined evaluation metrics (established in Task 1.1) and ideally based on a
single dataset.

Comparative analysis highlighting the strengths and weaknesses of each
approach for regional modeling.

Informed decision-making regarding model selection and further development
based on comprehensive evaluation of different model approaches, contributing to
advancements in data-driven forecasting techniques

RMI, MET Norway, KNMI, Météo-France, MeteoSwiss

8

1.8

Multiple encoders architecture

It will be common that the global and regional datasets are defined with
different variables, resolution, vertical coordinates system or different levels. A single encoder configuration would require a coherent definition of variables and vertical levels.
Instead of having to data reshape / interpolate all regional datasets to a form compatible with the global datasets, in this task we implement functionality to support multiple encoders such that the global and regional datasets are processed into the hidden mesh
through different encoders.

Version of Anemoi that can be configured with various encoders for different datasets.

The multiple encoder system is tested and evaluated with at least one configuration of datasets

KNMI, MET Norway, MeteoSwiss, DMI, DWD, DMN, LVGMC, ARSO

9

1.9

Training on CERRA dataset

The CERRA dataset is a pan-European reanalysis with very high
horizontal resolution (5.5 km) forced by the global ERA5 reanalysis. A grid-stretching training of ERA5 with CERRA in the European will provide a pre-trained model of high value for all participants of the pilot project for fine tuned training on high resolution regional datasets. It has been shown that a fine tuning based on a pre-trained ERA5 model improves scores for high-resolution regional datasets that do not contain as much
historical data as ERA5.

A pre-trained model using ERA5 + CERRA available to all members of the
project.

The data will be used for further fine tuning and learn about the benefits of
using a pre-trained model for high-resolution regional fine-tuning.

KNMI, RMI, MET Norway

10

1.10

Transfer Learning

Use the output of Task 1.9 to adapt the pre-trained model and fine-tune to various high-resolution datasets. In this task we will develop methods and approaches to fine-tune on high resolution regional datasets. In particular, we will design approaches to generate fine tune models for hourly trained data from the 3 hourly temporal resolution of
the CERRA dataset.

Documentation of approaches for transfer learning from the pre-trained model to high resolution regional datasets using grid-stretching or LAM models.

The capability to use pre-trained models can reduce training effort or increase
the performance of data-driven models. This task can also be seen as a step towards
fine-tuning future foundation models for regional data-driven forecasting.

MET Norway, AEMET, KNMI, DMN, GeoSphere, LVGMC,

11

1.11

Scaling Efforts

Improve the scalability and performance of data-driven models across
multiple GPUs. Where global, LAM and stretched-grid approaches share a codebase, collective efforts across these approaches will improve the training performance and scale to larger ML models (both in terms of complexity and resolution) across multiple GPUs. It includes optimizing model architectures, algorithms, and distributed computing techniques to improve scalability and performance.

Infrastructure to train and test models with improved scalability and
performance.

Enhanced scalability and performance of data-driven forecasting models
across multiple GPUs will lead to a capability to tackle higher-resolution datasets and
reduce training time, allowing for broader ablation studies.

MET Norway, RMI, SMHI

12

1.12

Hourly / sub-hourly resolution forecasting

Several approaches will be explored:

-          time interpolation

-          new solution relying on multiple decoders to forecast all time steps

Report describing the results of each method, including evaluation of advantages/disadvantages

Increased knowledge and implementation of approaches for increased temporal forecasting resolution

MET Norway, MeteoSwiss, RMI, DMN, LVGMC

WP 2

This work package aims to develop a world-leading regional and global ensemble forecasting system using data-driven models, focusing on building reliable ensembles by addressing initial condition and model uncertainty. Leveraging the low inference cost of data-driven models, it explores enhancing deterministic models to probabilistic ones. Key tasks include utilizing NWP-derived uncertainties, investigating ML-specific error sources and enriching ensembles with generative methods. Model forecast evaluation will emphasize reliability and extreme event detection, with milestones including a literature review, baseline system establishment, error source investigation, verification tests definition and testing ensemble generation methods.

The tasks outlined in the following encapsulate a diverse array of activities essential for achieving the project objectives. All Beneficiaries contribute to all tasks of Work Package 2 with some acting as Lead Contributors, as specified for each task

13

2.1

Establish Evaluation Metrics and Headline Variables (jointly with task 1.1)

The objective is to define the evaluation metrics and headline variables that will be used to assess the performance of probabilistic approaches for data-driven forecasting. Thorough evaluation of these data-driven ensemble systems will be critical, as initial work has highlighted that data-driven models can score well on metrics such as the Continuous Ranked Probability Score, while remaining significantly overconfident.
Evaluation should emphasise the value of data-driven ensembles for the detection of extreme events.

Document outlining the chosen evaluation metrics and headline variables for data-driven ensemble verification.

Clear criteria established for assessing the performance of probabilistic data-driven forecasting models, enabling objective evaluation and comparison of different
approaches.

DMI, GeoSphere

14

2.2

Approaches for Building Reliable Ensembles with Data-driven Models

This task will interact with WP1 to conduct a review to identify optimal approaches for building reliable ensembles with data-driven models. This entails evaluating existing methodologies, focusing on their ability to address initial condition uncertainty and model uncertainty. The most promising approaches are selected for further exploration and adaptation in the subsequent tasks.

A document detailing the evaluation of existing methodologies for building
reliable ensembles with data-driven models.

A first set of approaches for building reliable ensembles, providing a clear direction for subsequent tasks focused on adapting and implementing these methodologies.

DMI, KNMI, DWD, ITAF MET Service, Météo-France

15

2.3

Representing Initial Condition Uncertainty

The aim is to establish a baseline ensemble system for global and local
approaches using initial condition uncertainty only, with deterministic data-driven model (WP1). To this end, initial condition uncertainty will be represented for forecast reliability across all timescales. As an initial approach, the task will reuse initial condition uncertainty from NWP models

A functional baseline ensemble forecasting system for global and local
approaches.

A reliable baseline system for data-driven forecasts across all timescales.

DMI, DWD

16

2.4

Incorporating Model Uncertainty

This task will investigate approaches for incorporating model uncertainty in the baseline ensemble system (Task 2.3). This includes training to minimise probabilistic skill scores, using Diffusion models or Generative Adversarial Network (GAN) approaches. Alternately, ensembles could be bypassed and probability distributions could be directly predicted.

A document outlining the investigated approaches for incorporating model
uncertainty in data-driven models.

Reliable ensemble forecasts through the effective incorporation of model
uncertainty.

KNMI, ITAF MET Service, DMI, Météo-France, LVGMC

17

2.5

ML ensemble diagnostics

Develop a framework for diagnostics that will allow investigation of error sources in ML-based models as these will probably be different from error sources in physics-based NWP models, comparing forecast errors in data-driven and NWP models, studying error growth properties of perturbations.

Relevant diagnostics established for improved
understanding of errors and uncertainties in ML-based ensembles.

A better understanding of uncertainty characteristics of ML-based ensembles
that will help us to develop more reliable data-driven ensemble forecasts.

DWD, ITAF MET Service

18

2.6

Enrichment of the Ensemble Using Generative Methods

This task will investigate a complement to forecast ensemble members,
where the ensemble is enriched from pre-existing members by exploring generative methods (e.g. GANs, Diffusion models). This means extension of these methods to many variables (precipitation, etc) and extreme events are given particular attention.

Integration of generative methods to enrich forecast ensemble members, with a focus on extreme events.

Improved ensemble forecasts, through the successful enrichment of the
ensemble using generative methods.

Météo-France, SMHI

19

2.7

Documentation and Reporting

This task entails documenting all research findings, methodologies and
outcomes in WP2 generated throughout the course of the project. It aims to ensure transparency, reproducibility, and dissemination of project results. It also involves organizing and presenting findings at conferences, workshops, and other relevant forums to engage with the scientific community.

Comprehensive report summarizing the research conducted, including results,
conclusions, and recommendations for further development.

Transparent documentation and dissemination of research findings, methodologies, and outcomes, facilitating knowledge sharing and contributing to the advancement of the scientific community.

Météo-France, KNMI, DWD, DMI, SMHI,
ITAF MET Service

20

2.8

Implementation of diagnostic framework in Anemoi

The diagnostic framework developed under Task 2.5 will be integrated into the Anemoi framework to enable systematic evaluation of machine-learning (ML) ensemble forecasts. This implementation will make diagnostic tools and methods accessible to the broader Anemoi community, supporting reproducibility, transparency, and shared model assessment practices.

Diagnostic framework implemented and available within the Anemoi environment


Documentation and examples demonstrating use of diagnostic tools for ML ensemble evaluation

Enhanced community capacity to evaluate and compare ML ensemble performance through shared, accessible diagnostic tools integrated within the Anemoi framework.

DWD, ITAF MET Service

21

2.9

Comparative studies as a follow-up on Task 2.2

Demonstration of pros and cons, both in terms of meteorological quality and technical performance, of (1) CRPS-based models vs diffusion models, and (2) spectral vs multi-scale loss as a means to reduce noise in predicted fields, e.g. precipitation.

A report documenting pros and cons of (1) CRPS-based vs diffusion models and (2) spectral vs multi-scale losses.

Better understanding and possibly a recommendation for when to use one or the other method in data-driven ensemble forecasting

RMI, MET Norway, MeteoSwiss

22

2.10

Diagnose and improve physical realism and consistency for ensemble forecasts

Even though verification scores may look fine, a closer look reveals that the forecast fields might be unrealistic, and that there is not always consistency between different parameters (e.g. rain but no clouds).  These sorts of problems are more pronounced for stochastic models than for deterministic ones where the RMSE loss smooths the fields and to some extent hides the problems.

Develop and apply diagnostics that quantify physical realism and consistency.

Documentation of physical realism and consistency, and - if relevant - recommendations on how to improve physical realism and consistency.

MET Norway, DWD, RMI, DMI, GeoSphere

23

2.11

Higher resolution in time

Most AIWP models developed in WP1 have been trained using 6-hourly time steps with limited experiments exploring 3-hourly training. Generating reliable hourly forecasts while avoiding error growth remains an open challenge. Met Norway has successfully developed a neural time interpolator, enabling 6-hourly trained deterministic models to produce high-quality hourly outputs. Extending this deterministic approach to ensemble models presents a promising opportunity to generate hourly forecasts with well-calibrated uncertainty estimates.

Adaptation of the deterministic neural time interpolator to ensemble models.

Hourly, well-calibrated data-driven ensemble forecasts.

MET Norway, MeteoSwiss, KNMI

24

2.12

Pretraining methodologies for fine-tuning to high-resolution datasets

Without available high-resolution global datasets spanning multiple decades two pretraining strategies have been investigated in WP1 in order to bridge the gap between global low-resolution and regional high-resolution datasets: (1) pretraining using global ERA5 reanalysis data before fine-tuning it with higher-resolution regional data, and (2) direct training of models from scratch using global ERA5 reanalysis and a long-term, regional dataset (e.g. CERRA covering 37 years at 5.5km spatial resolution). The extension of these two approaches for ensemble models is needed to provide ready-to-fine-tune ensemble models to meteorological institutes but also yield valuable insights into optimal pretraining methodologies for ensemble models.

Pretraining methodologies available for ensemble models

Ready-to-fine-tune ensemble models.

KNMI, RMI

WP 3

The main objective of this work package is to establish cutting-edge data assimilation approaches in order to enable a fully data-driven NWP forecasting process. Data assimilation is crucial in guaranteeing the quality of the numerical weather forecasts and has been a major pillar of their continuous improvements over the last decades. With the advent of AI-based frameworks in NWP, new possibilities for enhancing the performance of data assimilation systems arise while at the same time allowing for increased quality in the data assimilation process. As data-driven NWP forecasts are likely to become a part of the operational NWP chain, the traditional data assimilation systems may become a bottleneck in terms of performance. Therefore, research and development efforts are needed to establish new AI-based data assimilation schemes which will enable a fully data-driven NWP process.

Therefore the following tasks are carried out in the project utilizing new machine learning techniques in order to enhance data assimilation systems and to address the objective of WP3. All Beneficiaries contribute to all tasks of Work Package 3 with some acting as Lead Contributors, as specified for each task

25

3.1

Emulation of the classical variational assimilation scheme and its components

4D-Var is a state-of-the-art assimilation method that is used by many weather centers for research and operational purposes. However, the development and maintenance of 4D-Var components (e.g., linearized model operators) are difficult and require a significant amount of human and machine resources. This task is designed  to build expertise and start constructing a prototype 4D-Var (using the available data-driven model in Anemoi software) in which the tangent-linear and adjoint operators are obtained by automatic differentiation. The quality of those TL/AD operators should be assessed and might be penalized during training. Depending on the quality, the possibility to plug AIFS TL/AD in existing FORTRAN framework (e.g. OOPS) can be envisaged. We also aim to build the necessary functionality in Anemoi software that is generic for all available models. In this task, other hybrid ML-DA approaches are possibly included (e.g., ML-based bias correction). ML methods (e.g. architectures like Convolutional Neural Networks) can be used to emulate the linear and adjoint model of 4D-Var in a computationally efficient way. Initial studies will be conducted to develop a 4D-Var ML Prototype and evaluate the algorithm's performance to the original version (focusing on high spatial resolution).

Available technical implementation (incorporated into Anemoi if successful)

For automatic differentiation, a technical framework that is working with basic conventional observation types (e.g., aircraft observations) and a case study that demonstrates the performance of the obtained system or operators.

DWD, Météo-France, KNMI

26

3. 2

Flow dependent data assimilation using Bris or other Data-driven ensembles

Flow-dependent data assimilation methods, such as Ensemble Variational (EnVar) and Ensemble Kalman Filter-based methods, aim to incorporate evolving error statistics from a numerical model's ensemble forecasts. Unlike static background error covariances used in traditional 3D-Var or 4D-Var, these methods dynamically estimate covariances from an ensemble of model states, allowing them to better represent the current state of the atmosphere and its uncertainties. One of the major limitations has been the relatively small ensemble sizes that could be afforded operationally in NWP. Here we explore the potential of using ensembles from data driven models to increase the ensemble size

Paper on the Bris ensemble and the potential for use in flow dependent DA.  Software for interactive analysis of ensemble covariance. A working EnKF prototype using conventional observations with the Bris ensemble for exploration of new ideas. Implementation of OOPS in the Harmonie Cy49Th2 branch, validation report Implementation and validation of the Envar formulation based on the Meteo France implementation. New cost function objects in OOPS. New saddle point formulation for the linearized inner loop problem. TL and AD code mapping Harmonie model levels to Bris pressure levels.

Analysis of the suitability of the covariance structure of  the Bris ensemble for potential used in flow dependent DA algorithms.  Prototype of simple EnKF for quick exploration of new ideas using conventional observations in the Bris Ensemble. For Envar, implementation of OOPS in Harmonie cy49Th2 branch, validation of OOPS code against MASTERODB code for all  obs types (SYNOP, AIRCRAFT, BOUY, TEMP, GNSS, AMSUA, AMSUB, MWHS ), implementation of 3D Envar in Harmonie based on Météo-France implementation. Development and  implementation of tangent linear and adjoint code of interpolation operators to map Harmonie states to Bris states. New cost function object in OOPS and implementation of a saddle point formulation for the linearized inner loop problem

MET Norway,


27

3.3

AI/ML bias correction

Adaptive bias correction using Variational Bias Correction (VarBC) has long been an integral part of numerical weather prediction (NWP). While effective, VarBC faces limitations in addressing complex, non-linear biases — for example, those linked to cloud and precipitation processes or arising from overlapping observational sources.

This task explores a machine-learning–based approach to replace or augment the current VarBC method. Training datasets will be developed to enable the ML model to learn bias patterns dynamically from multiple observation types and model states.

Training datasets for ML-based bias correction

Evaluation of different ML model architectures and training strategies

Prototype implementation of ML bias correction within the regional LAM 4D-Var system (Q3 2025)

Improved representation and correction of complex, non-linear observation and model biases, leading to enhanced analysis quality and forecast reliability in the regional LAM 4D-Var suite.

KNMI


28

3.4

Direct integration of data assimilation into neural networks

In this task, we seek to integrate the recently published AI-Var approach (https://arxiv.org/abs/2406.00390) into a comprehensive data assimilation framework (AIDA) in order to streamline the development of innovative AI-based weather prediction technologies. A first version will be implemented in a new streamlined framework, that makes it easy for users to utilize the system. The preliminary framework will then be tested through specific data assimilation experiments. These will support gaining knowledge and experience about AI-Var for a subsequent extension of the system towards a completely AI-based data assimilation cycle.

An AI-based data assimilation framework with the capability of processing observational data and producing initial conditions for AI-based NWP models.

The implementation of an AI-based data assimilation approach is expected to yield substantial advancements towards a fully data-driven NWP chain, with the potential to set a new standard for operational weather forecasting and to demonstrate the transformative potential of AI in enhancing weather prediction capabilities.

DWD, ITAF MET Service

29

3.5

Fully data-driven data assimilation cycle

A fully AI-based data assimilation cycle integrates the DA scheme from Task 3.4 with a forecasting model to perform analysis as well as forecasting step using AI-based component with the goal of achieving a seamless, data-driven workflow. The observations are digested by the DA scheme in combination with the background state provided from the model. The system will then create an analysis state from which the next forecast can be issued. Both components are operated by a control script system that ensures a seamless data flow.

A fully AI-driven NWP data assimilation cycle based on a DA scheme, an NWP model and a control script system

A robust, fully AI‐driven data assimilation and forecasting system that delivers high‐accuracy, low‐latency atmospheric analyses and forecasts.

DWD, ARSO

30

3.6

Integration of AI-Var into Anemoi

In this task, we seek to integrate the developed AI-based data assimilation framework AIDA (Task 3.4) into the Anemoi framework in order to streamline the development of innovative AI-based weather prediction technologies with a focus on future operational employment. Subsequently, the Anemoi framework will be extended by a “Data assimilation” as well as a “Observations” module in order to facilitate data assimilation. The AIDA framework will then be modified or extended, building on the existing structures from the Anemoi “Models”, “Training” and “Datasets” modules. By incorporating the AI-Var approach into Anemoi, we expect to extend its capabilities to enable the integration of data assimilation with AI-based forecast models (WP1) and ensemble approaches (WP2).

An AI-based data assimilation module in the Anemoi framework with the capability of processing observational data and producing initial conditions for AI-based NWP models.

The incorporation of the AIDA within the Anemoi framework is expected to yield substantial advancements towards a fully data-driven NWP chain, with the potential to set a new standard for operational  weather forecasting and to demonstrate the transformative potential of AI in enhancing weather prediction capabilities.

DWD, ITAF MET Service, Météo-France

31

3.7

Observation-driven models and multi-encoder-decoder architectures

In this task, the ultimate goal is to exploit observations either alone or in combination with reanalysis or NWP data sources. The method called AI-DOP (https://arxiv.org/abs/2412.15687) is largely considered here and the use of a new multi-source data handler in Anemoi. For the Member States, the adaptation of DOP to LAM or stretched-grid approaches are foreseen using additional local observational datasets. In order to ensure good quality of input datasets, quality control and data filtering are also needed. This exercise requires a similar data screening routine than the traditional data assimilation has without the dependency of NWP model fields (e.g., the background check). Additionally, the encode-decode blocks can be flexible and independent allowing irregular and variable input-output mesh. Preliminary tests with NetAtmo dataset have been carried out, but developments were not merged into Anemoi. In this task, we will closely follow the corresponding developments of multi-source data handlers and we plan to contribute with local adaptations and observations. RADAR observations are one data source that is considered to be tested in this framework. Also it is planned to exploit ML-ready observations that are prepared by E-AI program (namely SEVIRI and OPERA ZARR datasets) when Anemoi multi-encoder-decoder architecture is ready to utilise these observations.

The adaptation of Anemoi DOP for regional purposes; A multi-encoder-decoder framework in Anemoi that can digest various data sources without any pre-processing or projection of the input observations.

The effective use of multi-source observations improving the short-range skill of ML weather predictions.

MET Norway, DWD, MeteoSwiss, SMHI, LVGMC

32

3.8

Densification, multi-objective loss, and nowcasting applications

Alternative ways of obtaining good short-term forecast quality using pre-processed data, e.g. sparse grid for station data, projection on the same grid with spatial interpolation for radar and satellite. Different sources are handled through a multi-objective loss comparing pointwise data with RMSE-type loss, and dense grid data with spatial loss (e.g. Log spectral distance or Fourier Correlation loss). Attempt to produce nowcasts directly from observation, comparison with runs with NWP forecasts as additional input.

A quantification of nowcast quality lost when NWP data is removed from the model input. A SOTA nowcasting model with NWP inputs at least for surface variables (precip, temperature, dewpoint and wind).

Understanding of how GNN can actually learn physics if no physical reasoning is provided (through NWP runs).

MeteoSwiss

WP4                                        

Integrating data-driven models (as a component) into future forecasting systems requires assessing differences between traditional NWP models and ML-based models. NWP models use mathematical formulations and supercomputers, while ML models need large datasets and specialized GPU infrastructure. This shift demands careful integration of different software stacks, services, and infrastructure. Implementing DevOps practices, such as continuous integration and continuous deployment (CI/CD), is crucial for seamless development and updates. This work package will identify and address software and service gaps. It will also establish best practices for shared infrastructure. The work package explores the possibility of establishing joint CI/CD test infrastructure (for example on Atos, as in Destination Earth) integrated with the software repositories (e.g. on GitHub) used for sharing developments within the community. Further, it will showcase how MLOps best practices can be enforced through specialized services and various degrees of automation. Finally, it will engage with other work packages to define their technical requirements so to enable active collaboration on ML development.

The tasks outlined in the following are cross-cutting and vital for the success of WPs 1 to 3. It is therefore expected that most of the project partners will contribute to the tasks of WP4 during the project period with some acting as Lead Contributors, as specified for each task. The initialisation of the tasks will be performed by MeteoSwiss and DMI. Each task in this work package is focused on specific objectives and outcomes, contributing to the overall goal of integrating data-driven models into existing forecasting systems through improved infrastructure and MLOps practices. This structured approach ensures clear deliverables and measurable outcomes, facilitating efficient project management and successful implementation..

33

4.1

Assess MLOps Maturity and Promote Internal Knowledge Sharing

Develop a shared understanding of MLOps maturity and practices across the ML Pilot Project. This task focuses on surveying partners, documenting current practices, and sharing insights internally to foster alignment and convergence toward best practices in operational ML workflows.

Activities:

  • Design and conduct an MLOps survey across all project partners.
  • Maintain a public status page summarizing MLOps tools and practices across member states.
  • Organize a community workshop on MLflow, covering setup, maintenance, and experiment management.

Report on MLOps maturity in the ML Pilot Project, based on survey responses and the public status page.

Conducted MLflow community workshop with recorded materials and shared best-practice documentation.

Improved consistency, transparency, and adoption of MLOps practices within the ML Pilot Project, enabling more efficient and reliable deployment of machine-learning models.

MeteoSwiss, DWD, DMI, UK Met Office

34

4.2

Promote external knowledge sharing and alignment/ interoperabily of approaches

Align the ML Pilot Project’s MLOps strategies with broader European initiatives and facilitate knowledge exchange with the wider community. This task ensures that best practices are shared externally and that the pilot benefits from synergies with other projects

Activities:

●        Coordinate with the E-AI Working Group 3 to share insights and align strategies.
●        Document exchanges, lessons learned, and recommendations for broader dissemination.

Documented exchanges and recommendations developed jointly with E-AI Working Group 3.

Enhanced interoperability and alignment of MLOps practices across European initiatives, promoting community-wide adoption and collaboration in operational ML workflows.

MeteoSwiss, UK Met Office, DWD, DMI

35

4.3

Develop joint ML infrastructure (a.k.a. “Scaling Anemoi”)

This task supports the joint development and maintenance of the infrastructure needed to train and run AI-based forecasting models, centered around the Anemoi framework. The work focuses on strengthening the reliability, interoperability, and usability of shared ML infrastructure across participating organizations.

Activities:

  • Develop and run integration tests across components of the Anemoi ecosystem.
  • Improve DevOps practices, including automation and standardization of package releases, CI/CD pipelines, and development workflows.
  • Develop onboarding resources (demos, tutorials, documentation) for new users.
  • Define and document plugin interfaces and configuration standards to enable single- or multi-organization extensions while preserving interoperability.
  • Publish a roadmap to guide the future development of the Anemoi project.
  • Define and maintain a process for writing and reviewing Anemoi Decision Records (ADRs).

Integration tests and documented results ensuring component compatibility.

Standardized and automated DevOps workflow, including CI/CD pipelines and consistent package release processes.

Onboarding materials (demos, tutorials, documentation) hosted on a shared platform.

Modular system architecture supporting custom plugins, with published interface standards.

Public Anemoi Development Roadmap outlining scope, philosophy, and key development themes.

Repository of Anemoi Decision Records (ADRs) documenting architectural and infrastructure decisions.

Implemented caching option for downloaded integration test data.

A robust, scalable, and interoperable ML infrastructure within the Anemoi framework that enables collaborative development, efficient deployment, and long-term sustainability of AI-based forecasting systems across participating organizations.

MeteoSwiss, RMI, Météo-France, DMI, UK Met Office

36

4.4

Establish Best Practices for Shared Data, Infrastructure, and Services

Develop and disseminate comprehensive best practices for effectively utilizing shared data, infrastructure resources, and services. This includes:
Resource Allocation: Creating guidelines on the optimal allocation of resources such as ECMWF’s Atos supercomputer, the European Weather Cloud, the MLflow server, and the Anemoi Catalogue.
Efficient Usage: Establishing protocols for the efficient use of these shared resources to avoid conflicts and ensure maximum performance.
Collaboration Protocols: Developing collaboration protocols to enhance teamwork and information sharing among different teams and stakeholders.
Documentation and Training: Providing detailed documentation and training materials to ensure all users can effectively follow the established best practices.

Guidelines and templates for the usage of shared infrastructure and service usage.

Optimized utilization of shared resources, fostering improved collaboration,
efficiency, and overall effectiveness within the forecasting community.

MeteoSwiss,  KNMI, UK Met Office, MET Norway

37

4.5

Governance and Lifecycle Support of ML artifacts for Anemoi Versions
Joint task with WP1

Establish a governance framework for maintaining, modernising, and deprecating Anemoi model artifacts — including checkpoints, datasets, configurations, and workflows — to ensure long-term sustainability, reproducibility, and efficient integration within the shared infrastructure and across releases. The initial focus will be on checkpoints and models produced under the Machine Learning Pilot Project, followed by production-ready models, ensuring that all artifacts remain traceable, reproducible, and compatible with supported Anemoi versions.

Activities:

●             Define governance policies for ML artifacts (support levels, maintenance rules, deprecation criteria).

●             Review and catalogue existing MLPP model artifacts and assess their compatibility with current Anemoi releases.

●             Update Anemoi codebase to enable version-controlled interfaces for loading, saving, and managing artifacts.

●             Implement CI/CD validation pipelines to automatically check artifact compatibility and lifecycle status.

●             [Stretch goal] Develop monitoring and alerting tools for detecting data drift in maintained datasets and models.

Anemoi Lifecycle Governance Document detailing definitions of support, maintenance, and deprecation rules.

Updated Anemoi codebase with enforced lifecycle and compatibility interfaces.

Maintained and validated artifact repository (datasets, checkpoints, configurations) with explicit version tags.

Automated CI/CD checks for lifecycle and compatibility enforcement.

[Stretch goal] Prototype monitoring and alerting tools for data drift detection.

A governed, reproducible, and version-controlled ecosystem of ML artifacts within Anemoi, ensuring sustainability, interoperability, and operational efficiency across releases and participating partners.

MeteoSwiss,  Météo-France, MET Norway, UK Met Office

38

4.6

Development of the “Anemoi Platform” (Anemoi Catalogue and Model Zoo)

Develop and self-host the Anemoi Catalogue as a community-driven “model zoo” and registry for shared ML models, metadata, and related resources.

Operational “Anemoi Platform” (Catalogue + Model Zoo)

Contribution and governance framework

Technical documentation and API references

Demonstration use cases across partners

An open and collaborative Anemoi Platform enabling transparent sharing, discovery, and reuse of ML models and resources, fostering community engagement and accelerating adoption of AI-based forecasting across partners.

AEMET, FMI, MeteoSwiss, UK Met Office

39

4.7

Showcase ML Pipelines from the ML Pilot Project

This task highlights practical experiences from member states in operationalizing machine-learning (ML) models, with the goal of fostering shared learning, transparency, and reuse of good practices across the community. By documenting and communicating lessons learned, the task helps lower barriers for future operational ML implementations and strengthens collaboration within and beyond the project consortium.

Activities:

●        Engage with project partners to identify and document ML pipelines and their automation or operationalization workflows.

●        Collect and curate lessons learned, including technical challenges, solutions, and success factors.

●        Draft and publish blog-style technical articles to communicate insights and promote knowledge sharing beyond the consortium.

Collection of documented ML pipeline and automation case studies from member states (internal report or shared repository).

One or more published blog posts or technical articles (e.g., AIFS blog, Medium, Substack) highlighting engineering insights and lessons from operational ML deployments.

Increased transparency, knowledge sharing, and cross-institutional learning on the operationalization of ML forecasting models, supporting broader adoption of effective and reproducible practices across the community.

MET Norway, UK Met Office, MeteoSwiss

WP5 - Training and support

Training and support on Machine Learning (ML) is clearly needed across the meteorological community to understand everything from the basics of ML in Numerical Weather Prediction (NWP), to how to utilise ML forecasts when providing forecasts and warnings to the public, to how ML NWP models may impact forecasting in the future. There are many general courses on machine learning available therefore it is proposed that ML training from ECMWF and the EMI focuses on the use of ML in the domain of Earth Sciences particularly meteorology with potential expansion to other related areas e.g. climatology, hydrology and air quality in the future.

Training content will largely cover topics of the other work packages of this project, but may also be expanded based on needs. Possible training formats cover webinars, Jupyter Notebooks, e-learning modules, in-person multi-day courses, and documentation linked to results of the other work packages of this project.

Work package 5, as a cross-cutting effort through all other work packages of the pilot project, will be coordinated by the project leads in close collaboration with ECMWF.

  • No labels