Global Nr. | Task Nr. | Task | Description | Output | Outcome | Lead Contributors |
WP 1 The Data-driven Forecasting Work Package (WP1) focuses on advancing meteorological forecasting using data-driven approaches with traditional numerical weather prediction systems. This WP aims to develop models specifically also for regional forecasting at kilometer-scale resolution. The project will create data-driven limited area models (LAMs), develop global stretched-grid systems, and investigate downscaling techniques to establish a leading forecasting system that provides accurate and reliable forecasts across different spatial and temporal scales. Key activities include ensuring dataset access and sharing, defining evaluation metrics, developing and testing different models, and collaborating on scaling efforts. By the end of the project, WP1 aims to evaluate the strengths and weaknesses of various model approaches and contribute to the advancement of data-driven forecasting technologies, ensuring the delivery of high-quality forecasts for diverse applications and sectors. We expect that multiple project participants will have real-time data-driven forecasts running in parallel with their current operational systems. The tasks outlined in the following encapsulate a diverse array of activities essential for achieving the project objectives. All Beneficiaries contribute to all tasks of Work Package 1 with some acting as Lead Contributors, as specified for each task. | ||||||
1 | 1.1 | Evaluation Metrics | Define evaluation metrics and headline variables that will be used to assess the performance of data-driven forecasting models. This task is essential for establishing clear criteria for model evaluation. The task involves a thorough review of existing evaluation metrics used in meteorological forecasting and determining which metrics are most relevant for assessing the performance of data-driven models, especially at the km-scale and for evaluating performance for extremes. | Document outlining the chosen evaluation metrics and headline variables. | Clear criteria established for assessing the performance of data-driven forecasting models, enabling objective evaluation and comparison of different approaches. | Met Éireann, RMI, Météo-France, LVGMC, GeoSphere |
2 | 1.2 | Datasets for high-resolution regional analysis data. | Ensure anemoi-datasets can deal with the diversity of data we have in the community and evaluate and extend its design and functionality. In this task we generate anemoi-datasets for various regional and global datasets in order to test and evaluate anemoi-dataset as a common tool that generates datasets that are compatible for all models and architectures employed in the project. | Various plugins of Anemoi for various sources of regional datasets (e.g. | anemoi-datasets can generate ML-ready datasets for regional datasets of interest in the pilot project, allowing to more easily compare training and inference using different datasets. | KNMI, DMI, MeteoSwiss, AEMET, RMI, |
3 | 1.3 | Extension of Anemoi to AICON | Integration of the relevant functionality of the AICON model into the Anemoi infrastructure such that we can easily evaluate and compare this approach with the other existing model architectures. In particular the task integrates the ICON data grid mesh into the hidden mesh encoder, as well as the nested hidden graph approach of AICON. | Version of Anemoi that can be used to train directly AICON on the ICON native grid and with the nested hidden graph of AICON. | Infrastructure of Anemoi can be used to train AICON on ICON global and LAM datasets. | DWD |
4 | 1.4 | Data-driven LAM Model | Build a data-driven limited area model(s) (LAM) using existing exemplar models as references. The aim is to create a model tailored for regional forecasting at kilometer-scale resolution. A special focus will be given to variables important for high-resolution LAM applications (e.g., precip, low cloudiness, wind gusts, …). | A system for training, validating and testing data-driven LAM models with | Successful development of a tailored LAM model capable of producing accurate forecasts at kilometer-scale resolution, demonstrating the potential of data-driven approaches for regional forecasting. | DMI, MeteoSwiss, DWD, RMI, Météo-France, Met Éireann, GeoSphere |
5 | 1.5 | Stretched-grid Data-driven Model | Continue the design and implementation of a global stretched-grid system using suitable architectures, such as those used in the AIFS model based on Anemoi. The objective is to develop a model that seamlessly transitions from coarse resolution to high-resolution grids over specific domains of interest. The task requires exploring different approaches for constructing stretched grids and adapting existing models like AIFS to support this configuration. | A system for training, validating and testing stretched-grid data-driven models tested over at least two regions. | Creation of a global stretched-grid system using suitable architectures, with improved resolution over specific domains of interest, showcasing the effectiveness of data-driven models for transitioning between different grid resolutions. | MET Norway, KNMI, MeteoSwiss, AEMET, FMI, , GeoSphere |
6 | 1.6 | Downscaling / Superresolution | Downscaling using generative AI is an alternative approach to auto-regressive model emulation, which has the potential to ease the generation of convection resolving ensembles (e.g. SEEDS from Google or CorrDiff from NVIDIA). | Downscaling method tested over at least one region, demonstrating improved | Successful implementation of downscaling techniques to enhance the resolution of global forecasts for better accuracy at the regional level, contributing to the refinement of data-driven forecasting systems. Potential applications go beyond forecasting and can also be applied to reanalysis datasets. | Météo-France, MeteoSwiss |
7 | 1.7 | Intercomparison of Model Approaches | Compare the performance of the different model approaches in Tasks 1.4, 1.5, and 1.6 for at least one geographical region. The aim is to evaluate the relative strengths and weaknesses of each approach. The task involves conducting a comprehensive analysis comparing the performance of different model approaches based on predefined evaluation metrics (established in Task 1.1) and ideally based on a | Comparative analysis highlighting the strengths and weaknesses of each | Informed decision-making regarding model selection and further development | RMI, MET Norway, KNMI, Météo-France, MeteoSwiss |
8 | 1.8 | Multiple encoders architecture | It will be common that the global and regional datasets are defined with | Version of Anemoi that can be configured with various encoders for different datasets. | The multiple encoder system is tested and evaluated with at least one configuration of datasets | KNMI, MET Norway, MeteoSwiss, DMI, DWD, DMN, LVGMC, ARSO |
9 | 1.9 | Training on CERRA dataset | The CERRA dataset is a pan-European reanalysis with very high | A pre-trained model using ERA5 + CERRA available to all members of the | The data will be used for further fine tuning and learn about the benefits of | KNMI, RMI, MET Norway |
10 | 1.10 | Transfer Learning | Use the output of Task 1.9 to adapt the pre-trained model and fine-tune to various high-resolution datasets. In this task we will develop methods and approaches to fine-tune on high resolution regional datasets. In particular, we will design approaches to generate fine tune models for hourly trained data from the 3 hourly temporal resolution of | Documentation of approaches for transfer learning from the pre-trained model to high resolution regional datasets using grid-stretching or LAM models. | The capability to use pre-trained models can reduce training effort or increase | MET Norway, AEMET, KNMI, DMN, GeoSphere, LVGMC, |
11 | 1.11 | Scaling Efforts | Improve the scalability and performance of data-driven models across | Infrastructure to train and test models with improved scalability and | Enhanced scalability and performance of data-driven forecasting models | MET Norway, RMI, SMHI |
12 | 1.12 | Hourly / sub-hourly resolution forecasting | Several approaches will be explored: - time interpolation - new solution relying on multiple decoders to forecast all time steps | Report describing the results of each method, including evaluation of advantages/disadvantages | Increased knowledge and implementation of approaches for increased temporal forecasting resolution | MET Norway, MeteoSwiss, RMI, DMN, LVGMC |
WP 2 This work package aims to develop a world-leading regional and global ensemble forecasting system using data-driven models, focusing on building reliable ensembles by addressing initial condition and model uncertainty. Leveraging the low inference cost of data-driven models, it explores enhancing deterministic models to probabilistic ones. Key tasks include utilizing NWP-derived uncertainties, investigating ML-specific error sources and enriching ensembles with generative methods. Model forecast evaluation will emphasize reliability and extreme event detection, with milestones including a literature review, baseline system establishment, error source investigation, verification tests definition and testing ensemble generation methods. The tasks outlined in the following encapsulate a diverse array of activities essential for achieving the project objectives. All Beneficiaries contribute to all tasks of Work Package 2 with some acting as Lead Contributors, as specified for each task | ||||||
13 | 2.1 | Establish Evaluation Metrics and Headline Variables (jointly with task 1.1) | The objective is to define the evaluation metrics and headline variables that will be used to assess the performance of probabilistic approaches for data-driven forecasting. Thorough evaluation of these data-driven ensemble systems will be critical, as initial work has highlighted that data-driven models can score well on metrics such as the Continuous Ranked Probability Score, while remaining significantly overconfident. | Document outlining the chosen evaluation metrics and headline variables for data-driven ensemble verification. | Clear criteria established for assessing the performance of probabilistic data-driven forecasting models, enabling objective evaluation and comparison of different | DMI, GeoSphere |
14 | 2.2 | Approaches for Building Reliable Ensembles with Data-driven Models | This task will interact with WP1 to conduct a review to identify optimal approaches for building reliable ensembles with data-driven models. This entails evaluating existing methodologies, focusing on their ability to address initial condition uncertainty and model uncertainty. The most promising approaches are selected for further exploration and adaptation in the subsequent tasks. | A document detailing the evaluation of existing methodologies for building | A first set of approaches for building reliable ensembles, providing a clear direction for subsequent tasks focused on adapting and implementing these methodologies. | DMI, KNMI, DWD, ITAF MET Service, Météo-France |
15 | 2.3 | Representing Initial Condition Uncertainty | The aim is to establish a baseline ensemble system for global and local | A functional baseline ensemble forecasting system for global and local | A reliable baseline system for data-driven forecasts across all timescales. | DMI, DWD |
16 | 2.4 | Incorporating Model Uncertainty | This task will investigate approaches for incorporating model uncertainty in the baseline ensemble system (Task 2.3). This includes training to minimise probabilistic skill scores, using Diffusion models or Generative Adversarial Network (GAN) approaches. Alternately, ensembles could be bypassed and probability distributions could be directly predicted. | A document outlining the investigated approaches for incorporating model | Reliable ensemble forecasts through the effective incorporation of model | KNMI, ITAF MET Service, DMI, Météo-France, LVGMC |
17 | 2.5 | ML ensemble diagnostics | Develop a framework for diagnostics that will allow investigation of error sources in ML-based models as these will probably be different from error sources in physics-based NWP models, comparing forecast errors in data-driven and NWP models, studying error growth properties of perturbations. | Relevant diagnostics established for improved | A better understanding of uncertainty characteristics of ML-based ensembles | DWD, ITAF MET Service |
18 | 2.6 | Enrichment of the Ensemble Using Generative Methods | This task will investigate a complement to forecast ensemble members, | Integration of generative methods to enrich forecast ensemble members, with a focus on extreme events. | Improved ensemble forecasts, through the successful enrichment of the | Météo-France, SMHI |
19 | 2.7 | Documentation and Reporting | This task entails documenting all research findings, methodologies and | Comprehensive report summarizing the research conducted, including results, | Transparent documentation and dissemination of research findings, methodologies, and outcomes, facilitating knowledge sharing and contributing to the advancement of the scientific community. | Météo-France, KNMI, DWD, DMI, SMHI, |
20 | 2.8 | Implementation of diagnostic framework in Anemoi | The diagnostic framework developed under Task 2.5 will be integrated into the Anemoi framework to enable systematic evaluation of machine-learning (ML) ensemble forecasts. This implementation will make diagnostic tools and methods accessible to the broader Anemoi community, supporting reproducibility, transparency, and shared model assessment practices. | Diagnostic framework implemented and available within the Anemoi environment
| Enhanced community capacity to evaluate and compare ML ensemble performance through shared, accessible diagnostic tools integrated within the Anemoi framework. | DWD, ITAF MET Service |
21 | 2.9 | Comparative studies as a follow-up on Task 2.2 | Demonstration of pros and cons, both in terms of meteorological quality and technical performance, of (1) CRPS-based models vs diffusion models, and (2) spectral vs multi-scale loss as a means to reduce noise in predicted fields, e.g. precipitation. | A report documenting pros and cons of (1) CRPS-based vs diffusion models and (2) spectral vs multi-scale losses. | Better understanding and possibly a recommendation for when to use one or the other method in data-driven ensemble forecasting | RMI, MET Norway, MeteoSwiss |
22 | 2.10 | Diagnose and improve physical realism and consistency for ensemble forecasts | Even though verification scores may look fine, a closer look reveals that the forecast fields might be unrealistic, and that there is not always consistency between different parameters (e.g. rain but no clouds). These sorts of problems are more pronounced for stochastic models than for deterministic ones where the RMSE loss smooths the fields and to some extent hides the problems. | Develop and apply diagnostics that quantify physical realism and consistency. | Documentation of physical realism and consistency, and - if relevant - recommendations on how to improve physical realism and consistency. | MET Norway, DWD, RMI, DMI, GeoSphere |
23 | 2.11 | Higher resolution in time | Most AIWP models developed in WP1 have been trained using 6-hourly time steps with limited experiments exploring 3-hourly training. Generating reliable hourly forecasts while avoiding error growth remains an open challenge. Met Norway has successfully developed a neural time interpolator, enabling 6-hourly trained deterministic models to produce high-quality hourly outputs. Extending this deterministic approach to ensemble models presents a promising opportunity to generate hourly forecasts with well-calibrated uncertainty estimates. | Adaptation of the deterministic neural time interpolator to ensemble models. | Hourly, well-calibrated data-driven ensemble forecasts. | MET Norway, MeteoSwiss, KNMI |
24 | 2.12 | Pretraining methodologies for fine-tuning to high-resolution datasets | Without available high-resolution global datasets spanning multiple decades two pretraining strategies have been investigated in WP1 in order to bridge the gap between global low-resolution and regional high-resolution datasets: (1) pretraining using global ERA5 reanalysis data before fine-tuning it with higher-resolution regional data, and (2) direct training of models from scratch using global ERA5 reanalysis and a long-term, regional dataset (e.g. CERRA covering 37 years at 5.5km spatial resolution). The extension of these two approaches for ensemble models is needed to provide ready-to-fine-tune ensemble models to meteorological institutes but also yield valuable insights into optimal pretraining methodologies for ensemble models. | Pretraining methodologies available for ensemble models | Ready-to-fine-tune ensemble models. | KNMI, RMI |
WP 3 The main objective of this work package is to establish cutting-edge data assimilation approaches in order to enable a fully data-driven NWP forecasting process. Data assimilation is crucial in guaranteeing the quality of the numerical weather forecasts and has been a major pillar of their continuous improvements over the last decades. With the advent of AI-based frameworks in NWP, new possibilities for enhancing the performance of data assimilation systems arise while at the same time allowing for increased quality in the data assimilation process. As data-driven NWP forecasts are likely to become a part of the operational NWP chain, the traditional data assimilation systems may become a bottleneck in terms of performance. Therefore, research and development efforts are needed to establish new AI-based data assimilation schemes which will enable a fully data-driven NWP process. Therefore the following tasks are carried out in the project utilizing new machine learning techniques in order to enhance data assimilation systems and to address the objective of WP3. All Beneficiaries contribute to all tasks of Work Package 3 with some acting as Lead Contributors, as specified for each task | ||||||
25 | 3.1 | Emulation of the classical variational assimilation scheme and its components | 4D-Var is a state-of-the-art assimilation method that is used by many weather centers for research and operational purposes. However, the development and maintenance of 4D-Var components (e.g., linearized model operators) are difficult and require a significant amount of human and machine resources. This task is designed to build expertise and start constructing a prototype 4D-Var (using the available data-driven model in Anemoi software) in which the tangent-linear and adjoint operators are obtained by automatic differentiation. The quality of those TL/AD operators should be assessed and might be penalized during training. Depending on the quality, the possibility to plug AIFS TL/AD in existing FORTRAN framework (e.g. OOPS) can be envisaged. We also aim to build the necessary functionality in Anemoi software that is generic for all available models. In this task, other hybrid ML-DA approaches are possibly included (e.g., ML-based bias correction). ML methods (e.g. architectures like Convolutional Neural Networks) can be used to emulate the linear and adjoint model of 4D-Var in a computationally efficient way. Initial studies will be conducted to develop a 4D-Var ML Prototype and evaluate the algorithm's performance to the original version (focusing on high spatial resolution). | Available technical implementation (incorporated into Anemoi if successful) | For automatic differentiation, a technical framework that is working with basic conventional observation types (e.g., aircraft observations) and a case study that demonstrates the performance of the obtained system or operators. | DWD, Météo-France, KNMI |
26 | 3. 2 | Flow dependent data assimilation using Bris or other Data-driven ensembles | Flow-dependent data assimilation methods, such as Ensemble Variational (EnVar) and Ensemble Kalman Filter-based methods, aim to incorporate evolving error statistics from a numerical model's ensemble forecasts. Unlike static background error covariances used in traditional 3D-Var or 4D-Var, these methods dynamically estimate covariances from an ensemble of model states, allowing them to better represent the current state of the atmosphere and its uncertainties. One of the major limitations has been the relatively small ensemble sizes that could be afforded operationally in NWP. Here we explore the potential of using ensembles from data driven models to increase the ensemble size | Paper on the Bris ensemble and the potential for use in flow dependent DA. Software for interactive analysis of ensemble covariance. A working EnKF prototype using conventional observations with the Bris ensemble for exploration of new ideas. Implementation of OOPS in the Harmonie Cy49Th2 branch, validation report Implementation and validation of the Envar formulation based on the Meteo France implementation. New cost function objects in OOPS. New saddle point formulation for the linearized inner loop problem. TL and AD code mapping Harmonie model levels to Bris pressure levels. | Analysis of the suitability of the covariance structure of the Bris ensemble for potential used in flow dependent DA algorithms. Prototype of simple EnKF for quick exploration of new ideas using conventional observations in the Bris Ensemble. For Envar, implementation of OOPS in Harmonie cy49Th2 branch, validation of OOPS code against MASTERODB code for all obs types (SYNOP, AIRCRAFT, BOUY, TEMP, GNSS, AMSUA, AMSUB, MWHS ), implementation of 3D Envar in Harmonie based on Météo-France implementation. Development and implementation of tangent linear and adjoint code of interpolation operators to map Harmonie states to Bris states. New cost function object in OOPS and implementation of a saddle point formulation for the linearized inner loop problem | MET Norway, |
27 | 3.3 | AI/ML bias correction | Adaptive bias correction using Variational Bias Correction (VarBC) has long been an integral part of numerical weather prediction (NWP). While effective, VarBC faces limitations in addressing complex, non-linear biases — for example, those linked to cloud and precipitation processes or arising from overlapping observational sources. This task explores a machine-learning–based approach to replace or augment the current VarBC method. Training datasets will be developed to enable the ML model to learn bias patterns dynamically from multiple observation types and model states. | Training datasets for ML-based bias correction Evaluation of different ML model architectures and training strategies Prototype implementation of ML bias correction within the regional LAM 4D-Var system (Q3 2025) | Improved representation and correction of complex, non-linear observation and model biases, leading to enhanced analysis quality and forecast reliability in the regional LAM 4D-Var suite. | KNMI |
28 | 3.4 | Direct integration of data assimilation into neural networks | In this task, we seek to integrate the recently published AI-Var approach (https://arxiv.org/abs/2406.00390) into a comprehensive data assimilation framework (AIDA) in order to streamline the development of innovative AI-based weather prediction technologies. A first version will be implemented in a new streamlined framework, that makes it easy for users to utilize the system. The preliminary framework will then be tested through specific data assimilation experiments. These will support gaining knowledge and experience about AI-Var for a subsequent extension of the system towards a completely AI-based data assimilation cycle. | An AI-based data assimilation framework with the capability of processing observational data and producing initial conditions for AI-based NWP models. | The implementation of an AI-based data assimilation approach is expected to yield substantial advancements towards a fully data-driven NWP chain, with the potential to set a new standard for operational weather forecasting and to demonstrate the transformative potential of AI in enhancing weather prediction capabilities. | DWD, ITAF MET Service |
29 | 3.5 | Fully data-driven data assimilation cycle | A fully AI-based data assimilation cycle integrates the DA scheme from Task 3.4 with a forecasting model to perform analysis as well as forecasting step using AI-based component with the goal of achieving a seamless, data-driven workflow. The observations are digested by the DA scheme in combination with the background state provided from the model. The system will then create an analysis state from which the next forecast can be issued. Both components are operated by a control script system that ensures a seamless data flow. | A fully AI-driven NWP data assimilation cycle based on a DA scheme, an NWP model and a control script system | A robust, fully AI‐driven data assimilation and forecasting system that delivers high‐accuracy, low‐latency atmospheric analyses and forecasts. | DWD, ARSO |
30 | 3.6 | Integration of AI-Var into Anemoi | In this task, we seek to integrate the developed AI-based data assimilation framework AIDA (Task 3.4) into the Anemoi framework in order to streamline the development of innovative AI-based weather prediction technologies with a focus on future operational employment. Subsequently, the Anemoi framework will be extended by a “Data assimilation” as well as a “Observations” module in order to facilitate data assimilation. The AIDA framework will then be modified or extended, building on the existing structures from the Anemoi “Models”, “Training” and “Datasets” modules. By incorporating the AI-Var approach into Anemoi, we expect to extend its capabilities to enable the integration of data assimilation with AI-based forecast models (WP1) and ensemble approaches (WP2). | An AI-based data assimilation module in the Anemoi framework with the capability of processing observational data and producing initial conditions for AI-based NWP models. | The incorporation of the AIDA within the Anemoi framework is expected to yield substantial advancements towards a fully data-driven NWP chain, with the potential to set a new standard for operational weather forecasting and to demonstrate the transformative potential of AI in enhancing weather prediction capabilities. | DWD, ITAF MET Service, Météo-France |
31 | 3.7 | Observation-driven models and multi-encoder-decoder architectures | In this task, the ultimate goal is to exploit observations either alone or in combination with reanalysis or NWP data sources. The method called AI-DOP (https://arxiv.org/abs/2412.15687) is largely considered here and the use of a new multi-source data handler in Anemoi. For the Member States, the adaptation of DOP to LAM or stretched-grid approaches are foreseen using additional local observational datasets. In order to ensure good quality of input datasets, quality control and data filtering are also needed. This exercise requires a similar data screening routine than the traditional data assimilation has without the dependency of NWP model fields (e.g., the background check). Additionally, the encode-decode blocks can be flexible and independent allowing irregular and variable input-output mesh. Preliminary tests with NetAtmo dataset have been carried out, but developments were not merged into Anemoi. In this task, we will closely follow the corresponding developments of multi-source data handlers and we plan to contribute with local adaptations and observations. RADAR observations are one data source that is considered to be tested in this framework. Also it is planned to exploit ML-ready observations that are prepared by E-AI program (namely SEVIRI and OPERA ZARR datasets) when Anemoi multi-encoder-decoder architecture is ready to utilise these observations. | The adaptation of Anemoi DOP for regional purposes; A multi-encoder-decoder framework in Anemoi that can digest various data sources without any pre-processing or projection of the input observations. | The effective use of multi-source observations improving the short-range skill of ML weather predictions. | MET Norway, DWD, MeteoSwiss, SMHI, LVGMC |
32 | 3.8 | Densification, multi-objective loss, and nowcasting applications | Alternative ways of obtaining good short-term forecast quality using pre-processed data, e.g. sparse grid for station data, projection on the same grid with spatial interpolation for radar and satellite. Different sources are handled through a multi-objective loss comparing pointwise data with RMSE-type loss, and dense grid data with spatial loss (e.g. Log spectral distance or Fourier Correlation loss). Attempt to produce nowcasts directly from observation, comparison with runs with NWP forecasts as additional input. | A quantification of nowcast quality lost when NWP data is removed from the model input. A SOTA nowcasting model with NWP inputs at least for surface variables (precip, temperature, dewpoint and wind). | Understanding of how GNN can actually learn physics if no physical reasoning is provided (through NWP runs). | MeteoSwiss |
WP4 Integrating data-driven models (as a component) into future forecasting systems requires assessing differences between traditional NWP models and ML-based models. NWP models use mathematical formulations and supercomputers, while ML models need large datasets and specialized GPU infrastructure. This shift demands careful integration of different software stacks, services, and infrastructure. Implementing DevOps practices, such as continuous integration and continuous deployment (CI/CD), is crucial for seamless development and updates. This work package will identify and address software and service gaps. It will also establish best practices for shared infrastructure. The work package explores the possibility of establishing joint CI/CD test infrastructure (for example on Atos, as in Destination Earth) integrated with the software repositories (e.g. on GitHub) used for sharing developments within the community. Further, it will showcase how MLOps best practices can be enforced through specialized services and various degrees of automation. Finally, it will engage with other work packages to define their technical requirements so to enable active collaboration on ML development. The tasks outlined in the following are cross-cutting and vital for the success of WPs 1 to 3. It is therefore expected that most of the project partners will contribute to the tasks of WP4 during the project period with some acting as Lead Contributors, as specified for each task. The initialisation of the tasks will be performed by MeteoSwiss and DMI. Each task in this work package is focused on specific objectives and outcomes, contributing to the overall goal of integrating data-driven models into existing forecasting systems through improved infrastructure and MLOps practices. This structured approach ensures clear deliverables and measurable outcomes, facilitating efficient project management and successful implementation.. | ||||||
33 | 4.1 | Assess MLOps Maturity and Promote Internal Knowledge Sharing | Develop a shared understanding of MLOps maturity and practices across the ML Pilot Project. This task focuses on surveying partners, documenting current practices, and sharing insights internally to foster alignment and convergence toward best practices in operational ML workflows. Activities:
| Report on MLOps maturity in the ML Pilot Project, based on survey responses and the public status page. Conducted MLflow community workshop with recorded materials and shared best-practice documentation. | Improved consistency, transparency, and adoption of MLOps practices within the ML Pilot Project, enabling more efficient and reliable deployment of machine-learning models. | MeteoSwiss, DWD, DMI, UK Met Office |
34 | 4.2 | Promote external knowledge sharing and alignment/ interoperabily of approaches | Align the ML Pilot Project’s MLOps strategies with broader European initiatives and facilitate knowledge exchange with the wider community. This task ensures that best practices are shared externally and that the pilot benefits from synergies with other projects Activities: ● Coordinate with the E-AI Working Group 3 to share insights and align strategies. | Documented exchanges and recommendations developed jointly with E-AI Working Group 3. | Enhanced interoperability and alignment of MLOps practices across European initiatives, promoting community-wide adoption and collaboration in operational ML workflows. | MeteoSwiss, UK Met Office, DWD, DMI |
35 | 4.3 | Develop joint ML infrastructure (a.k.a. “Scaling Anemoi”) | This task supports the joint development and maintenance of the infrastructure needed to train and run AI-based forecasting models, centered around the Anemoi framework. The work focuses on strengthening the reliability, interoperability, and usability of shared ML infrastructure across participating organizations. Activities:
| Integration tests and documented results ensuring component compatibility. Standardized and automated DevOps workflow, including CI/CD pipelines and consistent package release processes. Onboarding materials (demos, tutorials, documentation) hosted on a shared platform. Modular system architecture supporting custom plugins, with published interface standards. Public Anemoi Development Roadmap outlining scope, philosophy, and key development themes. Repository of Anemoi Decision Records (ADRs) documenting architectural and infrastructure decisions. Implemented caching option for downloaded integration test data. | A robust, scalable, and interoperable ML infrastructure within the Anemoi framework that enables collaborative development, efficient deployment, and long-term sustainability of AI-based forecasting systems across participating organizations. | MeteoSwiss, RMI, Météo-France, DMI, UK Met Office |
36 | 4.4 | Establish Best Practices for Shared Data, Infrastructure, and Services | Develop and disseminate comprehensive best practices for effectively utilizing shared data, infrastructure resources, and services. This includes: | Guidelines and templates for the usage of shared infrastructure and service usage. | Optimized utilization of shared resources, fostering improved collaboration, | MeteoSwiss, KNMI, UK Met Office, MET Norway |
37 | 4.5 | Governance and Lifecycle Support of ML artifacts for Anemoi Versions | Establish a governance framework for maintaining, modernising, and deprecating Anemoi model artifacts — including checkpoints, datasets, configurations, and workflows — to ensure long-term sustainability, reproducibility, and efficient integration within the shared infrastructure and across releases. The initial focus will be on checkpoints and models produced under the Machine Learning Pilot Project, followed by production-ready models, ensuring that all artifacts remain traceable, reproducible, and compatible with supported Anemoi versions. Activities: ● Define governance policies for ML artifacts (support levels, maintenance rules, deprecation criteria). ● Review and catalogue existing MLPP model artifacts and assess their compatibility with current Anemoi releases. ● Update Anemoi codebase to enable version-controlled interfaces for loading, saving, and managing artifacts. ● Implement CI/CD validation pipelines to automatically check artifact compatibility and lifecycle status. ● [Stretch goal] Develop monitoring and alerting tools for detecting data drift in maintained datasets and models. | Anemoi Lifecycle Governance Document detailing definitions of support, maintenance, and deprecation rules. Updated Anemoi codebase with enforced lifecycle and compatibility interfaces. Maintained and validated artifact repository (datasets, checkpoints, configurations) with explicit version tags. Automated CI/CD checks for lifecycle and compatibility enforcement. [Stretch goal] Prototype monitoring and alerting tools for data drift detection. | A governed, reproducible, and version-controlled ecosystem of ML artifacts within Anemoi, ensuring sustainability, interoperability, and operational efficiency across releases and participating partners. | MeteoSwiss, Météo-France, MET Norway, UK Met Office |
38 | 4.6 | Development of the “Anemoi Platform” (Anemoi Catalogue and Model Zoo) | Develop and self-host the Anemoi Catalogue as a community-driven “model zoo” and registry for shared ML models, metadata, and related resources. | Operational “Anemoi Platform” (Catalogue + Model Zoo) Contribution and governance framework Technical documentation and API references Demonstration use cases across partners | An open and collaborative Anemoi Platform enabling transparent sharing, discovery, and reuse of ML models and resources, fostering community engagement and accelerating adoption of AI-based forecasting across partners. | AEMET, FMI, MeteoSwiss, UK Met Office |
39 | 4.7 | Showcase ML Pipelines from the ML Pilot Project | This task highlights practical experiences from member states in operationalizing machine-learning (ML) models, with the goal of fostering shared learning, transparency, and reuse of good practices across the community. By documenting and communicating lessons learned, the task helps lower barriers for future operational ML implementations and strengthens collaboration within and beyond the project consortium. Activities: ● Engage with project partners to identify and document ML pipelines and their automation or operationalization workflows. ● Collect and curate lessons learned, including technical challenges, solutions, and success factors. ● Draft and publish blog-style technical articles to communicate insights and promote knowledge sharing beyond the consortium. | Collection of documented ML pipeline and automation case studies from member states (internal report or shared repository). One or more published blog posts or technical articles (e.g., AIFS blog, Medium, Substack) highlighting engineering insights and lessons from operational ML deployments. | Increased transparency, knowledge sharing, and cross-institutional learning on the operationalization of ML forecasting models, supporting broader adoption of effective and reproducible practices across the community. | MET Norway, UK Met Office, MeteoSwiss |
WP5 - Training and support Training and support on Machine Learning (ML) is clearly needed across the meteorological community to understand everything from the basics of ML in Numerical Weather Prediction (NWP), to how to utilise ML forecasts when providing forecasts and warnings to the public, to how ML NWP models may impact forecasting in the future. There are many general courses on machine learning available therefore it is proposed that ML training from ECMWF and the EMI focuses on the use of ML in the domain of Earth Sciences particularly meteorology with potential expansion to other related areas e.g. climatology, hydrology and air quality in the future. Training content will largely cover topics of the other work packages of this project, but may also be expanded based on needs. Possible training formats cover webinars, Jupyter Notebooks, e-learning modules, in-person multi-day courses, and documentation linked to results of the other work packages of this project. Work package 5, as a cross-cutting effort through all other work packages of the pilot project, will be coordinated by the project leads in close collaboration with ECMWF. | ||||||