Flow Matching for Efficient and Scalable Data Assimilation

Jun 15, 2026·
Taos Transue
,
Bohan Chen
,
So Takao
,
Bao Wang
· 1 min read
Abstract
Data assimilation (DA) estimates a dynamical system’s state from noisy observations. Recent generative models like the ensemble score filter (EnSF) improve DA in high-dimensional nonlinear settings but are computationally expensive. We introduce the ensemble flow filter (EnFF), a training-free, flow matching (FM)-based framework that accelerates sampling and offers flexibility in flow design. EnFF uses Monte Carlo estimators for the marginal flow field, localized guidance for observation assimilation, and utilizes a novel flow path that exploits the Bayesian DA formulation. It generalizes classical filters such as the bootstrap particle filter and ensemble Kalman filter. Experiments on high-dimensional benchmarks demonstrate EnFF’s improved cost-accuracy tradeoffs and scalability, highlighting FM’s potential for efficient, scalable DA.
Type
Publication
SIAM/ASA Journal on Uncertainty Quantification (accepted for publication)

The ensemble flow filter (EnFF) is a training-free data-assimilation framework that uses flow matching to transform a forecast ensemble into samples from the filtering distribution. Its Monte Carlo flow-field estimator and localized observation guidance avoid model training while retaining the flexibility of generative flow design.

The paper introduces a filtering-to-predictive (F2P) flow that uses the previous filtering distribution, rather than a standard Gaussian, as its reference. This path is better aligned with sequential Bayesian filtering and improves efficiency and robustness when only a small number of sampling steps is available. The analysis also shows how EnFF recovers the bootstrap particle filter and ensemble Kalman filter under appropriate choices and assumptions.

Experiments span Lorenz-63, Lorenz-96, the one-dimensional Kuramoto–Sivashinsky system, and two-dimensional Navier–Stokes equations, including state dimensions up to a 256×256 grid. Across these benchmarks, EnFF provides a strong accuracy–cost trade-off and scales to nonlinear, high-dimensional data-assimilation problems.