🏛️ Company / Organization | German Aerospace Center (DLR) Remote Sensing Technology Institute, IBM Research, EuropeForschungszentrum Jülich GmbH (FZJ) Jülich Supercomputing Centre (JSC), KP Labs sp. z o. o. |
📆 Contract Duration with ESA Φ-lab | January 2024 – July 2025 |
🌍 Project Title | FAST-EO — Fostering Advancements in Foundation Models via Unsupervised and Self-supervised Learning for Downstream Tasks in Earth Observation |
🌐 Project Website | |
📍 Project GitHub | |
🔨 Project Linkedin |
Project Description
Abstract
- FAST-EO developed TerraMind, the first any-to-any generative, large-scale multimodal foundation model for Earth observation, pre-trained on approximately 500 billion tokens from global geospatial data.
- The model processes inputs simultaneously at pixel-level, token-level, and as sequences, ingesting modalities such as Sentinel-2 L2A, Sentinel-1 RTC/GRD, Digital Elevation Models (DEM), NDVI, Land-Use/Land-Cover (LULC) maps, image captions, and geographic coordinates.
- TerraMind introduces the Thinking-in-Modalities (TiM) paradigm, enabling the model to first generate auxiliary modalities and then leverage them to improve downstream predictions.
- TerraMind outperforms existing deep learning models for Earth observation across community benchmarks such as PANGAEA, establishing new state-of-the-art results.
- The project produced the TerraMesh dataset — a planetary-scale multimodal EO dataset comprising approximately nine million globally distributed samples.
- All models, code, and data are released open-source under an Apache-2.0 license, fostering community adoption and reproducibility.
Detailed description
Overview
FAST-EO (Fostering Advancements in Foundation Models via Unsupervised and Self-supervised Learning for Downstream Tasks in Earth Observation) is an ESA Φ-lab–funded project that brings together five European partners — DLR (lead), IBM Research Europe, Forschungszentrum Jülich (FZJ-JSC), KP Labs, and ESA Φ-lab — to advance foundation models for Earth observation.
The project's central output is TerraMind, the first any-to-any generative multimodal foundation model for Earth observation. TerraMind was pre-trained on approximately 500 billion tokens drawn from globally distributed geospatial data, making it the largest openly available model of its kind.
Figure 1. The TerraMind pipeline. On the left, diverse Earth observation modalities (Sentinel-2, Sentinel-1, DEM, NDVI, LULC, text captions, and coordinates) are ingested. In the center, these are converted through modality-wise tokenization into a shared token representation. The encoder–decoder learns cross-modal correlations via masked token prediction. On the right, the learned representation enables multimodal generation, native multimodal fine-tuning, and the Thinking-in-Modalities paradigm.
TerraMind Architecture and Capabilities
TerraMind digests inputs at three levels simultaneously:
- Pixel-level input: Raw satellite imagery (Sentinel-2 L2A, L1C, RGB; Sentinel-1 GRD, RTC; DEM) is processed through learned vector quantization into discrete patch tokens.
- Token-level input: Image patches are encoded by modality-specific tokenizers, then fed into the transformer backbone.
- Sequence input: Text captions and geographic coordinates are tokenized using standard text tokenization and concatenated with image tokens.
The model's dual-scale fusion approach combines a generative decoder (for any-to-any modality generation) with a discriminative encoder (for downstream classification and segmentation tasks).
The Thinking-in-Modalities (TiM) paradigm is a key innovation: during fine-tuning and inference, TerraMind first generates auxiliary modalities (e.g., synthesizing a SAR image from an optical input), then uses these generated modalities as additional context to improve the quality of its final predictions.
Any-to-Any Generation
TerraMind's generative capabilities allow it to translate between arbitrary combinations of Earth observation modalities. Given any subset of available modalities, the model can generate the missing ones - for instance, producing a synthetic radar view from optical imagery or reconstructing optical data from a SAR-only input.
Figure 2. Thinking-in-Modalities (TiM) example. From the original cloudy Sentinel-2 L2A image (right), TerraMind generates a realistic Sentinel-1 RTC synthetic view (left), preserving the continuity of the river basin and coastal features.
Figure 3. Cross-modal generation example. From the original Sentinel-2 L2A image (right), TerraMind generates a Sentinel-1 GRD view (left), capturing terrain relief and forest cover patterns.
This technology opens new opportunities to reconstruct missing observations, fill cloud gaps, and provide richer multi-modal inputs for downstream applications. It can augment scarce training data and even simulate "virtual sensors" or future missions to test algorithms and monitoring strategies in advance.
TerraMesh Dataset
To train TerraMind, the consortium created TerraMesh — a planetary-scale, multimodal Earth observation dataset comprising approximately nine million globally distributed samples and hundreds of billions of tokens across aligned EO modalities. TerraMesh integrates Sentinel-1, Sentinel-2, DEM, NDVI, LULC, text captions, and coordinates into a unified data structure, providing the community with an unprecedented resource for pre-training and benchmarking multimodal EO models.
Downstream Performance
TerraMind establishes new state-of-the-art results on community benchmarks such as PANGAEA, outperforming previous models, including DINOv2-ViT-L, DOFA-ViT-L, Prithvi v1/v2, and SSL4EO-ResNet50, across segmentation, classification, and change detection tasks. The TiM paradigm consistently improves performance over uni-modal baselines, with gains of up to +5.28% F₁ on challenging tasks such as artisanal mining detection in Ghana.
Research and Development Tools
TerraMind was developed in five core phases: (i) dataset generation, (ii) learned quantization, (iii) tokenization of data, (iv) correlation learning, and (v) evaluations on downstream applications.
- Data processing and dataset construction: The TerraMesh dataset was built using xarray, rasterio, and zarr for large-scale geospatial data processing and storage. Jobs were parallelized across a large CPU cluster, with zarr's compressed format minimizing storage bottlenecks.
- Deep learning framework: All model components — learned quantization, tokenization, correlation learning, and downstream evaluation — were implemented in PyTorch and distributed across GPUs via Distributed Data Parallel (DDP) and Fully Sharded Data Parallel (FSDP) strategies.
- HPC infrastructure: Training was executed on the JUWELS Booster at the Jülich Supercomputing Center, equipped with NVIDIA A100 GPUs in multi-GPU nodes. SLURM and custom launch scripts (torchrun_jsc) coordinated automated job submission and distributed training across cluster nodes. LLview was used to track cluster performance and manage resources.
- Data management and transfer: Jülich DataPub served as the public repository for AI-ready datasets, with JUDAC (Jülich Data Access) providing optimized, high-volume, and secure data transfers to and from the supercomputing facility.
- Benchmarking and evaluation: Downstream evaluations relied on community benchmark settings provided by PANGAEA and related benchmarks.
Project Outcomes
Models and Weights
Pretrained TerraMind models are released in multiple sizes (Tiny, Small, Base, Large) under a permissive Apache-2.0 license.
- FAST-EO HuggingFace: https://huggingface.co/FAST-EO
- IBM–ESA Geospatial Models: https://huggingface.co/ibm-esa-geospatial
Datasets
- TerraMesh — large-scale multimodal EO pretraining dataset (~9M samples): https://huggingface.co/datasets/ibm-esa-geospatial/TerraMesh
Code and Tooling
- TerraMind GitHub Repository (model code, fine-tuning, TiM workflows): https://github.com/IBM/terramind
- TerraTorch Integration (PyTorch Lightning–based fine-tuning and evaluation): https://github.com/terrastackai/terratorch
- TerraMind Website: https://ibm.github.io/terramind/
Community Challenges
- TerraMind Blue-Sky Challenge — IBM–ESA Φ-lab open innovation challenge: https://huggingface.co/spaces/ibm-esa-geospatial/challenge
Publications
- J. Jakubik, F. Yang, B. Blumenstiel, E. Scheurer, R. Sedona, S. Maurogiovanni, J. Bosmans, N. Dionelis, V. Marsocci, N. Kopp, R. Ramachandran, P. Fraccaro, T. Brunschwiler, G. Cavallaro, J. Bernabe-Moreno, N. Longépé, "TerraMind: Large-Scale Generative Multimodality for Earth Observation", ICCV 2025, pp. 7383–7394. Paper
- C. T. Marimo, B. Blumenstiel, M. Nitsche, J. Jakubik, T. Brunschwiler, "Beyond the Visible: Multispectral Vision-Language Learning for Earth Observation (MS-CLIP)", ECML PKDD 2025, pp. 516–532. arXiv
- R. S. Kuzu, G. Cavallaro, T. Brunschwiler, J. Nalepa, C. O. Dumitru, A. Zappacosta, D. Espinoza Molina, R. Kienzler, J. Jakubik, B. Blumenstiel, P. Fraccaro, R. Sedona, E. Scheurer, S. Maurogiovanni, A. M. Wijata, D. Marek, J. Sadel, L. Tulczyjew, N. Dionelis, N. Longépé, "FAST-EO: Transforming Earth Observation Through Multi-Modal Foundation Models", Living Planet Symposium, Vienna, Austria, 2025. elib.dlr.de
- B. Blumenstiel, P. Fraccaro, V. Marsocci, J. Jakubik, S. Maurogiovanni, M. Czerkawski, R. Sedona, G. Cavallaro, T. Brunschwiler, J. Bernabe-Moreno, N. Longépé, "TerraMesh: A Planetary Mosaic of Multimodal Earth Observation Data", CVPR EarthVision 2025, pp. 2394–2402. Paper
- S. Ofori-Ampofo, A. Zappacosta, R. S. Kuzu, P. Schauer, M. Willberg, X. X. Zhu, "Assessing Cocoa Landuse Change and Exposure to Artisanal and Small-Scale Gold Mining Using Satellite Imagery", IEEE JSTARS, vol. 19, pp. 10083–10095, 2026. DOI
- S. Ofori-Ampofo, A. Zappacosta, R. S. Kuzu, P. Schauer, M. Willberg, X. X. Zhu, "SmallMinesDS: A Multimodal Dataset for Mapping Artisanal and Small-Scale Gold Mines", IEEE GRSL, vol. 22, pp. 1–5, 2025. DOI
- A. Banze, T. Stassin, N. A. A. Braham, R. S. Kuzu, S. Besnard, M. Schmitt, "HyBiomass: Global Hyperspectral Imagery Benchmark Dataset for Evaluating Geospatial Foundation Models in Forest Aboveground Biomass Estimation", IEEE GRSL, vol. 22, pp. 1–5, 2025. DOI
- L. Chiarabini, D. Espinoza Molina, A. Zappacosta, R. S. Kuzu, A. Camero, "Multimodal Learning for Earth Observation: Automating Satellite Image Captioning with Geo-FMs", Helmholtz AI Conference, 2025. elib.dlr.de
- J. Sadel, L. Tulczyjew, A. M. Wijata, M. Przeliorz, J. Nalepa, "Monitoring Forest Changes With Foundation Models and Sentinel-2 Time Series", IEEE GRSL, 2025. DOI
- J. Nalepa, L. Tulczyjew, B. Le Saux, N. Longépé, B. Ruszczak, A. M. Wijata, K. Smykala, M. Myller, M. Kawulok, R. S. Kuzu, F. Albrecht, C. Arnold, M. Alasawedah, S. Angeli, D. Nobileau, A. Ballabeni, A. Lotti, A. Locarini, D. Modenini, P. Tortora, M. Gumiela, "Estimating Soil Parameters From Hyperspectral Images: A Benchmark Dataset and the Outcome of the HYPERVIEW Challenge", IEEE GRSM, vol. 12, no. 3, pp. 35–63, 2024. DOI