📆 Project Period | January - June, 2026 |
👤 CIN Visiting Researcher |
Project Summary
- This project compares two geospatial foundation models (GFMs) developed under ESA Φ-lab, THOR and TerraMind, on ten Earth Observation (EO) use cases. The aim was to understand why the models differ, not only which one ranks first.
- Headline finding: within the explored configuration space, patch size and decoder type explain more performance variance than model identity.
- The models exhibit complementary investment strategies. TerraMind invests in pre-training and delivers strong features at a fixed resolution. THOR invests in finer tokenization at inference time, at a higher compute cost. Neither wins everywhere.
- The project delivers a diagnostic methodology that combines controlled ablations (over 800 runs) with dataset-level characterization, along with practical guidance for EO practitioners.
- Released outputs: experiment configuration files for all runs (GitHub) and a paper describing the study.
Development Tools
- TerraTorch: shared framework for fine-tuning and evaluating both models.
- THOR and TerraMind checkpoints (Tiny and Base): the models compared; TerraMesh was used for the partial TerraMind re-pretraining.
- NVIDIA RTX A6000 GPU (46 GB): hardware for all experiments.
Development Outputs
- Project repo: Terramind-vs-Thor-ESA-PhiLab contains the experiment configuration files for all runs.
- Model weights: THOR and TerraMind weights are available from their original publications (Forgaard et al., CVPR Workshops 2026; Jakubik et al., ICCV 2025).
- Datasets: the FAST-EO benchmark datasets used are previously published. The FM4CS datasets, in the configurations used here, are access-restricted; researchers can contact the Norwegian Computing Center.
Publications
- F. Schindlegger, K. Bounegta, E. Gmelich Meijling, J. Jakubik, A.-B. Salberg, T. Forgaard, N. Longépé and V. Marsocci, "Downstream Deployment Architecture Outweighs Model Choice in Geospatial Foundation Model Performance." [TO BE COMPLETED: submitted to Scientific Reports / preprint link / DOI]
Project Description
Context and motivation
GFMs are large neural networks pretrained on unlabelled, multi-sensor EO data, enabling them to transfer to many downstream tasks. Benchmarks usually rank them by aggregate score. A ranking says which model scores highest, but not why: a gap can come from the pretrained backbone, the decoder (the task-specific head), the fine-tuning regime, or spatial tokenization.
The number of GFMs has grown rapidly over the past few years, and there is still no consensus on how to benchmark them. THOR and TerraMind embody two contrasting design philosophies, so we treated them as a controlled natural experiment to move from ranking to attribution on a shared evaluation framework (TerraTorch).
Objectives
- Compare THOR and TerraMind under a shared protocol on ten EO use cases, varying one design axis at a time.
- Quantify how much performance variation comes from patch size, decoder, fine-tuning regime, input modality, and model scale, compared with model identity.
- Distill the results into hypotheses and practical deployment guidance.
The two models
Design axis | THOR | TerraMind |
Origin | FM4CS consortium (NR, UiT, ESA Φ-lab) | FAST-EO consortium (DLR, Jülich, IBM, KP Labs, ESA Φ-lab) |
Core philosophy | Compute-adaptive, native-resolution sensor fusion | Generative cross-modal representation |
Model type | Vision Transformer (FlexiViT) | Transformer encoder-decoder |
Sensors and modalities | Sentinel-1, -2, -3 (OLCI/SLSTR) | Sentinel-1, -2, land-use/land-cover maps, elevation, text |
Patch size | Variable at inference: 4, 8, 16, 32 | Fixed: 16 |
Ground sampling distance (GSD) handling | Native 10 m to 1 km, GSD-aware | Nominal 10–20 m (224² minimum input) |
Modality fusion | Cross-modal ViT fusion; token-level mean pooling | Modality-specific projections; token-level mean pooling |
Distinctive capability | Dynamic inference-time cost/detail trade-off | Thinking-in-Modalities (TiM): generating missing modalities as intermediate steps |
A patch is the image tile that a Vision Transformer encodes into a single token. Halving the patch size quadruples the token count and roughly quadruples encoder compute. THOR can change patch size at inference without retraining; TerraMind cannot.
Use cases
Ten use cases were adopted from the two ESA consortia, five each. The FAST-EO cases include classical benchmarks and multimodal setups. The FM4CS cases target newer problems, with an Arctic focus, more radar (SAR) data, Sentinel-3 data, and more varied task types.
Consortium | Dataset | Sensor(s) | GSD | Task | Metric |
FAST-EO | Sen1Floods11 | Sentinel-1 + Sentinel-2 | 10 m | Segmentation (flood / no flood) | mIoU |
FAST-EO | HLS Burn Scars | Sentinel-2 (HLS) | 30 m | Segmentation (burned / unburned) | mIoU |
FAST-EO | CocoaMining | S1 + S2 + elevation | 10 m | Segmentation (mines/cocoa/background) | mIoU |
FAST-EO | Methane Leaks | Airborne hyperspectral | – | Classification (methane plume) | Overall accuracy |
FAST-EO | HYPERVIEW | Airborne hyperspectral | – | Regression (soil properties: K, Mg, P₂O₅, pH) | RMSE |
FM4CS | Flood Zone | Sentinel-1 | 10 m | Change detection (flood / no flood) | F1 (pixel-wise) |
FM4CS | Iceberg Detection | Sentinel-1 | 10 m | Object detection | F1 (instance-wise) |
FM4CS | Sea Ice | Sentinel-1 | 240 m | Segmentation (3 classes) | mIoU |
FM4CS | Snow Monitoring | Sentinel-3 SLSTR | 500 m | Regression (snow cover) | RMSE |
FM4CS | Mires/Wetlands Mapping | Sentinel-2 | 10 m | Segmentation (6 classes) | mIoU |
mIoU is the mean intersection-over-union across classes; RMSE is the root-mean-square error.
Approach and experimental design
All hyperparameters not intrinsic to a model's architecture were fixed, following the PANGAEA evaluation philosophy. Each model was run at its best achievable configuration on a single NVIDIA RTX A6000 GPU (46 GB), yielding more than 800 runs in total. The ablation axes were:
- Patch size: 4, 8, 16, and 32 for THOR; 16 for TerraMind.
- Decoder: Linear (a single projection, isolating backbone quality) and UperNet (a multi-scale decoder with pyramid pooling). Flood Zone and Iceberg Detection used lightweight PixelShuffle and MLP decoders.
- Backbone regime: frozen (only the decoder is trained) versus fine-tuned (the whole model is trained).
- Input modality, model scale (Tiny vs. Base), and training-data fraction, the last compared against a UNet trained from scratch.
THOR was given the true physical extent of each input tile so that its position encoding matches the ground resolution. TerraMind received tiles within or near its pretraining size range, and smaller inputs were resized. The two models, therefore, do not always see geometrically identical inputs, most notably on CocoaMining.
Key results
Results hold within the explored configuration space and under the present training recipe. Not every ablation was run on every use case.
Use case (metric) | TerraMind, best | THOR, best |
Sen1Floods11 (mIoU, fine-tuned) | 0.910 | 0.904 |
HLS Burn Scars (mIoU, fine-tuned) | 0.891 | 0.865 |
CocoaMining (mIoU, fine-tuned) | 0.817 | 0.821 |
Mires/Wetlands (mIoU, frozen) | 0.646 | 0.636 |
Flood Zone (pixel-wise F1, frozen) | 0.33 | 0.73 |
Iceberg Detection (instance-wise F1, frozen) | 0.823 | 0.857 |
Sea Ice (mIoU; TerraMind with TiM, THOR fine-tuned) | 0.755 | 0.873 |
Snow (RMSE, lower is better; fine-tuned) | 8.956 | 5.922 |
1. Design choices explain more variance than model identity. A one-way analysis of variance (ANOVA) on four datasets (Sen1Floods11, HLS Burn Scars, CocoaMining, Snow) measured the share of variation explained by each factor. Patch size and decoder type jointly explained more than the choice of model.
2. Spatial tokenization is the dominant lever. For THOR, the step from patch size 32 to 16 gave the largest gain on most datasets (2–7 percentage points, pp). CocoaMining, where mining sites occupy a small part of each tile, gained 15.5 pp from ps32 to ps4. For compact, minority-class targets in SAR data, coarse patches were catastrophic. On Flood Zone, one THOR decoder configuration collapsed to F1 = 0.004 at ps32 and recovered only at ps4 (0.73). The best patch size is task-dependent: Sea Ice gained the most from ps16 to ps8 (+26 pp; UperNet, fine-tuned), whereas Wetlands was non-monotonic.
At matched patch size 16, the two encoders cost nearly the same (62 vs. 56 GMACs, billions of multiply-accumulate operations). TerraMind led on Sen1Floods11 (by 3.3 pp), HLS Burn Scars (by 4.5 pp), and Wetlands (0.646 vs. 0.578), while THOR already led on Sea Ice. TerraMind pays at pretraining time; THOR pays at inference time. Upsampling training images did not reproduce the benefit of a finer token grid.
3. Decoder and resolution interact. For THOR, UperNet's advantage over Linear shrinks as patches get finer: on Sen1Floods11 from about 12 pp at ps16 to about 5 pp at ps8. For Wetlands, the gap remains at 9.3 pp at ps4, so convergence is task-dependent. For TerraMind, UperNet consistently helps (0.6 pp on CocoaMining to 14.2 pp on Wetlands). Among lightweight decoders, no single head wins across models and tasks.
Cost matters too. On Sen1Floods11, ps4 with Linear costs about 1,004 GMACs, against about 375 GMACs for ps8 with UperNet, roughly 2.7 times cheaper. Ps8 with UperNet is the recommended THOR operating point when compute is constrained.
4. Frozen versus fine-tuned. For TerraMind, fine-tuning gave a modest lift on most datasets (for example, +6.3 pp on HLS Burn Scars), but a frozen backbone fell far behind on Sea Ice (linear decoder: 0.092 vs. 0.627 fine-tuned). For THOR, fine-tuning generally helped by 2–5 pp, but at coarse patch sizes, frozen sometimes won (Sen1Floods11 at ps8: 0.894 vs. 0.889). At ps4 on Sea Ice, fine-tuning added 23 pp.
5. Multimodal fusion: a negative result. Simple Sentinel-1 + Sentinel-2 fusion did not reliably beat Sentinel-2 alone in the full-data regime. On Sen1Floods11, TerraMind reached 0.910 with Sentinel-2 only against 0.896 with both. The original THOR and TerraMind papers report small SAR gains with less data, consistent with a data-regime effect. Elevation data added no measurable benefit to CocoaMining.
6. Cross-sensor generalization on Snow. TerraMind was not pretrained on Sentinel-3 SLSTR. We mapped SLSTR bands to the closest Sentinel-2 wavelengths and used its existing Sentinel-2 input head. It achieved an RMSE of 8.956, compared to 5.922 for THOR, which natively handles Sentinel-3. Wavelength-based band mapping may be a route to extend a GFM to new sensors without retraining an input head.
7. Model scale and cost. Base generally beat Tiny, with a task-dependent gap. TerraMind-Tiny's encoder costs about 3.7 GMACs, compared to 56.0 for Base, yet achieves 94–99% of Base's mIoU with UperNet, making it the most compute-efficient configuration in the benchmark. Sea Ice showed some Tiny-over-Base reversals, but its splits look unstable, so we draw no firm conclusion.
8. Low-data regime. On Sen1Floods11, fine-tuned THOR at ps4 reached 0.888 mIoU with 5% of the training data (13 images), above the 0.871 mIoU a UNet reached with 50% (126 images). On HLS Burn Scars, the UNet led at most fractions, and frozen GFMs trailed it by up to 11.4 pp.
9. A TerraMind patch-size-4 probe. With IBM Research, we obtained a TerraMind-Tiny variant whose patch embedding was partially re-trained at a patch size of 4 (48 epochs on TerraMesh). It reached 0.919 on Sen1Floods11 (vs. 0.910 for TerraMind at ps16 and 0.904 for THOR at ps4) and 0.894 on HLS Burn Scars (vs. 0.891 and 0.865). This is a preliminary, uncontrolled signal, but it suggests fine-patch benefits are not exclusive to THOR's architecture.
Practical guidance for EO practitioners
- Fix the compute budget, tokenization, and decoder first; these mattered at least as much as the model itself.
- Under a compute constraint, use THOR at ps8 with UperNet.
- Use UperNet with TerraMind, and with THOR at ps16 or coarser.
- Avoid coarse patches for rare or compact targets.
- Do not assume adding Sentinel-1 to Sentinel-2 helps when optical data are abundant.
- With very little labeled data, we still compare against a task-specific UNet.
Limitations and caveats
- All main experiments used a single seed (0); gaps below about 1 pp may not be reproducible.
- THOR's search space is larger than TerraMind's, giving it more chances to find a well-tuned point.
- Sea Ice and Flood Zone show split instability, and Flood Zone, Iceberg Detection, and Wetlands were evaluated only with frozen backbones.
- The patch-size-4 TerraMind result is an uncontrolled preliminary probe.
- Findings should not be assumed to generalize beyond the evaluated models and tasks.
Impact and future developments
The study offers a reusable template for GFM comparison: controlled ablation combined with dataset-level characterization. Future directions are gated or cross-attention fusion, full TerraMind re-pretraining at finer patch sizes, applying the methodology to other GFMs, and adding multiple seeds and statistical robustness analysis.