📆 Project Period | March - July 2026 |
👤 CIN Visiting Researcher |
Project Summary
- Assessed whether TerraMind transfers its representation from Sentinel-2 to real $\Phi$-sat-2 imagery.
- Built and analyzed a co-registered dataset of real $\Phi$-sat-2, Sentinel-2B, and simulated $\Phi$-sat-2 patches.
- Developed a cross-architecture knowledge-distillation approach with a lightweight U-Net student and domain-specific batch normalization.
Development Tools
- TerraMind: multimodal geospatial foundation model used as the frozen teacher and for latent-space analysis.
- PyTorch and TerraTorch: model training, fine-tuning, knowledge distillation, DSBN, and segmentation experiments.
- U-Net: compact CNN student architecture selected for segmentation and on-board deployment constraints.
- OrbitalAI degradation simulator: physics-based simulation of $\Phi$-sat-2-like spectral, spatial, noise, band-misalignment, and optical effects.
- eo-learn: modular Earth observation processing tasks used to structure the degradation pipeline.
- LightGlue: deep feature matching used to establish correspondence between $\Phi$-sat-2 and Sentinel-2B scenes.
- ESA Insula Perception platform: access point for $\Phi$-sat-2 acquisitions.
- ESA HPC infrastructure and HDF5: storage of triplet data, provenance information, and split manifests.
- OpenVINO and Myriad 2 deployment constraints: used to guide the compact student architecture.
Development Outputs
- Project repository: terra-sat-drift, containing the research code, configuration files, data-processing components, model-training tasks, and domain-gap analysis tools.
- Co-registered triplet dataset: 253,228 patches from 1,299 $\Phi$-sat-2 products, stored on ESA HPC infrastructure under the applicable data-sharing conditions.
- Physics-degraded Sen1Floods11 dataset: simulated $\Phi$-sat-2-like flood scenes and degradation ladders for controlled segmentation experiments.
- Research documentation: the completed master's thesis, including methodology, results, limitations, and reproducibility information. No publication has been released yet. A manuscript based on the results of this collaboration is currently under discussion with the ESA supervisor and the home-organization supervisor of the thesis.
Outcome of the collaboration
The collaboration produced a reproducible study of sensor domain shift for GeoFM-based on-board processing, a co-registered real/simulated dataset for measuring that shift, a physics-degraded flood-segmentation benchmark, and a sensor-adaptive knowledge-distillation method for a compact student model.
Project Description
On-board processing can reduce the delay between Earth observation acquisition and the generation of actionable information, which is important for time-critical applications such as flood and wildfire response. Geospatial foundation models (GeoFMs) are promising for this setting because their large-scale pre-training can provide reusable representations for downstream tasks with comparatively few labeled examples. However, a model trained on archived data from one satellite sensor may not transfer reliably to a different sensor in orbit.
This collaboration with the ESA $\Phi$-lab investigated that problem using TerraMind and the $\Phi$-sat-2 mission as a case study. TerraMind is a multimodal GeoFM with a Vision Transformer encoder. Its architecture is not directly compatible with the space-qualified Intel Movidius Myriad 2 processor used by the $\Phi$-sat-2 on-board payload, and its multispectral pre-training is based on Sentinel-2 data, which is similar to $\Phi$-sat-2 data but contains a shift. The project addressed two challenges: measuring the sensor domain gap and transferring useful knowledge into a compact, hardware-compatible model.
The collaboration enabled work with real $\Phi$-sat-2 acquisitions accessed via the Insula Perception platform. These acquisitions were matched with corresponding Sentinel-2B scenes using spatio-temporal filtering and deep feature matching. The resulting triplets contain three views of the same underlying scene: real $\Phi$-sat-2 imagery, the corresponding Sentinel-2B imagery, and a physics-simulated $\Phi$-sat-2 rendering of the Sentinel-2 scene. The dataset contains 253,228 co-registered patches from 1,299 $\Phi$-sat-2 products. Land-cover labels were derived from ESA WorldCover 2021 and rasterized to the $\Phi$-sat-2 footprint.
Representative triplet patches showing corresponding observations from the real $\Phi$-sat-2, simulated $\Phi$-sat-2, and Sentinel-2 domains.
The triplets enabled a direct measurement of the domain gap. Comparing paired sensor views showed that TerraMind's latent representations are substantially affected by changes in sensors. Paired cosine similarity was 0.098, and a linear probe trained on the source domain retained approximately 33% of its accuracy when transferred to the real $\Phi$-sat-2 domain. The discrepancy was already visible in the early encoder layers, consistent with a low-level radiometric origin related to the sensors' different spectral responses rather than a change in semantic scene content. The physics-based simulator also provided a controlled way to study noise, band misalignment, and optical blur separately.
Paired latent representations from the teacher model; links connect two sensor views of the same ground patch.
The second part focused on knowledge distillation. A TerraMind teacher was fine-tuned on Sentinel-2 land-cover data and then frozen. Its predictions supervised a compact U-Net student processing both Sentinel-2 and $\Phi$-sat-2 branches. The student shared its convolutional weights across domains, while domain-specific batch normalization (DSBN) provided separate normalization statistics and affine parameters for each sensor. This allowed the student to adapt to the known input sensor without duplicating the full network. A paired-contrastive objective was also implemented and evaluated, but it did not improve performance beyond DSBN alone.
The student was evaluated against a supervised U-Net baseline with the same architecture and parameter count, trained directly on labeled $\Phi$-sat-2 data. At the 10,000-patch label budget, KD + DSBN achieved a test mIoU of 0.4481, compared with 0.3739 for the supervised baseline. At the 1,000-patch budget, the supervised baseline achieved 0.2707 compared with 0.2635 for KD + DSBN. These results show that sensor-aware distillation can be beneficial when enough target-domain labels are available, but it is not automatically superior in the low-label regime.
Representative land-cover ground-truth and prediction masks from the U-Net student model trained with KD + DSBN on 10,000 $\Phi$-sat-2 labels.
The project also produced a physics-degraded version of the Sen1Floods11 flood-segmentation benchmark. This dataset was used to isolate individual degradation factors in the absence of a labeled real $\Phi$-sat-2 flood event. At full simulated severity, noise caused the largest drop in student IoU, followed by band misalignment and optical blur. These results complement the real triplet analysis while remaining subject to the simulator's limitations.
Simulator schema: Sentinel-2 products are transformed through sensor-specific spectral, spatial, alignment, noise, and optical stages.
Example degradation ladder used to isolate the effect of individual simulated sensor factors on flood segmentation.
The main outcome is that a GeoFM can contribute to on-board processing across a sensor gap, but its representation cannot be assumed to be sensor-invariant. The sensor gap should be measured with paired acquisitions where possible, and domain adaptation should be integrated with compression rather than treated as an unrelated post-processing step. Future work includes product-level rather than patch-level dataset splits, validation on a real labeled $\Phi$-sat-2 flood event, comparisons with mission-specific models such as $\Phi$-satNet, and replication across other sensor pairs and missions.
Downstream performance comparison showing that the benefit of distillation depends on the available target-domain label budget.