📆 Project Period | April – June, 2026 |
👤 CIN Visiting Researcher |
Project Summary
- Investigated whether meaningful land-cover information can be extracted directly from Level-0 Synthetic Aperture Radar (SAR) measurements, before completing the conventional SAR focusing pipeline. The broader motivation was to explore whether downstream AI tasks can operate on earlier signal representations, reducing processing requirements for future onboard Earth Observation systems.
- Developed an experimental workflow using the Maya4 multi-level Sentinel-1 SAR dataset to study machine-learning-based classification across early SAR representations. The work addressed practical challenges associated with large-scale raw SAR data, representation design, normalization, patch extraction, and model training.
- The collaboration established a technical foundation for task-driven SAR processing, in which the amount of signal processing performed is determined by the downstream AI task rather than by automatically producing a fully focused image. Future work will extend the study through systematic cross-level benchmarking, lightweight model design, and deployment-oriented evaluation of accuracy, latency, memory, and computational cost.
Development Tools
- Maya4: Multi-level Sentinel-1 SAR dataset and software framework providing access to representations across the SAR processing chain, from raw measurements through intermediate processing stages to focused data. It formed the primary data infrastructure for investigating learning from Level-0 and early SAR representations.
- Python and PyTorch: Used to develop the machine-learning experimentation and classification pipeline, including data preparation, model training, and evaluation.
- Zarr-based data access: Used through the Maya4 ecosystem to work efficiently with large SAR datasets without requiring complete acquisitions to be loaded into memory.
- Hugging Face / cloud-based dataset access: Supported access to the large-scale Maya4 data resources and reproducible dataset handling.
- Scientific Python ecosystem: Supporting numerical processing, experimental analysis, and visualization was used throughout the development and evaluation workflow.
Development Outputs
- Level-0 SAR classification experimental pipeline: A research prototype for loading, preparing, and evaluating early-stage SAR representations using machine-learning-based land-cover classification.
- Cross-processing-level evaluation framework: An experimental structure that can be extended to systematically compare raw, intermediate, and focused SAR representations under a common downstream task.
- Experimental analyses and trained model artifacts: Initial experiments developed during the Visiting Researcher period provided the foundation for continued investigation after the visit.
- Maya4: The project made extensive use of the publicly available Maya4 multi-level Sentinel-1 SAR dataset and software ecosystem developed at ESA Φ-lab. Public resources: Maya4 GitHub repository, Maya4 Python package, and ESA Φ-lab Maya4 data resources.
Upcoming outputs
- Extended benchmarking of land-cover classification across multiple SAR processing levels.
- Evaluation of the trade-off between classification performance and avoided signal-processing cost.
- Investigation of lightweight and deployment-oriented architectures for early-stage SAR representations.
- End-to-end profiling incorporates processing latency, memory, and computational requirements.
- Preparation of the work for publication following completion of the extended experimental study.
Publications
No dedicated publication resulting from this Visiting Researcher project has been released at the time of submission. The work is being extended with the objective of developing a publication around machine learning from early/Level-0 SAR representations and its implications for onboard Earth Observation processing.
Project Description
Motivation
Synthetic Aperture Radar (SAR) is an important Earth Observation modality because it can acquire information independently of daylight and under most weather conditions. However, the measurements collected by a SAR sensor are not immediately available as conventional images. Raw radar echoes undergo a sequence of signal-processing operations before becoming the focused products typically used by computer vision and machine-learning algorithms.
This processing chain is appropriate when the objective is to create a high-quality image for human interpretation or general downstream use. For an autonomous satellite, however, the end goal may instead be a particular decision: identifying a type of terrain, detecting an object, recognizing an anomaly, or deciding whether to prioritize an acquisition for downlink.
This raises an important question:
Does an onboard AI system always need a fully focused SAR image to make a useful semantic decision?
The Visiting Researcher project at ESA Φ-lab investigated this question using land-cover classification as a downstream task. The central objective was to determine whether discriminative information could be learned directly from Level-0 or other early SAR representations before the conventional focusing chain was completed.
The project, therefore, sits at the intersection of SAR signal processing, machine learning, and onboard AI. Rather than considering SAR image formation and AI inference as two completely independent stages, the work explored whether they can be considered jointly as part of a task-driven processing pipeline.
Working with multi-level SAR data
A key enabler of the project was Maya4, a multi-level SAR dataset and processing framework developed around Sentinel-1 acquisitions. Maya4 exposes data at different stages of the SAR processing chain, including raw measurements and progressively processed representations leading towards focused SAR imagery.
This structure is particularly valuable for machine-learning research because most commonly available SAR datasets provide only already processed products. Such datasets make it difficult to investigate what information is available at intermediate stages or whether a downstream neural network could operate before conventional image formation is complete.
The project used Maya4 to develop an experimental workflow that enabled SAR data to be accessed, transformed into suitable model inputs, and evaluated using land-cover classification as a semantic probe.
Working with Level-0 SAR also introduces a different set of challenges from conventional image classification. Raw SAR measurements cannot simply be treated as standard grayscale photographs. Their structure is determined by radar acquisition physics and by the signal-processing operations that have or have not yet been applied. Appropriate data extraction, normalization, representation, and sampling, therefore, form an important part of the machine-learning problem itself.
A significant part of the collaboration consequently involved understanding these representations and creating a practical bridge between the SAR processing domain and a conventional deep-learning training pipeline.
Machine learning as a probe of information content
The goal of the classification experiments was not simply to maximize land-cover classification accuracy. Classification was instead used as a controlled downstream task for asking a more fundamental question: at what point in the SAR processing chain does sufficient semantic information become accessible to a learning algorithm?
The experimental framework was developed to enable the study of different SAR representations within a common machine-learning workflow. This provides a basis for examining the relationship between the amount of signal processing performed and the usefulness of the resulting representation for a downstream task.
This perspective is important for onboard AI. Conventional SAR processing stages have been designed primarily around image reconstruction quality. An autonomous system may have a different optimization objective. If the purpose of an acquisition is to answer a particular semantic question, a representation that is less suitable for human visual interpretation may still contain sufficient information for a neural network.
The project, therefore, explored the possibility of moving from an image-first pipeline
Raw SAR → complete focusing → image → AI model → decision
towards a more task-driven pipeline
Raw/intermediate SAR → task-specific representation → AI model → decision
where the processing chain can potentially be shortened or adapted according to the downstream objective.
Task-driven trade-off across SAR processing levels. Earlier SAR representations potentially reduce preprocessing requirements for onboard inference, while progressively processed representations provide increasingly image-like information for downstream land-cover classification. Ongoing work is quantifying this trade-off under a common experimental protocol.
Technical outcomes of the collaboration
One of the main outcomes of the Visiting Researcher period was the establishment of an experimental framework for studying machine learning directly on early SAR representations.
The work provided practical experience with large-scale Level-0 SAR data and highlighted several considerations that are less visible when working only with conventional Level-1 imagery. These include efficient access to very large signal-domain datasets, selecting meaningful samples from early-processing representations, determining appropriate input representations for neural networks, and ensuring that comparisons across different processing levels remain methodologically meaningful.
The collaboration also reinforced the importance of evaluating onboard AI as a complete computational pipeline. Reducing the cost of a neural network alone does not necessarily yield an efficient onboard system if substantial processing must be performed first to generate its input.
Conversely, moving a machine-learning task earlier in the SAR processing chain introduces its own trade-offs. Less-processed representations may require the neural network to learn invariances or transformations that would normally be produced explicitly by the signal-processing pipeline.
The resulting research question is therefore not simply whether Level-0 classification is possible, but whether a useful accuracy–processing–deployment trade-off exists between different points in the SAR processing chain.
This represents a natural connection between the work conducted at Φ-lab and broader research on efficient, deployment-ready AI for Earth Observation.
Value of the Φ-lab collaboration
The Visiting Researcher period provided a particularly useful environment for investigating this problem by bringing together expertise in SAR processing, machine learning, and onboard computing.
The collaboration with Roberto Del Prete and the ESA Φ-lab team helped refine the problem from a conventional computer vision formulation to a question about the complete SAR processing and inference pipeline.
An important outcome was therefore both methodological and technical. The project encouraged consideration of machine-learning performance alongside signal-processing requirements, data movement, computational constraints, and the eventual onboard use case.
This perspective has subsequently influenced the direction of my PhD research at Deakin University, which focuses on efficient and deployment-aware AI for SAR and Earth Observation.
Future developments
The Visiting Researcher project established the initial framework for a broader investigation of AI directly on early SAR representations.
The next stage is to perform a systematic comparison across processing levels under controlled experimental conditions. Rather than comparing only Level-0 and fully focused data, intermediate representations can be evaluated to determine where useful semantic information emerges and where additional SAR processing provides diminishing benefit to the downstream task.
A second direction is the development of lightweight architectures specifically designed for early SAR representations. Models originally designed for natural imagery may not make optimal use of signal-domain structure. Architectures incorporating suitable inductive biases or efficient signal-processing components could potentially provide a better balance between accuracy and computational complexity.
A third direction is deployment-oriented evaluation. Future experiments will complement classification metrics with measurements such as:
- preprocessing cost;
- inference latency;
- peak memory requirements;
- model size;
- computational complexity; and
- end-to-end processing requirements.
This would allow different configurations to be compared based on the total cost of producing a semantic decision, rather than on neural-network inference in isolation.
In the longer term, the same principle can be extended beyond land-cover classification to other Earth Observation tasks such as maritime monitoring, disaster detection, and anomaly recognition.
The broader research direction is therefore towards task-driven SAR processing, in which the satellite does not necessarily reconstruct every acquisition into a single standard product before analysis. Instead, signal processing and machine learning could be jointly designed to meet the mission's information requirements.
Such an approach could contribute to future autonomous Earth Observation systems that can analyze acquisitions earlier in the processing chain, prioritize relevant observations, and transmit actionable information, while reducing unnecessary onboard computation and data movement.