Réunion


Recent Advances in Remote Sensing 2026

Date : 25 Septembre 2026
Horaire : 09h00 - 17h15
Lieu : Université Paris Cité, 45, rue des Saints-Pères, 75006 Paris

Axes scientifiques :
  • Fusion, multimodalité , réseaux de capteurs, traitement multicanal
  • Apprentissage machine

Organisateurs :

Nous vous rappelons que, afin de garantir l'accès de tous les inscrits aux salles de réunion, l'inscription aux réunions est gratuite mais obligatoire.

Inscriptions

21 personnes membres du GdR IASIS, et 72 personnes non membres du GdR, sont inscrits à cette réunion.

Capacité de la salle : 1 personnes. Nombre d'inscrits en présentiel : 1 ; Nombre d'inscrits en distanciel : 92
0 Places restantes

Annonce

Les inscriptions réalisées avant le 21 juillet 2026 ont bien été prises en compte.

Les inscriptions en distanciel restent possibles. Les inscriptions en présentiel sont closes.

Les demandes de prise en charge de mission formulées avant le 21 juillet 2026, continueront à être traitées à partir du 14 septembre. Il n’est plus possible de demander une prise en charge de mission.

Recent Advances in Remote Sensing 2026 is a one-day event that takes place on September 25th 2026 in Paris. It provides an opportunity for researchers from France and neighboring countries to meet, present their work, and discuss recent advances in remote sensing. Authors of journal articles, conference papers, as well as unpublished work in the field of remote sensing — particularly PhD students — will have the opportunity to present their work during an oral presentation and/or a poster session. 

Overall, Recent Advances in Remote Sensing aims to:

  • promote sustainability and inclusivity within the remote sensing community, given that major international conferences are often held outside of Europe,
  • highlight high-quality, long-term research (possibly published in scientific journals) by giving authors the opportunity to present their work to a wide audience.

Invited speakers

  • Loïc Landrieu (Imagine, A3SI/LIGM, ENPC, IP-Paris) — AI for Forest Monitoring & Massively Multimodal Representation Learning
  • Florence Tupin (LTCI, Télécom Paris, IP-Paris) — Self-supervised learning for multi-channel SAR imaging

Organizers

  • Sylvain Lobry (LIPADE, Univ. Paris Cité)
  • Emilie Robert (CNES)
  • Ewelina Rupnik (LaSTIG, Univ. Eiffel-IGN-ENSG)
  • Romain Thoreau (MIA Paris Saclay, AgroParisTech)

Support
We thank the COMET TSI (https://www.comet-cnes.fr/tsi) for its financial support for lunch and coffee breaks.

Programme

9h - 9h30 : Accueil,

9h30 - 9h45 : Introduction

9h45 - 10h45 : Keynote - Florence Tupin - Self-supervised learning for multi-channel SAR imaging

10h45 - 11h : pause

11h - 11h30 : présentations orales

Swann Briand (ONERA, Université Paris-Saclay)

Sylvain Colomer (Université Bretagne Sud)

11h30 - 12h : présentations « flash » des posters

Guilherme Iablonovski (LASTIG, Université Gustave-Eiffel)

Thomas Hallopeau (ESPACE-DEV, IRD)

Guneet Mutreja (DLR)

Victor Sanchez (CESBIO, Université de Toulouse)

Imen Kaabachi (LIPADE, Université Paris-Cité)

12h - 14 h : déjeuner et session posters

14h - 15h : Keynote - Loic Landrieu - AI for Forest Monitoring & Massively Multimodal Representation Learning

15h - 15h30 : présentations orales

Nicolas Houdré (LIPADE, Université Paris-Cité)

Nils Foix-Colonier (LS2N, Ecole Centrale de Nantes)

15h30 - 16h45 : session posters

16h45 - 17h : clôture

Résumés des contributions

Swann Briand, Weakly supervised learning for snow cover segmentation in mountainous areas from Sentinel-1 SAR images using interpolated NDSI time series

Snow cover plays a fundamental role in climate regulation and hydrological processes. Existing snow products are based on optical imagery. Yet, snow monitoring remains challenging in mountainous regions due to frequent cloud cover. Synthetic Aperture Radar (SAR) imagery, unaffected by clouds, enables regular wet snow observations. However dry snow remains mostly transparent to SAR. In this study, we propose a fully automated framework that uses optical-derived labels to train a SAR-based model, combining the ability of optical sensors to detect both dry and

wet snow with the cloud-penetrating capability of SAR. At inference time, our method only needs SAR images to predict wet and dry snow maps. A convolutional neural network is trained to predict a binary snow cover map from a Sentinel-1 Single Look Complex (SLC) dual-pol amplitude image and a snow-free reference image. We generate binary training labels from thresholded MODIS Normalized Difference Snow Index (NDSI). Our model is trained in a weakly supervised manner by filling the cloud-induced gaps via temporal interpolation to generate pseudo-labels from sparse observations. We first evaluate the influence of the input SAR channels configuration and show that

concatenating the acquisition of the day with the reference image is preferable to more complex preprocessing. Then, we compare the Closest Neighbours Interpolation and the Kalman smoother to fill the cloud-induced gaps in the MODIS NDSI time series. We show that increasing the level of supervision improves the model performance. By removing all the gaps and the noise in the NDSI time series, the Kalman smoother yields the best model perfomance. However the regularization strength of the Kalman smoother is shown to be critical. To validate our method, we compare it to existing snow products. By comparing with the THEIA L2B Snow product, we show that our method gives comparable results to Sentinel-2 based snow cover maps. The comparison with the Copernicus Wet/Dry Snow product shows that our model can detect both wet and dry snow solely from Sentinel-1 dual-pol amplitude images. Overall, this study highlights how the snow detection capabilities of optical sensor can be transfered to SAR images thanks to a deep learning framework, to ensure a robust monitoring of wet and dry snow in cloud-prone alpine regions.



Sylvain Colomer, Large-scale individual tree segmentation across entire white spruce forests using UAV hyperspectral imagery and deep learning with ConvNeXt2



Guilherme Iablonovski , Spatially explicit feature importance for height estimation at the building level using freely accessible SAR and optical remote sensing

Accurate building height information at the individual footprint scale is essential for material stock accounting and post-disaster damage assessment, yet remains difficult to obtain at city scale in the Global South where airborne LiDAR coverage is rare and commercial very high resolution imagery is cost-prohibitive or unavailable. While recent works have demonstrated building height estimation using freely available Sentinel imagery, the resolution ceiling of resulting products was capped at 10m pixel. This study incorporates products derived from data freely accessible under scientific research licenses, TerraSAR-X StripMap and PlanetScope, alongside Sentinel-1 to predict building

heights in the medium-sized Brazilian city of Porto Alegre. To account for the significant spatial autocorrelation in building heights, features from all sources are integrated in a geographically weighted random forest mode, returning an RMSE of 5.34 m and R² of 0.756 against a LiDAR reference subset of 33,212 buildings. Each product's joint contribution to height prediction is evaluated through spatially varying local feature importance, which showed predictor dominance to vary consistently across intra-urban contexts. Footprint geometry dominates for low-rise residential buildings, shadow-derived height for taller and more isolated structures, and spectral reflectance for

the tallest buildings in the set. Sentinel-1 backscatter and TerraSAR-X InSAR-derived features occupy complementary spatial niches, with no single sensor uniformly preferable across the full building stock. Results demonstrate that the GWRF framework provides optioneering guidance and insight over satellite-derived products predictive relevance in distinct urban contexts, which global machine learning or neural network models cannot offer.



Thomas Hallopeau, Scale or Specificity? Pretraining Remote Sensing Foundation Models for Urban Environments

Although they cover only a small fraction of the Earth's surface, urban environments are home to the majority of the world's population and therefore constitute critical areas of study. Their complex spatial organization and rapid development dynamics make cities key environments for adapting to climate change. Earth observation plays an essential role in understanding these landscapes. Remote Sensing Foundation Models (RSFMs), deep learning models pretrained on satellite imagery at a global scale, provide rich and high-level representations of the Earth's surface. However, their pretraining data are largely dominated by natural and rural landscapes, such as forests, deserts, and agricultural areas. The distinctive materials and morphology of cities make urban environments fundamentally di erent from these dominant land cover types. This raises the question of whether urban remote-sensing tasks bene t more from the scale and diversity of general-purpose pretraining data or from smaller amounts of data with greater urban speci city. We investigate this trade-o by specializing an RSFM for urban environments through incremental self-supervised pretraining on increasingly speci c urban imagery. Using location metadata, we select the 30% most urban samples from the SSL4EO-S12 pretraining dataset, divide them into three subsets according to their degree of urbanization, and sequentially pretrain the model from the most to the least urban subset. We evaluate the resulting model on three urban remote-sensing tasks: Local Climate Zone classi cation

using So2Sat LCZ42, classi cation of the urban classes of EuroSAT, and informal settlement mapping using a dedicated dataset in Brazil. Our urban-specialized model performs slightly below the original globally pretrained model on So2Sat LCZ42 and EuroSAT, but outperforms it on the more ne-grained informal settlement mapping task. These results suggest that urban-focused specialization does not necessarily improve performance on broad urban land-cover classi cation, but can enhance representations for tasks that require capturing subtle intra-urban spatial and morphological patterns. Given the many factors involved in both pretraining and downstream adaptation, further work is needed to determine whether general empirical rules can be established for when urban-specic pretraining is benecial.


Guneet Mutreja, GeospatialVLM: Unified Vision-Language Reasoning and Spatial Grounding for Earth Observation

Vision-language models are creating new possibilities for interacting with Earth Observation (EO) imagery through natural language, but current approaches often separate high-level semantic understanding from precise spatial grounding and are commonly optimized for a limited set of tasks. This talk presents GeospatialVLM, a unified multimodal framework designed to support image-, region-, and pixel-level understanding within a single architecture, enabling tasks such as captioning, visual question answering, spatial reasoning, grounding, segmentation, and open-

vocabulary interpretation of geospatial scenes. A central focus of the work is the construction of geospatially rich instruction data. The training pipeline combines internally generated Esri datasets with open-source EO datasets, resulting in more than 2.6 million instruction pairs spanning captioning, VQA, grounding, segmentation, and referring tasks. Beyond conventional image-text supervision, the generated data incorporates spatial statistics and object-level relationships such as direction, distance, object size, land cover, and nearest-object context, with sampling and entropy-based filtering used to improve geographic and semantic diversity. GeospatialVLM combines a

vision encoder for global scene semantics, a grounding encoder for fine-grained spatial understanding, and a multimodal large language model. An adaptive high-resolution pipeline preserves both local details and global context across EO imagery. Text-guided attention improves query-aware visual reasoning, while Spatial Attention Prompting (SAP) supervises segmentation-token attention toward relevant object regions. Textual outputs are generated directly by the language model, whereas segmentation tokens are passed to a SAM2 decoder for dense pixel-level

prediction. A two-stage alignment and instruction-tuning strategy is used to integrate visual representations with the langage model while maintaining stable feature learning. The talk will also demonstrate how the shared vision-language embedding space can support large-scale text-to-image retrieval over EO archives, allowing users to discover relevant imagery using natural-language queries and then selectively apply tasks such as segmentation, counting, or detection. Together, these capabilities illustrate a path toward more general, interactive, and operational geospatial AI systems.



Victor Sanchez, Fusing Perceiver IO and Diffusion Model for Probabilistic Multimodal Remote Sensing

Remote Sensing applications involve a range of multimodal sensor such as optical, radar or thermal. Due to satellite properties, measurements have different spectral, spatial and temporal resolutions and are irregular and unaligned in time and space. Combining sensors increases temporal and spectral resolution and leads to higher accuracy in various application (e.g. land cover, object detection, cloud removal). This research introduces a probabilistic generative framework combining Perceiver IO with diffusion models. By encoding heterogeneous measurement into a unified latent space, Perceiver IO enables cross-sensor integration, while diffusion models provide uncertainty-

aware generation for tasks such as missing date synthesis or sensor translation. The approach leverages the strengths of both: Perceiver IO’s modality-agnostic processing and diffusion’s high-fidelity probabilistic outputs. We train our model following the ”Any-to-Any” paradigm i.e. learning any multimodal joint distribution. Preliminary results on data with various spectral and spatial resolution (Sentinel-2 10m, Sentinel-2 20m, Landsat8-15m, Landsat8-30m) demonstrate effective prediction performance while guaranteeing spectral consistency and spatial coherence. Moreover promising uncertainty quantification are observed offering a robust solution for data scarcity and heterogeneity in Earth observation.


Imen Kaabachi, Operational Benchmarking of Earth Observation Foundation Models Across Resolution Regimes

Earth observation (EO) foundation models are usually compared by their downstream performance, but this says little about the cost of adapting and running them in an operational setting. We use the PANGAEA benchmark to evaluate 13 pretrained encoders and two supervised baselines on five downstream tasks from four datasets, with ground sampling distances(GSDs) ranging from 0.1 m to 30 m. For pretrained models, we freeze the encoder and train only the downstream head. We measure runtime, peak GPU memory, floating-point operations (FLOPs) and energy use during downstream adaptatation and inference. We also estimate the corresponding operational carbon emissions. For inference, energy use and emissions are also normalized by mapped area, to enable fairer comparisons of energy and carbon efficiency across models.

We observe rankings variation across the evaluated tasks. Two pipelines can reach similar scores with very different computational costs. On the xView2 dataset, Scale-MAE reaches a mean intersection over union (mIoU) of 60.20, compared with 59.92 for U-Net, a gap of only 0.28 points. Yet, their operational profiles are quite different: Scale-MAE requires about 88% less energy for downstream adaptation, while U-Net uses about 73% less energy per square kilometer at inference. For these two models, the lower recurring inference cost of supervised UNet offsets its higher

training cost once the deployment area exceeds approximately 275 000 km². On Sen1Floods11 and HLS Burn Scars tasks, the supervised U-Net achieves the highest predictive performance while requiring less inference energy than the pipelines based on a frozen pretrained encoder. We additionally assess how sharing encoder features across multiple applications changes these comparisons. In practice, operational constraints can change which pipeline is the best choice for an EO practitioner and whether a pretrained foundation model is preferable to a task-specific supervised alternative.



Nicolas Houdré, RAMEN: Resolution-Adjustable Multimodal Encoder for Earth Observation

Earth observation (EO) data spans a wide range of spatial, spectral, and temporal resolutions, from high-resolution optical imagery to low resolution multispectral products or radar time series. While recent foundation models have improved multimodal integration for learning meaningful representations, they often expect fixed input resolutions or are based on sensor-specific encoders limiting generalization across heterogeneous EO modalities. To overcome these limitations we introduce RAMEN, a resolution-adjustable multimodal encoder that learns a shared visual

representation across EO data in a fully sensor-agnostic manner. RAMEN treats the modality and spatial and temporal resolutions as key input data features, enabling coherent analysis across modalities within a unified latent space. Its main methodological contribution is to define spatial resolution as a controllable output parameter, giving users direct control over the desired level of detail at inference and allowing explicit trade-offs between spatial precision and computational cost. We train a single, unified transformer encoder reconstructing masked multimodal EO data drawn from diverse sources, ensuring generalization across sensors and resolutions. Once pretrained, RAMEN transfers effectively to both known and unseen sensor configurations and outperforms larger state-of-the-art models on the community-standard PANGAEA benchmark, containing various multi-sensor and multi-resolution downstream tasks.



Nils Foix-Colonier, Multisolution Sparse Spectral Unmixing

Sparse spectral unmixing aims to decompose a measured spectrum into a linear combination of reference spectra, while explicitly limiting the number of nonzero components. Sparsity is particularly relevant when a large spectral library is available, e.g., built with reference spectra measured in laboratory. In the present work, we address this problem from a new perspective, aiming to enumerate the exhaustive set of sparse solutions – under an exact ℓ0-pseudonorm constraint – that are physically admissible with respect to a prescribed tolerance on the data

misfit. Contrary to the vast majority of existing methods based on optimization, which return a single solution, this multisolution approach provides a broader characterization of the solutions of interest. The proposed problem is solved using a tailored Branch-and-Bound algorithm, accelerated by a dedicated active-set algorithm for the computation of bounds. An optional minimum abundance constraint can additionally be enforced to improve the interpretability of the solutions, while considerably reducing the size of the solution set, and the numerical complexity in many cases. The proposed approach is evaluated through extensive numerical simulations. It systematically outperforms classical approaches in terms of retrieving the targeted solution, which is frequently found among the returned solution set, even when it does not correspond to the globally optimal solution. The analysis of the complete solution set also enables the extraction of higher-level information, such as the guaranteed presence or absence of some materials, or minimum and maximum admissible abundances for each component. An application to real data is

finally proposed using observations of the Europa – one of Jupiter’s moons – acquired by the Near-Infrared Mapping Spectrometer (NIMS) aboard the Galileo spacecraft. Finally, an extension of the method to hyperspectral images considering spatial coherence is discussed, using a binary total variation. The Python implementation of the multisolution Branch-and-Bound algorithm is made available.




Les commentaires sont clos.