Our paper, “SFR-Net: Learning Scale-Frustum Representations for Ultra-Wide Area Remote Sensing Image Segmentation,” was officially accepted by IEEE Transactions on Geoscience and Remote Sensing (IEEE TGRS)!
Ultra-wide area segmentation
SFR-Net
Learning Scale-Frustum Representations for Ultra-Wide Area Remote Sensing Image Segmentation
See the detail. Keep the context. SFR-Net aligns local, short-range, and long-range observations around one projection reference point, then fuses them into a continuous semantic view.
Project timeline
News
We updated the codebase, fixed known bugs, improved the inference, testing, and visualization scripts, and released trained weights for GID, FBPS, and Inria Aerial.
Get weightsWe received the first-round review decision from IEEE Transactions on Geoscience and Remote Sensing (IEEE TGRS), and the manuscript was invited for major revision.
Our paper, “SFR-Net: Learning Scale-Frustum Representations for Ultra-Wide Area Remote Sensing Image Segmentation,” was released on arXiv.
Read paperOur paper, “SFR-Net: Learning Scale-Frustum Representations for Ultra-Wide Area Remote Sensing Image Segmentation,” was submitted to IEEE TGRS.
We released the initial code version with training and testing scripts and pretrained weights.
01 · Motivation
One image. Many scales.
A very long story.
Ultra-wide area images combine fine spatial detail with city-scale coverage. Patch-only models see fragments; aggressive resizing erases small objects. The hard part is preserving both.
Extreme scale variation
The same semantic class can occupy a few pixels or dominate an entire region.
Broken continuity
Independent patches lose long-range structure at roads, rivers, boundaries, and settlements.
Bounded computation
Context must grow without feeding the full multi-gigapixel scene into the network.
02 · Method
Scale-frustum representations,
fused from near to far.
SFR-Net samples aligned observations, identifies each range with learnable scale embeddings, and uses cascaded cross-scale fusion to progressively enrich the local representation.
Construct
Center every observation on the same projection reference point.
Identify
Add learnable scale embeddings after resizing observations to a common resolution.
Fuse
Cascade cross-scale context toward the local branch while preserving detail.
03 · Performance
Consistent gains across
two demanding benchmarks.
The paper reports state-of-the-art segmentation quality on GID and FBPS, with the largest gains where scale variation and semantic discontinuity matter most.
GID · mIoU
74.67+1.72paper resultGID · OA
86.94+1.09paper resultFBPS · mIoU
77.24+4.29paper resultFBPS · OA
92.91+2.40paper result04 · Qualitative results
Slide across datasets.
Explore full-scene predictions. SFR-Net improves global consistency while retaining fine structures in dense, heterogeneous regions.
05 · Interactive demo
Zoom together.
Compare pixel by pixel.
Choose one of five ultra-wide scenes, then zoom or drag either pane. The original image and SFR-Net prediction stay perfectly synchronized.
06 · Ablations
Take it apart.
See what moves the needle.
Switch between the core studies to inspect transferability, cascaded fusion, feature quality, and overlap robustness.
SFR lifts three different segmentation families.
Adding scale-frustum representations produces large mIoU gains for PSPNet, DeepLabv3+, and UperNet. The convergence view shows the advantage persists through training.
Short-range and long-range cues are complementary.
Cascaded fusion improves both datasets, while the feature maps reveal sharper edges and cleaner foreground-background separation.
Only 0.84 mIoU separates overlap and no-overlap inference.
The smallest gap among compared methods indicates that SFR-Net learns stronger internal semantic continuity across patch boundaries.
07 · Model release
Ready to reproduce.
Download pretrained backbones and released checkpoints for GID, FBPS, and Inria Aerial. Released checkpoint metrics were reproduced with random seed 42 and therefore differ slightly from the paper.
Open on Hugging FaceGID
- OA
- 86.82
- mIoU
- 74.46
- mF1
- 85.73
FBPS
- OA
- 93.50
- mIoU
- 77.86
- mF1
- 66.72
Inria Aerial
- OA
- 96.91
- IoU*
- 83.96
- F1*
- 91.28
08 · Citation
Build on SFR-Net.
If this project supports your research, please cite the paper and star the repository.
@article{zhong2026sfr,
title={SFR-Net: Learning Scale-Frustum Representations for Ultra-Wide Area Remote Sensing Image Segmentation},
author={Zhong, Chuyu and Chen, Keyan and Yang, Qinzhe and Chen, Bowen and Zou, Zhengxia and Shi, Zhenwei},
journal={arXiv preprint arXiv:2605.25737},
year={2026}
}
One more thing
Meet Phoebe.
If you find this repository helpful, please give it a star. Finally, here is Phoebe. You are not allowed to bully her.
Star SFR-Net