TerraScope model logo Shanghai Jiao Tong University emblem

TerraScope · 天枢遥感

An Edge-Deployable Multimodal System
for Remote-Sensing Image Understanding

Hengtao Wu*, Xuanrui Cui*, Ningkai Wu*, Yanchen Li, Hao Li, Yujia Zhang, Haonan Ma, Yufei Yan, Yuting Jiang, Wenxuan Liu, Feiming Wei†, Tao Zhang†, Jin Ma†
* Equal contribution † Advisor
Shanghai Jiao Tong University

Abstract

TerraScope 是一套面向端侧部署的多模态遥感影像解译系统。系统以基于 MiniCPM-V 4.6 适配的 1.3B 视觉语言模型为核心,并发布合并后的 BF16 权重。面对单幅影像或双时相影像,用户无需添加任务前缀或结构化指令,即可完成场景识别、图像描述、视觉问答、目标计数、视觉定位与变化理解。对于像素级密集预测,系统集成了适配 LoveDA 七类地物的 SegFormer-B5 专家模型。轻量级路由层依据自然语言请求选择合适的模型,并通过统一的 Web 与 API 接口返回结果。整套系统可在离线、资源受限的环境中运行。

Highlights

  1. 01
    紧凑的遥感视觉语言模型。 TerraScope 面向航拍与卫星影像适配 MiniCPM-V 4.6,并提供合并 BF16 权重的独立 1.3B 模型。
  2. 02
    自然语言多任务交互。 同一套界面可处理单图与双时相输入,覆盖描述、识别、问答、计数、定位和变化理解。
  3. 03
    面向密集预测的专家协同。 SegFormer-B5 专家通过轻量路由和模型侧回退机制,提供指定类别及全场景语义分割。
  4. 04
    面向离线部署的工程集成。 项目包含推理、结果解析、定位与分割可视化、统一 Web/API 入口,以及可复现的运行环境。

System Overview

自然语言请求经轻量路由进入视觉语言模型或语义分割专家,随后被解析为统一的文本与可视化结果。

TerraScope system architecture
TerraScope 将任务感知的多模态模型、专业分割模型与轻量工程层组合为可离线部署的完整系统。

Qualitative Results

以下样例选自公开评测集。定位样例对比模型预测框与参考框,其余样例展示模型生成结果。

Complex Grounding XLRS-Bench
Harbor scene with predicted and reference boxes around a green boat
Prediction Reference IoU 0.665

“Locate the upper green boat in the central-right area; water is above it, a moving boat is to its left, and other green boats are below it.”

Referring Grounding VRSBench
Remote-sensing image with a grounded ship near the grassy shore
Prediction Reference IoU 1.000

“Locate the ship docked on the grassy shore near the center of the image.”

Aerial image containing four large storage tanks
Quantity VRSBench val

Simple Object Counting

“How many large storage tanks are visible in the image?”

Prediction
4
Reference
4
Exact Match
Image Description VRSBench
Residential area surrounded by vegetation with a swimming pool
Prediction

The image shows a high-resolution view of a residential area with houses surrounded by dense vegetation. A small swimming pool is located in the bottom-right of the image.

High-Resolution VQA MME-RealWorld-RS
High-resolution aerial view of Toronto

Where is the high sightseeing tower in the picture?

Prediction

B · bottom-right

High-resolution coastal scene used for route-planning VQA
Spatial Route Planning XLRS-Bench-lite

What is the best route from the large cruise terminal at the center top to the circular hotel complex at the top right?

Prediction

A. Move along the top edge to the right, then slightly downward.

Bi-Temporal Change LEVIR-CC test
Road-change scene at time A
Time A
Road-change scene at time B
Time B
Prediction

A road appears in the desert.

Bi-Temporal Change LEVIR-CC test
Building-change scene at time A
Time A
Building-change scene at time B
Time B
Prediction

A large building is built on the right side of the scene.

Task Demos

同一套自然语言交互界面支持六类典型遥感任务;GIF 将自动播放,点击可查看原始尺寸。

Remote-sensing scene classification demo
Scene Classification根据单幅影像识别遥感场景类别。
Remote-sensing image description demo
Image Description生成简洁且符合任务需求的遥感场景描述。
Object counting demo
QuantityObject Counting · 回答自然语言计数问题。
Visual grounding demo
Visual Grounding返回归一化坐标,并绘制目标预测框。
Bi-temporal change understanding demo
Change Understanding比较两个时相并描述其中发生的变化。
Semantic segmentation demo
Semantic Segmentation通过 SegFormer-B5 生成指定类别或全场景掩膜。

Models & Method

Stage 1 remote-sensing foundation model architecture
Stage 1 · Remote-Sensing Adaptation
Stage 2 high-precision localization enhancement architecture
Stage 2 · Localization Enhancement

Stage 2 · Method Details

Visual adaptation, Top-5 region selection, local grounding, coordinate mapping, and residual RL selection
Grounding Pipeline
Detail of pipeline step 4: residual RL candidate selection and training feedback
Pipeline Step 4 · Residual RL Selection

Downloads

M1
TerraScope-1.3BVision-language model

基于 MiniCPM-V 4.6 适配的遥感影像理解模型,提供合并后的 BF16 权重。

ModelScope
M2
TerraScope SegmentationSemantic-segmentation expert

面向 LoveDA 适配的 SegFormer-B5 模型,支持七类城乡地表覆盖分割。

ModelScope

Citation

@software{terrascope2026,
  title  = {TerraScope: An Edge-Deployable Multimodal System for
            Remote-Sensing Image Understanding},
  author = {Wu, Hengtao and Cui, Xuanrui and Wu, Ningkai and
            Li, Yanchen and Li, Hao and Zhang, Yujia and Ma, Haonan and
            Yan, Yufei and Jiang, Yuting and Liu, Wenxuan and Wei, Feiming and
            Zhang, Tao and Ma, Jin},
  year   = {2026},
  url    = {https://github.com/EternalWavee/challenge-cup-2026}
}