CrossTimeEdit

CrossTimeEdit

A Decade-Spanning Cross-View Dataset and Reward-Guided Editing for Historical Street-View Generation

Hanwen Lu1,* Jun He1,* Mingjia Yang1 Hao Wei1 Jinhao Huang1 Yi Lin1 Xiang Zhang1,†

1 Sun Yat-sen University

* Equal contribution    † Corresponding author

Abstract

Historical street-view imagery records urban evolution, but uneven coverage leaves substantial gaps in historical records. Generating plausible past appearances requires restoring changed structures while preserving persistent scene content. We construct VIGOR-his, a decade-spanning cross-view dataset containing 43,653 location-level quadruplets across 11 cities on three continents. Its automated pipeline performs spatial pairing, consistency screening, change classification, and the generation and validation of satellite-based change descriptions and local editing instructions. Based on VIGOR-his, we propose CrossTimeEdit, a model that reformulates historical street-view generation as editing, using recent street views to constrain viewpoint and unchanged appearance and temporal satellite differences as change evidence. Starting from FLUX.2 [Klein] 4B, we train CrossTimeEdit through supervised fine-tuning (SFT) followed by online reinforcement learning (RL). We design three street-view editing criteria, namely Instruction Alignment (IA), Background Preservation (BP), and Quality and Physical Plausibility (QP), as both RL reward dimensions and evaluation metrics. We optimize this multi-reward objective using Within Group Relative Policy Optimization for flow-matching models (Flow-GRPO) with Group reward-Decoupled Normalization Policy Optimization (GDPO), which normalizes each reward dimension before aggregation. CrossTimeEdit improves overall performance across the three editing criteria by 17.12% over the pretrained baseline and outperforms cross-view generation models in scene consistency, visual realism, and perceptual quality.

Comparison with existing cross-view datasets. VIGOR-his dataset provides four features to our CrossTimeEdit: (1) paired earlier and recent street and satellite views spanning approximately a decade; (2) acquisition-period matching and Gemini-based screening for cross-view scene consistency; (3) geographic coverage across 11 cities on three continents; and (4) change categories, satellite-based change descriptions, and local editing instructions, supporting historical street-view generation from cross-view change evidence.
Figure 1. Comparison with existing cross-view datasets. VIGOR-his dataset provides four features to our CrossTimeEdit: (1) paired earlier and recent street and satellite views spanning approximately a decade; (2) acquisition-period matching and Gemini-based screening for cross-view scene consistency; (3) geographic coverage across 11 cities on three continents; and (4) change categories, satellite-based change descriptions, and local editing instructions, supporting historical street-view generation from cross-view change evidence.

VIGOR-his

VIGOR-his provides 43,653 location-level cross-view quadruplets across 11 cities on three continents, pairing earlier and recent street and satellite views approximately a decade apart. Change categories, satellite-based change descriptions, and local editing instructions support historical street-view generation and location-level urban spatio-temporal analysis.

Table 1. Comparison with related cross-view and temporal datasets. ✓: provided; ✗: not provided in the cited release.

Table 1, cropped directly from the paper PDF. Table 1. Comparison with related cross-view and temporal datasets. ✓: provided; ✗: not provided in the cited release.
VIGOR-his dataset construction pipeline, including data preprocessing, viewpoint screening and change classification, cross-view and temporal consistency screening, satellite-based change description generation, local editing instruction generation, and instruction consistency validation. Temporal satellite differences guide instructions describing the earlier states of changed regions, while recent street views anchor viewpoint and unchanged appearance.
Figure 2. VIGOR-his dataset construction pipeline, including data preprocessing, viewpoint screening and change classification, cross-view and temporal consistency screening, satellite-based change description generation, local editing instruction generation, and instruction consistency validation. Temporal satellite differences guide instructions describing the earlier states of changed regions, while recent street views anchor viewpoint and unchanged appearance.
Stage 1

Data collection and pre-processing

Collect earlier and recent observations, pair images using panorama identifiers, capture times, and GPS coordinates, and align panoramas to a common north-facing convention.

Stage 2

Change classification and consistency screening

Screen viewpoint alignment, classify building and road changes, and retain quadruplets that pass the applicable cross-view and temporal consistency checks.

Stage 3

Satellite-grounded instruction generation

Describe temporal satellite changes along four viewing directions and translate them into local instructions specifying earlier states and unchanged content.

Stage 4

Instruction validation

Validate instructions against paired street views to check whether requested edits agree with earlier observations and preserve unchanged regions.

Representative VIGOR-his samples and their complete local editing instructions. Each sample contains a temporal cross-view quadruplet of co-located earlier/recent satellite images and street-view panoramas. From top to bottom, the examples cover no change (two locations), building changes, road changes, and combined building and road changes. The instructions organize the panorama by viewing direction and specify region-level `MODIFY' or `PRESERVE' operations.
Figure 6. Representative VIGOR-his samples and their complete local editing instructions. Each sample contains a temporal cross-view quadruplet of co-located earlier/recent satellite images and street-view panoramas. From top to bottom, the examples cover no change (two locations), building changes, road changes, and combined building and road changes. The instructions organize the panorama by viewing direction and specify region-level `MODIFY' or `PRESERVE' operations.

CrossTimeEdit

CrossTimeEdit formulates historical street-view generation as local image editing. A recent street view constrains viewpoint and unchanged appearance, while a satellite-derived instruction specifies the earlier states of changed regions. Earlier street views provide supervision and evaluation targets, but are not generation inputs.

01 / Supervised initialization

Learn editing and preservation

We train LoRA adapters on FLUX.2 [Klein] 4B using VIGOR-his. Changed samples teach historical editing, while no-change samples provide identity supervision for content preservation.

02 / Three-criterion rewards

Assess complementary requirements

Instruction Alignment (IA), Background Preservation (BP), and Quality and Physical Plausibility (QP) provide distinct RL reward dimensions and street-view editing evaluation metrics.

03 / Online reinforcement learning

Optimize multiple rewards

Flow-GRPO updates the SFT-initialized LoRA policy using VLM-scored candidates. GDPO normalizes each reward dimension independently before aggregation to account for their unequal variability.

Results

Table 2. Quantitative comparison with general image editing models on VIGOR-his.

Table 2, cropped directly from the paper PDF. Table 2. Quantitative comparison with general image editing models on VIGOR-his.

Table 3. Quantitative comparison with cross-view generation models on VIGOR-his.

Table 3, cropped directly from the paper PDF. Table 3. Quantitative comparison with cross-view generation models on VIGOR-his.

Table 4. Ablations of CrossTimeEdit under the same VLM protocol.

Table 4, cropped directly from the paper PDF. Table 4. Ablations of CrossTimeEdit under the same VLM protocol.
Qualitative comparison on five VIGOR-his cases. For readability, the satellite-based instructions shown in the figure are simplified for visualization.
Figure 3. Qualitative comparison on five VIGOR-his cases. For readability, the satellite-based instructions shown in the figure are simplified for visualization.
Qualitative ablation comparison on three matched locations. Each row shows earlier and recent satellite images, the recent street view, the earlier street-view target, and outputs from the Baseline, w/o SFT, w/o GT, w/o DGN, w/o RL, and Full variants. For readability, the satellite-based instructions shown in the figure are simplified for visualization.
Figure 7. Qualitative ablation comparison on three matched locations. Each row shows earlier and recent satellite images, the recent street view, the earlier street-view target, and outputs from the Baseline, w/o SFT, w/o GT, w/o DGN, w/o RL, and Full variants. For readability, the satellite-based instructions shown in the figure are simplified for visualization.
Qualitative comparison with cross-view generation models. Each row shows an earlier satellite image, outputs from Sat2Density and ControlS2S, the CrossTimeEdit output, and the corresponding earlier street-view ground truth.
Figure 8. Qualitative comparison with cross-view generation models. Each row shows an earlier satellite image, outputs from Sat2Density and ControlS2S, the CrossTimeEdit output, and the corresponding earlier street-view ground truth.

Resources

Citation

@article{lu2026crosstimeedit,
  title   = {{CrossTimeEdit}: A Decade-Spanning Cross-View Dataset and Reward-Guided Editing for Historical Street-View Generation},
  author  = {Lu, Hanwen and He, Jun and Yang, Mingjia and Wei, Hao and Huang, Jinhao and Lin, Yi and Zhang, Xiang},
  journal = {arXiv preprint arXiv:2609.36616},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.36616}
}