Damage-TriageFormer (primary crop-mode checkpoint)

Current revision checkpoint. This is the primary crop-mode model reported in the revised manuscript. It was trained from scratch on the corrected photo-atomic split, which keeps every tile from a source photo in one fold. It replaces the previously published tile-mode checkpoint trained on the superseded tile-level split.

Damage-TriageFormer predicts a five-class building damage typology from a single post-event RGB image. Each input is a crop centered on one footprint-derived building instance. The crop includes a 32-pixel context margin and is resized to 224×224 pixels.

Class Name
0 Undamaged
1 Partial Roof Damage
2 Total Roof Damage
3 Partial Structural Damage
4 Total Structural Collapse

Files

  • best.pth — the primary revised checkpoint, selected by validation macro F1. It is a PyTorch state_dict containing the full fine-tuned backbone and all three prediction heads.
  • config.json — architecture, preprocessing, training, split, and result metadata for this checkpoint.

Checkpoint SHA-256: 49761e951f52d7ee14eaf9f7f4e4d8a81e5daaaee04d285dd55004076bcf9430.

Architecture

The model fine-tunes a DINOv3 ViT-L/16 backbone end to end. The class-token embedding feeds three linear heads:

  1. an any-damage gate;
  2. a four-way damaged-class leaf head; and
  3. an auxiliary severity-regression head used only during training.

The five-class probability distribution is formed by combining the gate and conditional leaf probabilities. Unlike the tile-mode architecture, this checkpoint does not use a Simple Feature Pyramid or mask pooling. Building footprints are used to define the input crops before inference.

Preprocessing

  • Start from a footprint-derived building bounding box in a 1024×1024 tile.
  • Add 32 pixels of context on each side, clipped to the tile boundary.
  • Resize the crop to 224×224 pixels.
  • Convert to RGB and normalize with ImageNet mean [0.485, 0.456, 0.406] and standard deviation [0.229, 0.224, 0.225].
  • Exclude labeled footprint fragments smaller than 30 pixels.

Training

  • Split: corrected photo-atomic DamageTriage-Bench split
  • Epochs: 20
  • Optimizer: AdamW
  • Learning rate: 3e-5
  • Weight decay: 1e-4
  • Batch size: 48 on one A100 40 GB GPU
  • Schedule: cosine annealing
  • Seed: 42
  • Augmentation: random horizontal and vertical flips
  • Gate loss: binary cross-entropy
  • Leaf loss: inverse-square-root class-weighted cross-entropy, weight 2.0, label smoothing 0.1
  • Auxiliary loss: Smooth L1, weight 0.5, targets [0.0, 0.3, 0.7, 0.5, 1.0]

Results

The checkpoint was selected on the validation fold. The test fold was not used for optimization or checkpoint selection.

Metric Validation Test
Macro F1 0.646 0.616
Undamaged F1 0.925 0.919
Partial Roof Damage F1 0.539 0.563
Total Roof Damage F1 0.338 0.272
Partial Structural Damage F1 0.581 0.480
Total Structural Collapse F1 0.848 0.845

These are single-seed, within-event results. The paper reports substantially weaker transfer to held-out disasters and discusses the small support and label ambiguity of the roof-damage classes.

Usage

Use the crop-mode script in the companion code repository: github.com/YimingXiao98/Damage-TriageFormer. The script defines the exact model class and preprocessing pipeline used for this checkpoint.

git clone https://github.com/YimingXiao98/Damage-TriageFormer.git
cd Damage-TriageFormer

export DINOV3_RUNS_ROOT=/path/to/runs
mkdir -p "$DINOV3_RUNS_ROOT/crop_dtf"

hf download Ymx1025/DamageTriageFormer-model best.pth \
  --local-dir "$DINOV3_RUNS_ROOT/crop_dtf"

python scripts/train_crop.py \
  --arch dinov3_vitl16 --gated --dump-only \
  --crops /path/to/prepared/crops \
  --index-dir instance_index \
  --splits photo_splits.json \
  --batch 48 --name crop_dtf

The checkpoint contains a raw PyTorch state dictionary. Loading it requires the GatedDinoCrop architecture defined in scripts/train_crop.py.

Training data

DamageTriage-Bench (CC-BY-NC-4.0), using the released photo-atomic split.

Intended use and limitations

This release supports research and reproducibility. It is not a substitute for engineering inspection and has not been validated for autonomous habitability, safety, inspection-priority, or resource-allocation decisions. It assumes that registered building footprints are available. The three study events do not establish generalization to a new disaster, and all reported configurations use a single seed. Rare and visually ambiguous roof-damage classes remain the least reliable.

License

Released under CC-BY-NC-4.0, consistent with DamageTriage-Bench, and subject to the redistribution terms of the underlying NOAA Emergency Response Imagery and source building-footprint layers.

Citation

@misc{xiao2026damagetriageformerfoundationmodelframeworktypologybased,
      title={Damage-TriageFormer: A Foundation-Model Framework for Typology-Based Building Damage Assessment from Mono-Temporal Imagery}, 
      author={Yiming Xiao and Yu-Hsuan Ho and Sanjay Thasma and Junwei Ma and Ali Mostafavi},
      year={2026},
      eprint={2606.12248},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2606.12248}, 
}
Downloads last month
64
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Ymx1025/DamageTriageFormer-model

Dataset used to train Ymx1025/DamageTriageFormer-model

Paper for Ymx1025/DamageTriageFormer-model