PanoSAMic is a multi-modal semantic segmentation model for panoramic (360°) images. It integrates the frozen Segment Anything Model (SAM) encoder, modified to output multi-stage features, with a spatio-modal fusion...
Modellquelle
Quellenbeschreibung
PanoSAMic is a multi-modal semantic segmentation model for panoramic (360°) images. It integrates the frozen Segment Anything Model (SAM) encoder, modified to output multi-stage features, with a spatio-modal fusion module (MCBAM), a spherical-attention semantic decoder, and dual-view fusion to handle the distortion and edge discontinuity of equirectangular images.
Quellen
1 QuelleVerifiziert 1. Aug.
Modellartefakte
17 Artefaktematterport3d-vith-rgb/model.safetensors
safetensors · 753 MB · SHA-256 1518341bc181…8751 · Hugging Face
Herunterladenmatterport3d-vith-rgbd/model.safetensors
safetensors · 1.003 MB · SHA-256 9e5791bd27bd…f09d · Hugging Face
HerunterladenQuellenauszüge
2 AuszügeOnly the trainable PanoSAMic components are hosted here:
The full model state dict has two parts:
| Module prefix | Trainable | In Hub checkpoint |
|---|---|---|
feature_fuser.* | ✅ yes | ✅ yes |
semantic_decoder.* | ✅ yes | ✅ yes |
image_encoder.* | ❌ frozen (SAM ViT) | ❌ no |
prompt_encoder.* | ❌ frozen (SAM) | ❌ no |
mask_decoder.* | ❌ frozen (SAM) | ❌ no |
The frozen SAM ViT backbone is NOT hosted here. It is downloaded separately from Meta's official release (Apache-2.0) and combined at load time. This keeps each checkpoint small and avoids redistributing the SAM weights.
Each variant lives in its own subfolder of dfki-av/PanoSAMic
(e.g. stanford2d3ds-vith-rgbdn-fold1/model.safetensors).
3-fold checkpoints are published per fold so each can be evaluated on its held-out split.
| Checkpoint | Backbone | Modalities | Dataset | Split |
|---|---|---|---|---|
stanford2d3ds-vith-rgb-fold1 | ViT-H | RGB | Stanford2D3DS | Fold 1 |
stanford2d3ds-vith-rgb-fold2 | ViT-H | RGB | Stanford2D3DS | Fold 2 |
stanford2d3ds-vith-rgb-fold3 | ViT-H | RGB | Stanford2D3DS | Fold 3 |
stanford2d3ds-vith-rgbd-fold1 | ViT-H | RGB-D | Stanford2D3DS | Fold 1 |
stanford2d3ds-vith-rgbd-fold2 | ViT-H | RGB-D | Stanford2D3DS | Fold 2 |
Stanford2D3DS (3-fold validation), main table:
| Checkpoint | mIoU % | mAcc % | Trainable params (M) |
|---|---|---|---|
stanford2d3ds-vith-rgb | 59.62 | 74.11 | 178 |
stanford2d3ds-vith-rgbd | 60.90 | 73.95 | 184 |
stanford2d3ds-vith-rgbdn | 61.57 | 74.04 | 191 |
Encoder-size study (Stanford2D3DS, 3-fold, RGB-D-N):
| Checkpoint | mIoU % | mAcc % |
|---|---|---|
stanford2d3ds-vitb-rgbdn | 56.68 | 70.49 |
stanford2d3ds-vitl-rgbdn | 60.90 | 73.09 |
stanford2d3ds-vith-rgbdn | 61.57 | 74.04 |
Matterport3D (BEV360 splits):
| Checkpoint | mIoU % |
|---|---|
matterport3d-vith-rgb | 46.59 |
matterport3d-vith-rgbd | 48.43 |
ToF-360 (zero-shot, fold 1, instance-refined):
No checkpoint is trained on ToF-360 (only 179 real-world samples across 4 scenes — too little to train on). Instead, the Stanford2D3DS ViT-H fold-1 checkpoints are evaluated directly against it (single split, not averaged across folds), with ToF-360's own label ontology remapped to Stanford2D3DS's 13 classes at load time and SAM instance masks used to refine the semantic predictions. Adding depth/normals hurts zero-shot transfer here, likely due to the domain gap between ToF-360's real sensor depth and Stanford2D3DS's.
| Checkpoint | Modalities | mIoU % | mAcc % |
|---|---|---|---|
stanford2d3ds-vith-rgb-fold1 | RGB | 41.30 | 63.72 |
stanford2d3ds-vith-rgbd-fold1 | RGB-D | 37.33 | 60.05 |
stanford2d3ds-vith-rgbdn-fold1 | RGB-D-N | 36.05 | 59.06 |
uv sync from the GitHub repo (pyproject.toml pins dependencies)Download the official SAM weights from Meta and place them in sam_weights/:
sam_vit_h_4b8939.pthsam_vit_l_0b3195.pthsam_vit_b_01ec64.pthfrom panosamic.model import PanoSAMic
model = PanoSAMic.from_pretrained_panosamic(
"dfki-av/PanoSAMic",
subfolder="stanford2d3ds-vith-rgbdn-fold1",
config_path="config/config_stanford2d3ds_dv.json",
vit_model="vit_h",
modalities=("image", "depth", "normals"),
num_classes=13,
sam_weights_path="./sam_weights", # omit to auto-download from Meta's servers
)
from_pretrained_panosamic loads only the trainable weights from the Hub,
initialises the frozen SAM backbone from the local sam_weights/ directory
(auto-downloaded if not present), and returns the model in eval() mode.
import torch
from panosamic.model.instance_semantic_fusion import refine_semantic_with_instances
# batched_input: list of dicts, one per image.
# Each dict maps modality name → float tensor (3, H, W), values in [0, 255].
# Image must be equirectangular 2:1 (e.g. 512 × 1024).
batched_input = [{"image": image_tensor, "depth": depth_tensor, "normals": normals_tensor}]
with torch.no_grad():
outputs = model(batched_input)
sem_preds = outputs[0]["sem_preds"] # (num_classes, H, W) — logits
instance_masks = outputs[0]["instance_masks"]
# Instance-guided refinement: each SAM mask is assigned the majority
# semantic class within it, sharpening boundaries.
if instance_masks:
sem_preds = refine_semantic_with_instances(sem_preds, instance_masks)
seg_map = sem_preds.argmax(dim=0) # (H, W) — integer class indices
Use the exact splits reported in the paper:
panosamic/data_preparation/ into the processed structure documented in the
repo README.From a released Hub checkpoint (trainable weights only, SAM loaded separately):
python panosamic/evaluation/evaluate.py \
--dataset_path /path/to/processed/dataset \
--config_path config/config_stanford2d3ds_dv.json \
--checkpoint dfki-av/PanoSAMic \
--subfolder stanford2d3ds-vith-rgbdn-fold1 \
--sam_weights_path ./sam_weights \
--dataset stanford2d3ds \
--fold 1 \
--vit_model vit_h \
--modalities image,depth,normals \
--num_gpus 1
From a local training run (full checkpoint including frozen backbone):
python panosamic/evaluation/evaluate.py \
--dataset_path /path/to/processed/dataset \
--config_path config/config_stanford2d3ds_dv.json \
--experiments_path ./experiments \
--dataset stanford2d3ds \
--fold 1 \
--vit_model vit_h \
--modalities image,depth,normals \
--num_gpus 1
Repeat for folds 1–3 and average for the 3-fold numbers. For Matterport3D use
config/config_matterport3d_dv.json, --dataset matterport3d, and the
modalities for that row.
Zero-shot on ToF-360 (no fine-tuning, fold-1 Stanford2D3DS checkpoint):
python panosamic/evaluation/evaluate.py \
--dataset_path /path/to/processed/tof360 \
--config_path config/config_tof360_dv.json \
--checkpoint dfki-av/PanoSAMic \
--subfolder stanford2d3ds-vith-rgb-fold1 \
--sam_weights_path ./sam_weights \
--dataset tof360 \
--fold 1 \
--vit_model vit_h \
--modalities image \
--num_gpus 1
Swap --subfolder/--modalities (stanford2d3ds-vith-rgbd-fold1 +
image,depth, stanford2d3ds-vith-rgbdn-fold1 + image,depth,normals) to
reproduce the other rows above. Unlike the Stanford2D3DS/Matterport3D tables,
ToF-360 numbers are fold-1 only, not averaged across folds.
Indoor panoramic semantic segmentation with RGB / RGB-D / RGB-D-N input. Evaluated only on indoor datasets; outdoor generalization is not guaranteed.
@article{chamseddine2026panosamic,
title = {PanoSAMic: Panoramic Image Segmentation from SAM Feature Encoding and Dual View Fusion},
author = {Chamseddine, Mahdi and Stricker, Didier and Rambach, Jason},
journal = {arXiv preprint arXiv:2601.07447},
year = {2026}
}
Funded by the European Union as part of the projects HumanTech (Grant Agreement 101058236) and ShieldBOT (Grant Agreement 101235093).
stanford2d3ds-vitb-rgbdn-fold1/model.safetensors
safetensors · 791 MB · SHA-256 aae73775e3a7…873e · Hugging Face
stanford2d3ds-vitb-rgbdn-fold2/model.safetensors
safetensors · 791 MB · SHA-256 5112a6d5a975…35fd · Hugging Face
Herunterladenstanford2d3ds-vitb-rgbdn-fold3/model.safetensors
safetensors · 791 MB · SHA-256 8192553a729b…d01f · Hugging Face
Herunterladenstanford2d3ds-vith-rgb-fold1/model.safetensors
safetensors · 753 MB · SHA-256 d18ad880894f…34ac · Hugging Face
Herunterladenstanford2d3ds-vith-rgb-fold2/model.safetensors
safetensors · 753 MB · SHA-256 f370849047cb…d79e · Hugging Face
Herunterladenstanford2d3ds-vith-rgb-fold3/model.safetensors
safetensors · 753 MB · SHA-256 464855275815…bd3c · Hugging Face
Herunterladenstanford2d3ds-vith-rgbd-fold1/model.safetensors
safetensors · 1.003 MB · SHA-256 d6c3ce3291f3…9d88 · Hugging Face
Herunterladenstanford2d3ds-vith-rgbd-fold2/model.safetensors
safetensors · 1.003 MB · SHA-256 7ec8107249b4…7683 · Hugging Face
Herunterladenstanford2d3ds-vith-rgbd-fold3/model.safetensors
safetensors · 1.003 MB · SHA-256 ab5275afbdf8…131c · Hugging Face
Herunterladenstanford2d3ds-vith-rgbdn-fold1/model.safetensors
safetensors · 1,37 GB · SHA-256 f3d8d0e9b880…cd90 · Hugging Face
Herunterladenstanford2d3ds-vith-rgbdn-fold2/model.safetensors
safetensors · 1,37 GB · SHA-256 13804b146116…b613 · Hugging Face
Herunterladenstanford2d3ds-vith-rgbdn-fold3/model.safetensors
safetensors · 1,37 GB · SHA-256 8275f14d2062…2be3 · Hugging Face
Herunterladenstanford2d3ds-vitl-rgbdn-fold1/model.safetensors
safetensors · 1,04 GB · SHA-256 4347dcc3f826…de87 · Hugging Face
Herunterladenstanford2d3ds-vitl-rgbdn-fold2/model.safetensors
safetensors · 1,04 GB · SHA-256 e727842dc895…1f07 · Hugging Face
Herunterladenstanford2d3ds-vitl-rgbdn-fold3/model.safetensors
safetensors · 1,04 GB · SHA-256 3a96675a4b0b…5ddd · Hugging Face
Herunterladen--- license: cc-by-nc-sa-4.0 tags: - semantic-segmentation - panoramic - spherical-images - modality-fusion - sam library_name: panosamic pipeline_tag: image-segmentation datasets: - stanford2d3ds - matterport3d - tof360 --- # PanoSAMic PanoSAMic is a multi-modal semantic segmentation model for panoramic (360°) images. It integrates the **frozen** Segment Anything Model (SAM) encoder, modified to output multi-stage features, with a spatio-modal fusion module (MCBAM), a spherical-attention semantic decoder, and dual-view fusion to handle the distortion and edge discontinuity of equirectangular images. - **Paper:** PanoSAMic: Panoramic Image Segmentation from SAM Feature Encoding and Dual View Fusion (ICPR 2026) - **Code:** https://github.com/dfki-av/PanoSAMic - **arXiv:** https://arxiv.org/abs/2601.07447 - **Authors:** Mahdi Chamseddine, Didier Stricker, Jason Rambach (DFKI / RPTU Kaiserslautern-Landau) ## What is in this repository Only the **trainable** PanoSAMic components are hosted here: - **Feature fusion blocks (MCBAM)** — spatio-modal cross-attention applied to the branch features extracted by the frozen encoder - **Semantic decoder** — convolutional decoder with spherical attention and dual-view fusion head The full model state dict has two parts: | Module prefix | Trainable | In Hub checkpoint | |---|---|---| | `feature_fuser.*` | ✅ yes | ✅ yes | | `semantic_decoder.*` | ✅ yes | ✅ yes | | `image_encoder.*` | ❌ frozen (SAM ViT) | ❌ no | | `prompt_encoder.*` | ❌ frozen (SAM) | ❌ no | | `mask_decoder.*` | ❌ frozen (SAM) | ❌ no | The **frozen SAM ViT backbone is NOT hosted here.** It is downloaded separately from Meta's official release (Apache-2.0) and combined at load time. This keeps each checkpoint small and avoids redistributing the SAM weights. ## Available checkpoints Each variant lives in its own subfolder of `dfki-av/PanoSAMic` (e.g. `stanford2d3ds-vith-rgbdn-fold1/model.safetensors`). 3-fold checkpoints are published per fold so each can be evaluated on its held-out split. | Checkpoint | Backbone | Modalities | Dataset | Split | |---|---|---|---|---| | `stanford2d3ds-vith-rgb-fold1` | ViT-H | RGB | Stanford2D3DS | Fold 1 | | `stanford2d3ds-vith-rgb-fold2` | ViT-H | RGB | Stanford2D3DS | Fold 2 | | `stanford2d3ds-vith-rgb-fold3` | ViT-H | RGB | Stanford2D3DS | Fold 3 | | `stanford2d3ds-vith-rgbd-fold1` | ViT-H | RGB-D | S...
Source context: 0 downloads · 1 likes · Pipeline image-segmentation · Library panosamic · Repo dfki-av/PanoSAMic
stanford2d3ds-vith-rgbd-fold3 | ViT-H | RGB-D | Stanford2D3DS | Fold 3 |
stanford2d3ds-vith-rgbdn-fold1 | ViT-H | RGB-D-N | Stanford2D3DS | Fold 1 |
stanford2d3ds-vith-rgbdn-fold2 | ViT-H | RGB-D-N | Stanford2D3DS | Fold 2 |
stanford2d3ds-vith-rgbdn-fold3 | ViT-H | RGB-D-N | Stanford2D3DS | Fold 3 |
stanford2d3ds-vitl-rgbdn-fold1 | ViT-L | RGB-D-N | Stanford2D3DS | Fold 1 |
stanford2d3ds-vitl-rgbdn-fold2 | ViT-L | RGB-D-N | Stanford2D3DS | Fold 2 |
stanford2d3ds-vitl-rgbdn-fold3 | ViT-L | RGB-D-N | Stanford2D3DS | Fold 3 |
stanford2d3ds-vitb-rgbdn-fold1 | ViT-B | RGB-D-N | Stanford2D3DS | Fold 1 |
stanford2d3ds-vitb-rgbdn-fold2 | ViT-B | RGB-D-N | Stanford2D3DS | Fold 2 |
stanford2d3ds-vitb-rgbdn-fold3 | ViT-B | RGB-D-N | Stanford2D3DS | Fold 3 |
matterport3d-vith-rgb | ViT-H | RGB | Matterport3D | BEV360 |
matterport3d-vith-rgbd | ViT-H | RGB-D | Matterport3D | BEV360 |