Fine-tuned SAM3TrackerModel checkpoint and supporting artefacts from the M.Sc. thesis "Bridging the Annotation Distribution Gap in Medical Imaging: A Three-Stage Pipeline for Automated Detection, Segmentation, and...
Model source
Source excerpt
Fine-tuned SAM3TrackerModel checkpoint and supporting artefacts from the M.Sc. thesis "Bridging the Annotation Distribution Gap in Medical Imaging: A Three-Stage Pipeline for Automated Detection, Segmentation, and...
Sources
1 sourceVerified Jul 31
Model artifacts
6 artifactsSource excerpts
2 excerptsgeneric-baseline/best_by_iou.pth
pth · 1.00 MB · SHA-256 94b350564add…87c9 · Hugging Face
--- license: mit language: - en library_name: pytorch tags: - medical-imaging - segmentation - inpainting - sam - sam3-tracker - annotation-removal - thesis base_model: facebook/sam2-hiera-large pipeline_tag: image-segmentation --- # Medical Annotation Removal Pipeline Fine-tuned SAM3TrackerModel checkpoint and supporting artefacts from the M.Sc. thesis **"Bridging the Annotation Distribution Gap in Medical Imaging: A Three-Stage Pipeline for Automated Detection, Segmentation, and Removal of Visual Annotations from Medical Educational Imagery."** Technical University of Munich, M.Sc. Data Engineering and Analytics, 2026. ## What this model does Given an annotated educational image (an arrow, glyph, or freeform contour drawn over an underlying object) and a type-specific geometric prompt derived from the annotation mask, the model predicts a per-instance binary segmentation of the underlying object that the annotation refers to. The predicted object mask drives Stage 3 of the full pipeline (FLUX.1 Fill inpainting), which removes the annotation while preserving the underlying scene. ## Files in this repository | File | Description | |---|---| | `best_by_iou.pth` | **Production checkpoint** — epoch 23, val IoU 0.7143. The model deployed for downstream inference. | | `best_by_loss.pth` | Alternative checkpoint — epoch 16, val loss 0.3385. Use if you prefer best-by-loss selection. | | `RUN_INFO.md` | Training metadata snapshot. | | `loss_iou_curve.png` | Training curves. | | `config.json` | Lightweight metadata for downstream loaders. | ## Architecture | Property | Value | |---|---| | Base model | SAM3TrackerModel (Hiera vision encoder, video-pretrained on SA-V) | | Total parameters | 458 M | | Frozen parameters | 454 M (vision encoder) | | Trainable parameters | **4.2 M (0.9 %)** — prompt encoder + mask decoder only | | Precision | bfloat16 | ## Training | Property | Value | |---|---| | Dataset | In-house, 9,964 source images expanded to 82,875 annotated samples | | Annotation classes | arrow (25 %), number/letter (25 %), freeform_bbox (50 %) | | Split | 80 / 10 / 10 train/val/test, stratified by source image | | Optimizer | AdamW | | Learning rate | 5e-5 | | Weight decay | 1e-4 | | Batch size | 48 | | Epochs | 30 | | Loss | 20 · focal + Dice + IoU-MSE | | Hardware | Single NVIDIA A40 (48 GB), bfloat16 | | Wall-clock | ~1...
Source context: 19 downloads · 0 likes · Pipeline image-segmentation · Library pytorch · Repo ahmed275/medical-annotation-removal