← Mina Heinein

ORAL MICCAI 2026 · MSB EMERGE Workshop · Strasbourg, 27 Sept 2026

LocSAM3: Box-Supervised Adaptation of SAM3 for Text-Only Chest X-Ray Segmentation

K. Nashed*, M. Heinein*, M. Youssef*, T. Basha, H. M. T. Alam, A. M. Selim, O. S. Bhatti, D. Sonntag

*Equal contribution. In collaboration with DFKI, the German Research Center for Artificial Intelligence.

PAPER ↗ PDF ↗ CODE & SUPPLEMENTARY ↗

Abstract

PROBLEM

Pixel-level masks are expensive to annotate on chest X-rays, where anatomy overlaps and pathological boundaries are ill-defined. Bounding boxes are far cheaper, and already exist in several datasets.

APPROACH

Adapts SAM3's concept grounding and localization using only boxes and concept names — no pixel masks. The pretrained mask decoder and over 95% of the vision backbone stay frozen, and the box is supplied for only a random subset of training samples, so the model learns to localize from text alone.

RESULT

On a 10-class MIMIC-CXR benchmark, text-only mIoU improves over Medical-SAM3 from 0.48 to 0.69 for anatomical structures and from 0.09 to 0.53 for pathological findings — with no dense masks in training and no spatial prompt at inference.

Results

0.69ANATOMY mIoU · TEXT-ONLY, FROM 0.48
0.53PATHOLOGY mIoU · TEXT-ONLY, FROM 0.09
>95%OF THE VISION BACKBONE KEPT FROZEN

Cite

@inproceedings{nashed2026locsam3,
  title     = {LocSAM3: Box-Supervised Adaptation of SAM3 for Text-Only
               Chest X-Ray Segmentation},
  author    = {Nashed, K. and Heinein, M. and Youssef, M. and Basha, T. and
               Alam, H. M. T. and Selim, A. M. and Bhatti, O. S. and Sonntag, D.},
  booktitle = {MICCAI 2026 MSB EMERGE Workshop},
  year      = {2026},
  note      = {Oral presentation. Equal contribution: K. Nashed, M. Heinein, M. Youssef},
  url       = {https://openreview.net/forum?id=mWnH2hNROa}
}