ORAL MICCAI 2026 · MSB EMERGE Workshop · Strasbourg, 27 Sept 2026
*Equal contribution. In collaboration with DFKI, the German Research Center for Artificial Intelligence.
Pixel-level masks are expensive to annotate on chest X-rays, where anatomy overlaps and pathological boundaries are ill-defined. Bounding boxes are far cheaper, and already exist in several datasets.
Adapts SAM3's concept grounding and localization using only boxes and concept names — no pixel masks. The pretrained mask decoder and over 95% of the vision backbone stay frozen, and the box is supplied for only a random subset of training samples, so the model learns to localize from text alone.
On a 10-class MIMIC-CXR benchmark, text-only mIoU improves over Medical-SAM3 from 0.48 to 0.69 for anatomical structures and from 0.09 to 0.53 for pathological findings — with no dense masks in training and no spatial prompt at inference.
@inproceedings{nashed2026locsam3,
title = {LocSAM3: Box-Supervised Adaptation of SAM3 for Text-Only
Chest X-Ray Segmentation},
author = {Nashed, K. and Heinein, M. and Youssef, M. and Basha, T. and
Alam, H. M. T. and Selim, A. M. and Bhatti, O. S. and Sonntag, D.},
booktitle = {MICCAI 2026 MSB EMERGE Workshop},
year = {2026},
note = {Oral presentation. Equal contribution: K. Nashed, M. Heinein, M. Youssef},
url = {https://openreview.net/forum?id=mWnH2hNROa}
}