Project-MONAI / Project-MONAI/MONAI
Support for YOLO-style / modern object detection models (2D and 3D)
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 8.7k
- Forks
- 1.6k
- Avg merge
- 5d 1h
- Merged PRs (30d)
- 20
Description
Is your feature request related to a problem? Please describe.
MONAI currently ships a single object-detection architecture — RetinaNet —
under monai/apps/detection/ (retinanet_detector.py, retinanet_network.py),
with supporting anchor utilities, COCO-style mAP metrics, and box transforms. This
is a solid anchor-based, two-stage-style foundation and works in both 2D and 3D.
However, the detector zoo stops there. Users who want fast single-stage or
anchor-free detectors most notably the YOLO family — have no native path.
The existing guidance (see #903) is essentially "use MONAI transforms inside your
own YOLO pipeline," which leaves the detector itself, training loop, box-format
handling, and metrics outside MONAI's guarantees. #292 raised YOLO/COCO/Pascal-VOC
box-format support in transforms years ago but there is no dedicated tracking
issue for native modern detectors, and #8519 lists surgical instrument
localization/detection as a target without naming an architecture.
Describe the solution you'd like
Expand monai/apps/detection beyond RetinaNet to include modern detectors, with
YOLO as the flagship because of its strong fit for 2D medical / endoscopy /
surgical-tool and microscopy use cases (real-time inference, anchor-free variants,
mature ecosystem). Concretely:
- A YOLO-style detector network (e.g. an anchor-free YOLO head) integrated into
the existingDetectorNetwork/ detector API so it reuses MONAI's box
transforms, anchor/anchor-free utilities, ATSS-style matching, and mAP metrics. - Native handling of the YOLO box format (normalized cx, cy, w, h) in the
detection transforms, alongside the existing corner/CCWH conventions (follow-up
to #292). - A tutorial mirroring the existing RetinaNet LUNA16 / detection tutorial so the
new detector is a drop-in alternative. - (Optional / stretch) Room in the API for other modern detectors such as
DETR-family or FCOS, so this is an extensible "detector zoo" rather than a
one-off.
Note: interest in YOLO-inspired 3D detection
MONAI's biggest differentiator over general-purpose CV libraries is first-class
3D support — RetinaNet here already runs on volumetric data. If there is
community/maintainer interest, a YOLO-inspired 3D detector (a single-stage /
anchor-free volumetric detection head operating on 3D feature maps, predicting
6-DoF axis-aligned 3D boxes) would be a genuinely novel and high-value addition:
few libraries offer a fast single-stage 3D detector, and use cases like nodule /
lesion / landmark detection in CT and MR volumes would benefit directly. I'd be
happy to help scope and prototype this if the maintainers see value — please
comment if there's appetite for the 3D direction specifically.
Describe alternatives you've considered
- Continuing to use RetinaNet only (works, but no fast single-stage/anchor-free
option, and no YOLO ecosystem interop). - Running an external YOLO (ultralytics etc.) alongside MONAI purely for transforms
(the current #903 answer) — loses MONAI's metric/box/3D guarantees and
reproducibility.
Additional context
- Existing detection module:
monai/apps/detection/(RetinaNet). - Related prior discussions: #292 (box-format transforms), #903 (YOLO usage
question), #8519 (MICCAI submissions incl. surgical detection/localization). - Willing to contribute an implementation and tutorial, and specifically to
prototype the 3D variant if there's interest.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reading monai/apps/detection/retinanet_detector.py and retinanet_network.py, then inspect the supporting detection transforms, anchor utilities, metrics, and existing RetinaNet tutorial. Compare the current DetectorNetwork API and box conventions with the requested YOLO-style 2D and 3D scope. Done should include an agreed detector scope, API integration, box-format handling, tests, and a tutorial.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- computer-vision, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 30/100