NVIDIA / NVIDIA/NeMo-Retriever
[DOC]: add an architecture diagram for the pipeline including the NIMs
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 3k
- Forks
- 349
- Avg merge
- 1d 23h
- Merged PRs (30d)
- 116
Description
How would you describe the priority of this documentation request
None
Please provide a link or source to the relevant docs
https://github.com/NVIDIA/nv-ingest/blob/main/README.md
Describe the problems in the documentation
Please include an architecture diagram accompanied with some explanation of what the NIM services do and how they interact. It isn’t clear if the pipeline can be adjusted to use a subset of the services. The docs here show that deplot, yolox and PaddleOCR are used for chart extraction, it isn’t clear whether you can configure which service to use (ie. if they serve the same purpose or if they need to work in conjunction for a single extraction task).
(Optional) Propose a correction or improvement
No response
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with README.md and review the documented extraction services, including Deplot, YOLOX, and PaddleOCR. Done means adding an architecture diagram with an accompanying explanation of NIM roles, interactions, and whether a subset of services can be configured.
Written by the indexing model from the issue text.
Assessment
- Domain
- documentation
- Issue type
- Documentation
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 30/100