docling-project / docling-project/docling
How to fine-tune Docling Heron model?
Open
question
triage/close-stale
- Dominant language
- Python
- Stars
- 66.4k
- Forks
- 4.8k
- Avg merge
- 3d 4h
- Merged PRs (30d)
- 95
Description
I noticed the repo doesn't provide a general training script, so wondering where to go from here, and what the format my dataset should be in
Contributor guide
Research direction
Start by reviewing the repository's existing Docling Heron model and training-related material; the issue does not name specific files or entry points. Define what a general fine-tuning script should accept and document the expected dataset format, with completion demonstrated by a usable training path for the model.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100