LAION-AI / LAION-AI/Open-Assistant
Arxiv: Research papers
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 37.4k
- Forks
- 3.3k
- PR merge metrics
- No merged PRs in 30d
Description
I would like to contribute to the project by extracting data from Arxiv.
I would like to extract titles and abstracts or other metadata that might be helpul.
I think extracting the whole research paper text is not obvious as we cannot control the text length and we should extract text using OCR or other techniques.
Therefore, I would like to start with titles and abstracts extractions and I will think if we can also extract figures and tables in a futher step.
what do you think of the approach?
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
The issue does not identify files, tests, or an existing entry point. First clarify the intended Arxiv metadata scope and how the data should enter the project, then locate the relevant ingestion path; done should include a defined extraction workflow and tests for titles and abstracts.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- data
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100