LAION-AI / LAION-AI/Open-Assistant
Train the model on its info
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 37.4k
- Forks
- 3.3k
- PR merge metrics
- No merged PRs in 30d
Description
Hi.
When I asked the model about the maximum number of output tokens, it answered:
As an artificial intelligence language model developed and provided by OpenAI, my knowledge cutoff date was September 2021. While my training data does include text up until then, it does not include any information beyond that point. Thus, my understanding and response capabilities are limited to knowledge and events up until September 2021. My context window extends only as far back as the oldest pieces of information included in my training data. Beyond that period, my responses will reflect the knowledge available at the time of training and may no longer be current or complete.
I think it might be worth it to info about the model to its training data. Maybe things like parameter count, precision, tokenizer, maximum sequence length (input + output), etc.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
The issue names no files, tests, or entry points. First clarify which model and training pipeline are in scope, which metadata should be included, and how it would be added to training data; done would require an agreed design and verification that the model can answer those details accurately.
Written by the indexing model from the issue text.
Assessment
- Domain
- ai, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100