Lightning-AI / Lightning-AI/litgpt
Full finetuning crash and model outputs are random tokens.
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 13.7k
- Forks
- 1.5k
- Avg merge
- 15h 37m
- Merged PRs (30d)
- 1
Description
I followed the tutorial to setup Alpaca data and full fine-tuning with Falcon-7B-instruct.
Here are the results comparison:
- Original Falcon-7B-Instruct:
'What is the Premiere Pro?\nPremiere Pro is a video editing software developed by Adobe. It allows users to create and edit videos, edit audio, add effects to their videos, and export them in a variety of formats. It is commonly used by professional video editors and content'
- After fine-tuning with Alpaca data (`./finetune/full.py`):
'Below is an instruction that describes a task. Write a response that appropriately completes the request.\n\n### Instruction:\nWhat is the Premiere Pro?\n\n### Response: stro con con stro con stro con stro con stro�#\x01\x01 beginningGE th toax" managementmet$\x01\x01 beginning1660\x01)$\x01\x01 beginning1660\x01 up concerningac to to �be th cweight\x01\x01 sought\x01)$\x01\x01 beginningiek0\x01 infestationightess new of ctt th c hear th cbe thren black"ireess th c31$\x01\x01 beginning$ up Christightessown c Ser"il c claims$'
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the Alpaca setup described in the tutorial and the full fine-tuning entry point at ./finetune/full.py, using Falcon-7B-instruct. Reproduce the run and compare its generated output with the examples in the issue. Done means the crash and random-token output are explained and a correction is verified.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- ai, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100