ShareGPT conversation splits and "please continue"
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 39.5k
- Forks
- 4.8k
- PR merge metrics
- No merged PRs in 30d
Description
Take a look at [this conversation](https://gist.github.com/float-trip/05201a4c60fb7b785550aff97af20d6e), along with [the processed version](https://gist.github.com/float-trip/ff67036fd21531264a77fb64e2f2890e) from `split_long_conversation.py`.
Post-split, we have training data that looks like this:
```
{
"from": "gpt",
"value": "Here's an example of how to create a JAX-WS service and deploy the generated code using Wildfly.\n\n1. Create a new Java project in your favorite IDE (e.g., Eclipse, IntelliJ IDEA).\n2. Add the generated code to your project's source folder.\n3. Modify the `@WebService` annotation in the generated code to specify the correct endpoint interface and port name.\n4. Create a new JAX-WS service class that extends `javax.xml.ws.Service`.\n5. In the service class constructor, call the super constructor with the WSDL location and service name.\n6. Implement the service endpoint interface by overriding its methods.\n7. Deploy the service to Wildfly.\n\nHere's some sample code to get you started:\n```java\npackage com.example;\n\nimport java.util.List;\n\nimport javax.jws.WebService;\nimport javax.xml.ws.BindingType;\nimport javax.xml.ws.Endpoint;\nimport javax.xml.ws.soap.SOAPBinding;\n\nimport com.example.generated.Init;\nimport com.example.generated.InitResponse;\nimport com.example.generated.Seat;\nimport com.example.generated.SeatReservationImpl;\nimport com.example.generated.SeatReservationImplService;\n\n@WebService(endpointInterface = \"com.example.generated.SeatReservation\", portName = \"SeatReservationPort\")\n@BindingType(value = SOAPBinding.SOAP12HTTP_BINDING)\npublic class SeatReservationWebService implements com.example.generated.SeatReservation {\n\n private SeatReservationImpl serviceImpl;\n\n public SeatReservationWebService() {\n SeatReservationImplService service = new SeatReservationImplService(getClass().getResource(\"SeatReservation.wsdl\"));\n serviceImpl = service.getSeatReservationImplPort();\n }\n\n @Override\n public InitResponse init(Init parameters) {\n return serviceImpl.init(parameters);\n }\n\n @Override\n public List getAllSeats() {\n return serviceImpl.getAllSeats();\n }\n\n @Override\n public Seat getSeatStatus(int seatId) {\n return serviceImpl.getSeatStatus(seatId);\n }\n\n @Override\n public boolean lock(int seatId) {\n return serviceImpl.lock(seatId);\n }\n\n @Override\n public boolean unlock(int seatId) {\n return serviceImpl.unlock(seatId);\n }\n\n @Override\n public boolean reserve(int seatId) {\n return serviceImpl.reserve(seatId);\n }\n\n public static void main(String[] args) {\n SeatReservationWebService service = new SeatReservationWebService();\n Endpoint.publish(\"http://localhost:8080/SeatReservation\", service);\n }\n}\n```\nYou can deploy this service to Wildfly by building a WAR file and copying it to the `deployments` directory of your Wildfly installation. Alternatively, you can use the Wildfly Maven plugin to deploy the WAR file automatically. Here's an example `pom.xml` file that includes the necessary dependencies and configuration:\n```php\n\n com.example\n SeatReservation\n 1.0-SNAPSHOT\n\n \n \n com.sun.xml.ws\n jaxws-rt\n 2.3.5\n \n \n\n \n \n \n org.wildfly.plugins\n wildfly-maven-plugin\n 3.0.1.Final\n \n ${jboss.host}\n ${jboss.http.port}\n \n \n \n \n\n```\nYou"
}
]
}
{
"id": "wNBG8Gp_18",
"conversations": [
{
"from": "human",
"value": "continue"
},
```
GPT cuts itself off in its last message, so we probably don't want to train on that. @teknium1 [mentioned](https://github.com/PhoebusSi/Alpaca-CoT/issues/109#issuecomment-1538252439) the model has early stopping issues and I think this could be one contributor. This could be helped by looking for prompts like "please continue the above", and then dropping the last GPT response along with the rest of the conversation.
Moreover, training on a conversation that starts with "continue" seems odd. I think it'd make sense to ensure each split starts off with the last message from GPT. The old version of `split_long_conversation.py` might've done this(?), but the current one doesn't. For example -
[1] Human: ... (1500 tokens so far)
[2] GPT: ... (1800)
[3] Human: ... (2300, > 2048)
[4] GPT: ... (2800)
This could be split into two conversations, one with messages [1, 2], the other with messages [2, 3, 4]. Some data will be repeated during training, but by starting with a message from GPT, the model can learn that it's been dropped into the middle of a conversation and is missing context.
Also worth considering that the splits are currently divided very neatly, and won't perfectly emulate how LLaMA will see long conversations in practice. Naive question: why split at all? When processing something like a book, how do LLaMA and other GPT models normally handle long sequences? Is there a stride length of something like 512 tokens, or do they advance by the full context window?
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with split_long_conversation.py and compare its output with the linked original and processed conversations. Check how splits handle truncated final GPT messages, "continue" prompts, and the first message of each split. Done means the script's behavior matches an agreed policy for these cases, with the example conversation no longer producing unwanted training records.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100