alibaba / alibaba/graph-gpt

Custom Dataset

Open
#9 9 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
107
Forks
10
PR merge metrics
No merged PRs in 30d

Description

Thanks for your excellent work!
I would like to know how to pretrain and finetune GraphGPT in our custom dataset, and how to determine the tokenizer?
Any suggestion would be helpful.

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue asks how to pretrain and fine-tune GraphGPT on a custom dataset and determine its tokenizer, but it names no files, tests, or entry points. Start by locating the repository's training and tokenizer guidance, then document a complete custom-dataset workflow with clear completion criteria.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.