InternLM / InternLM/JanusCoder

Release JanusCoder models and JanusCode-800K dataset on Hugging Face

Đang mở
#4 0 bình luận 2 reaction 0 người được giao Xem trên GitHub
Ngôn ngữ chính
Jupyter Notebook
Star
80
Fork
8
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Mô tả

Hi @InternLM 🤗

Niels here from the open-source team at Hugging Face. I discovered your work on Arxiv and your paper page: https://huggingface.co/papers/2510.23538.

I saw the exciting announcement in your README regarding the future release of the full JanusCoder checkpoints and the JanusCode-800K dataset, once internal company policy requirements are met. It's fantastic that you are planning to open-source these valuable artifacts!

When you are ready to release them, we'd be thrilled to help you host both the JanusCoder models and the JanusCode-800K dataset on the Hugging Face Hub (https://huggingface.co/models and https://huggingface.co/datasets). Hosting on the Hub would greatly improve their discoverability and visibility within the AI community. We can add relevant tags to the model and dataset cards to make them easily searchable and link them directly to your paper page.

## Uploading models

See here for a guide: https://huggingface.co/docs/hub/models-uploading.

In this case, we could leverage the [PyTorchModelHubMixin](https://huggingface.co/docs/huggingface_hub/package_reference/mixins#huggingface_hub.PyTorchModelHubMixin) class which adds `from_pretrained` and `push_to_hub` to any custom `nn.Module`. Alternatively, one can leverages the [hf_hub_download](https://huggingface.co/docs/huggingface_hub/en/guides/download#download-a-single-file) one-liner to download a checkpoint from the hub.

We encourage researchers to push each model checkpoint to a separate model repository, so that things like download stats also work. We can then also link the checkpoints to the paper page.

## Uploading dataset

Would be awesome to make the JanusCode-800K dataset available on 🤗 , so that people can do:

```python
from datasets import load_dataset

dataset = load_dataset("your-hf-org-or-username/your-dataset")
```
See here for a guide: https://huggingface.co/docs/datasets/loading.

Besides that, there's the [dataset viewer](https://huggingface.co/docs/hub/en/datasets-viewer) which allows people to quickly explore the first few rows of the data in the browser.

Let me know if you're interested/need any help regarding this once your artifacts are fully released!

Cheers,

Niels
ML Engineer @ HF 🤗

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Hướng nghiên cứu

Bắt đầu bằng việc xem lại thông báo trong README và các hướng dẫn liên kết của Hugging Face về việc tải mô hình lên và tải dataset. Issue không nêu rõ điểm bắt đầu triển khai hay lịch phát hành; issue chỉ được hoàn thành khi các checkpoint JanusCoder và dataset JanusCode-800K được cung cấp công khai trên Hub, tuân theo các yêu cầu chính sách đã nêu.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
huggingface
Lĩnh vực
machine-learning
Loại issue
Tính năng
Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức độ hoạt động
Đình trệ
Độ rõ ràng
Cần làm rõ
Mức phù hợp với người mới
20/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.