facebookresearch / facebookresearch/CodeGen

Ablation on data size

Open
#66 2 comments 0 reactions 0 assignees View on GitHub
question
Dominant language
Python
Stars
777
Forks
144
PR merge metrics
No merged PRs in 30d

Description

Hi, appreciate the amazing work in unsupervised code translation!
I wonder if you have done ablation study on the training data size of TransCoder? Since the unsupervised model needs way much more training data (over 500M functions for 3 languages ) than the existing code PLMs, like CodeT5 (8.35M for 7 languages).
How's the performance of Transcoder if less data provided?

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.