4paradigm / 4paradigm/AutoX

Sample selection

未关闭
#17 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
主要语言
Jupyter Notebook
星标
551
派生
143
PR 合并指标
30 天内没有已合并 PR

描述

I would like to ask if AutoX has any plans for sample selection?

Now many data sets are so large that the computing power of individuals and small companies cannot afford.

Can a part of the data be selected for training to approximate the effect of full data training?

贡献指南

这个仓库没有索引到贡献指南

调研方向

Look at the AutoX codebase for data loading and preprocessing modules to understand the current pipeline. Investigate existing sampling techniques or if any are implemented. Determine how to integrate a sample selection feature that works with large datasets and evaluate its impact on training performance.

由索引模型根据 Issue 内容生成。

评估

技术栈
jupyter-notebook, machine-learning, python
领域
data-engineering, machine-learning
Issue 类型
功能
难度
4/5
预计耗时
3-5 天
活跃度
停滞
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。