dotnet / dotnet/machinelearning-modelbuilder

Request: Text Clustering

オープン
#2,901 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る
主要言語
Dockerfile
スター
285
フォーク
66
PR マージ指標
30日以内にマージされた PR はありません

説明

**Is your feature request related to a problem? Please describe.**
Sometimes we have a set of uncategorized texts. Either we do not care about categorization models, or we want to create a model to help categorization. Examples could be: products, credit card transaction, climate measurements, customer feedback etc.
In this request I am mainly talking about text based clustering.

**Describe the solution you'd like**
I would like to use Model Builder to cluster a set of texts, such as customer feedback.

**Describe alternatives you've considered**
Manually create the clustering code from samples. Perhaps, Model Builder could also help only with downloading sample code, or copy-paste snippets? Not ideal, but better than needing to search online.

**Additional context**
In contrast to other scenarios, this scenario would ideally also return the final clusters for each item in the input file.

This file could have multiple purposes, but considering building models, I think it is first step towards classification model.
However, if this were the only use of it, then users should use other data analysis tools. I believe the clusteration model itself will be useful for many users. We do not always need, or even can, categorize items.

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

調査の方向性

ファイル、テスト、エントリーポイントは特定されていません。まず、Model Builder の既存のシナリオと、それらがテキスト入力をどのように処理しているかを確認してください。完了とするには、各入力項目に対してクラスターを生成し、生成されたサンプルコードを含めるかどうかを明確にする、定義済みのクラスタリングワークフローが必要です。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
machine-learning
領域
machine-learning
issue の種類
機能追加
難易度
5/5
見積もり時間
1週間以上
活発さ
停滞
明瞭さ
説明が足りない
初心者へのやさしさ
25/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。