hackforla / hackforla/data-science

Text Analysis Tutorial: Featurization (BoW/TF-IDF vs subword embeddings)

Đang mở
#246 0 bình luận 0 reaction 0 người được giao Xem trên GitHub
feature: missing role: missing size: missing
Ngôn ngữ chính
Jupyter Notebook
Star
33
Fork
22
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Mô tả

---

# Title & Overview

**Template:** *Featurization: An Intermediate, End-to-End Analysis Tutorial*
**Overview (≤2 sentences):** This tutorial compares classical featureization (Bag-of-Words, TF-IDF) with dense subword embeddings (byte-level BPE transformers). It is intermediate because it highlights when each approach excels, how to quantify trade-offs, and how to design experiments with reproducibility and fairness in mind.

# Purpose

The value-add is teaching learners how to **choose features defensibly** rather than by default, and how to interpret differences in speed, memory, and accuracy. They will learn to contrast sparse n-gram features with dense embeddings, log trade-offs in metrics, and connect featureization to downstream error analysis.

# Prerequisites

* Skills: Python, Git, virtual envs; pandas; ML basics (splits, regularization).
* NLP: n-grams, TF-IDF weighting, embeddings, subword tokenization.
* Tooling: pandas, scikit-learn, Hugging Face Transformers, MLflow, FastAPI.

# Setup Instructions

* Environment: Conda/Poetry (Python 3.11), deterministic seeds.
* Install: pandas, scikit-learn, Hugging Face Transformers + Datasets, MLflow, FastAPI.
* Datasets:

* **Small:** SST-2 (sentiment classification).
* **Medium:** AG News (topic classification).
* Repo layout:

```
tutorials/t2-featurization/
├─ notebooks/
├─ src/
│ ├─ vectorizers.py
│ ├─ embeddings.py
│ ├─ utils.py
│ └─ config.yaml
├─ data/README.md
├─ reports/
└─ tests/
```

# Core Concepts

* **Sparse vs dense features:** TF-IDF captures frequency & local context; embeddings capture semantics & transfer learning.
* **Byte-level BPE embeddings:** robustness to OOV, subword granularity.
* **When to use which:** small data → TF-IDF may win; large data or domain transfer → embeddings.
* **Reproducibility:** store vocab size, vectorizer config, and tokenizer merges/vocab.
* **Error slicing:** features influence robustness to typos, rare words, and domain shifts.

# Step-by-Step Walkthrough

1. **Data intake & splits:** load SST-2/AG News; stratified splits.
2. **Classical features:**

* BoW counts (word/char n-grams).
* TF-IDF (tune n-gram ranges, min\_df, max\_features).
* Classifier: Logistic Regression or Linear SVM.
3. **Subword embeddings:**

* Byte-level BPE tokenizer (GPT-2).
* Dense representations: pooled embeddings or \[CLS].
* Model: DistilBERT fine-tuned on task.
4. **Evaluation:** macro-F1 + accuracy; calibration curves; per-slice metrics (short vs long text).
5. **Error analysis:** compare confusion matrices across feature sets; highlight failure categories unique to sparse vs dense.
6. **Reporting:** metrics tables, feature importance (TF-IDF top weights), embedding nearest neighbors; save to `reports/t2-featurization.md`.
7. *(Optional)* Serve: FastAPI endpoint that allows switching featureizers via query param.

# Hands-On Exercises

* Ablations: BoW vs TF-IDF; TF-IDF n-gram ranges; embedding pooling strategies.
* Robustness: add typos/rare chars; measure performance drop.
* Hybrid features: concatenate TF-IDF with embeddings; compare against each alone.
* Stretch: dimensionality reduction (PCA/UMAP) on embeddings, visualize clusters by class.

# Common Pitfalls & Troubleshooting

* **High-dimensional TF-IDF:** memory blow-up with big n-grams; mitigate with max\_features.
* **Embedding truncation:** long inputs chopped at max seq-len → hidden errors.
* **Mismatch between vectorizer training and inference:** missing preprocessing pipeline.
* **Metrics misuse:** TF-IDF often favors accuracy, embeddings improve macro-F1.
* **Silent drift:** new tokenizer vocab alters embeddings → log and version.

# Best Practices

* Always log vocab size, n-gram ranges, and embedding tokenizer artifacts.
* Compare features under identical splits for fairness.
* Keep a **Baseline → Improvement** narrative across featureizers.
* Use unit tests: TF-IDF vocabulary under fixed corpus, embedding dimensions consistent.
* Guardrails: check max input length, enforce schema at inference.

# Reflection & Discussion Prompts

* When is TF-IDF actually better than embeddings? Why?
* How do embeddings transfer across domains compared to TF-IDF?
* What do feature importance weights vs embedding similarity reveal about interpretability?

# Next Steps / Advanced Extensions

* Pre-train domain-specific embeddings (fine-tune GPT-2 on civic corpus).
* Explore dimensionality reduction for sparse vectors.
* Build semantic search index with embeddings (FAISS).
* Apply hybrid models in downstream tasks like topic modeling.

# Glossary / Key Terms

Bag-of-Words, TF-IDF, byte-level BPE, dense embeddings, \[CLS] token, calibration, PCA/UMAP.

# Additional Resources

* [[scikit-learn](https://scikit-learn.org/stable/)](https://scikit-learn.org/stable/)
* [[Hugging Face Transformers](https://huggingface.co/docs/transformers)](https://huggingface.co/docs/transformers)
* [[Hugging Face Datasets](https://huggingface.co/datasets)](https://huggingface.co/datasets)
* [[MLflow](https://mlflow.org/)](https://mlflow.org/)
* [[FastAPI](https://fastapi.tiangolo.com/)](https://fastapi.tiangolo.com/)

# Contributors

Author(s): TBD
Reviewer(s): TBD
Maintainer(s): TBD
Date updated: 2025-09-20
Dataset licenses: SST-2 (GLUE), AG News (CC).

# Issues Referenced

Epic: HfLA Text Analysis Tutorials (T0–T14).
This sub-issue: **T2: Featurization**.

---

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Hướng nghiên cứu

Bắt đầu với bố cục được đề xuất trong tutorials/t2-featurization/, đặc biệt là các notebook, src/vectorizers.py, src/embeddings.py, src/utils.py và src/config.yaml. Xem xét workflow của SST-2 và AG News, sau đó xác định các notebook và reports/t2-featurization.md dựa trên các so sánh và số liệu được liệt kê. Được xem là hoàn thành khi tutorial bao quát phần hướng dẫn từng bước, phân tích lỗi, hướng dẫn về khả năng tái lập và các unit test được chỉ định trong tests/.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
fastapi, jupyter-notebook, pandas, python, scikit-learn
Lĩnh vực
documentation, machine-learning, testing-qa
Loại issue
Tài liệu
Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức độ hoạt động
Đình trệ
Độ rõ ràng
Đặc tả rõ ràng
Mức phù hợp với người mới
35/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.