AOSSIE-Org / AOSSIE-Org/EduAid

[FEATURE] Long-Document Question Generation via Token-Aware Semantic Chunking

Aperta
#560 3 commenti 0 reazioni 0 assegnatari Vedi su GitHub
enhancement
Lingua principale
JavaScript
Stelle
171
Fork
425
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

### Feature and its Use Cases

## Feature and its Use Cases

### Problem

EduAid currently generates questions by directly passing the entire `input_text` to transformer-based generators (`MCQGenerator`, `ShortQGenerator`, `BoolQGenerator`). These generators rely on T5-based models, which typically have a strict input token limit of around **512 tokens**.

When users provide long input text such as:

- lecture notes
- textbook chapters
- large paragraphs
- extracted PDF content

the transformer model silently **truncates text beyond the token limit**.

As a result:

- Only the **first portion of the document** is used for question generation.
- Important information appearing later in the document is ignored.
- Generated questions **do not cover the entire document content**.

While the `QuestionGenerator` class includes a method called `_split_into_segments()` used internally for the `/get_*_hard` endpoints, the main generation pipelines (`MCQGenerator`, `ShortQGenerator`, and `BoolQGenerator`) currently **do not implement any long-document handling mechanism**.

---

### Proposed Feature

Introduce a **token-aware semantic chunking pipeline** that enables EduAid to generate questions from long documents without losing information due to transformer token limits.

Instead of sending the entire text directly to the model, the system will:

- detect long inputs
- split them into token-safe chunks
- generate questions per chunk
- merge results into a final response

---

### Proposed Pipeline

```
Input Text

Token Length Detection

Sentence Segmentation

Token-Aware Chunk Construction

Question Generation per Chunk

Merge Generated Questions

Semantic Deduplication

Final Response
```

---

### Implementation Approach

#### 1. Token Length Detection

Use the model tokenizer (for example `T5Tokenizer`) to determine the token length of the input text.

If the token count exceeds a safe threshold (around **400–450 tokens**), trigger the chunking pipeline.

---

#### 2. Sentence Segmentation

Split the text into sentences using NLP tokenizers such as:

- `nltk.sent_tokenize()`
- spaCy sentence parser

This ensures that chunk boundaries respect sentence structure.

---

#### 3. Token-Aware Chunk Construction

Construct chunks by grouping sentences while ensuring the total token count stays within the model limit.

Example chunking strategy:

```
Chunk 1 → tokens 0–400
Chunk 2 → tokens 350–750
Chunk 3 → tokens 700–1100
```

Optional small overlaps can be added to preserve context across chunk boundaries.

---

#### 4. Chunk-wise Question Generation

Each chunk is passed independently to the existing generators:

```
MCQGenerator.generate_mcq()
ShortQGenerator.generate_shortq()
BoolQGenerator.generate_boolq()
```

This allows questions to be generated from different sections of the document.

---

#### 5. Question Aggregation

Collect questions generated from each chunk.

Example:

```
Chunk 1 → 3 questions
Chunk 2 → 3 questions
Chunk 3 → 4 questions
```

Final merged output:

```
10 questions covering the entire document
```

---

#### 6. Semantic Deduplication

Chunk-based generation may produce similar questions.

Deduplication can be implemented using semantic similarity techniques such as:

- sentence embeddings
- cosine similarity thresholds

This helps remove duplicate or near-duplicate questions.

---

#### 7. Response Formatting

Return the final merged questions using the same API response structure currently used by EduAid.

This ensures **backward compatibility with existing frontend components and endpoints**.

---

### Benefits

This enhancement will:

- Enable EduAid to process **long educational documents**
- Prevent **information loss caused by transformer token limits**
- Improve **concept coverage across entire input text**
- Maintain compatibility with existing API endpoints
- Improve the quality and usefulness of generated quizzes

---

### Example Use Case

Input:

```
A multi-page lecture note explaining photosynthesis
```

Current behavior:

```
Questions are generated only from the first paragraph of the document
```

Expected behavior after enhancement:

```
Questions generated from multiple sections such as:
- plant cell structure
- chloroplast function
- light reactions
- Calvin cycle
```

---

### Additional Context

## Additional Context

The `QuestionGenerator` class already implements token-based segmentation through:

```
_split_into_segments()
```

However, this logic is currently limited to the `/get_*_hard` endpoints and is **not used by the main generators**:

```
MCQGenerator
ShortQGenerator
BoolQGenerator
```

This feature would extend similar segmentation logic to the **primary question generation pipelines**, enabling EduAid to properly handle long-form educational content.

### Code of Conduct

- [x] I have joined the [Discord server](https://discord.gg/hjUhu33uAn) and will post updates there
- [x] I have searched existing issues to avoid duplicates

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.