AOSSIE-Org / AOSSIE-Org/EduAid

[Enhancement]: Fix redundant model loading with Singleton ModelManager

未關閉
#520 0 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
enhancement
主要語言
JavaScript
星號
171
分支
423
PR 合併指標
30 天內沒有已合併 PR

描述

### Feature and its Use Cases

### The Problem Context

Currently in backend/Generator/main.py, massive machine learning models and NLP tools are being initialized independently inside the __init__ methods of MCQGenerator, ShortQGenerator, and ParaphraseGenerator.

Because server.py creates instances of these classes globally, it forces Python to load t5-large (which is ~3GB) and en_core_web_sm into memory three separate times at startup.

- This consumes roughly ~9GB of RAM/VRAM just to hold redundant models.
- It makes server startup painfully slow.
- It guarantees a CUDA "Out of Memory" (OOM) crash for anyone trying to run or host the backend on a smaller GPU (e.g., 8GB VRAM).

### Proposed Solution

I propose implementing a Singleton ModelManager to act as a single source of truth for the heavy NLP models.

**Implementation Steps:**

- Create a ModelManager class that uses the __new__ method to ensure it is only ever instantiated once.
- Load t5-large, spacy, Sense2Vec, and FreqDist exactly once inside this manager.
- Refactor the __init__ methods of the generator classes so they request the existing ModelManager instance and simply hold lightweight pointers to the shared models in memory.

Here is the exact implementation I have tested:

Image

(The __init__ methods of the generator classes are then refactored to simply call manager = ModelManager() and assign self.model = manager.qg_model).

### Impact & Testing

This architectural change is purely internal plumbing and does not change the external API behavior or endpoint logic.

- Memory footprint: Drops from ~9GB down to ~3GB.
- Server startup: Drastically faster since models are read from disk only once.
- Status: I have already built and tested this locally. All endpoints return 200 OK and pass the assertions in test_server.py with identical output logic.

### Code of Conduct

- [x] I have joined the [Discord server](https://discord.gg/hjUhu33uAn) and will post updates there
- [x] I have searched existing issues to avoid duplicates

貢獻指南

這個儲存庫沒有索引到貢獻指南

評估

這個 Issue 還沒有評估資料。

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。