AOSSIE-Org / AOSSIE-Org/EduAid

[Enhancement]: Fix redundant model loading with Singleton ModelManager

Ouverte
#520 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub
enhancement
Langage dominant
JavaScript
Étoiles
171
Forks
423
Métriques de merge des PR
Aucune PR mergée en 30 j

Description

### Feature and its Use Cases

### The Problem Context

Currently in backend/Generator/main.py, massive machine learning models and NLP tools are being initialized independently inside the __init__ methods of MCQGenerator, ShortQGenerator, and ParaphraseGenerator.

Because server.py creates instances of these classes globally, it forces Python to load t5-large (which is ~3GB) and en_core_web_sm into memory three separate times at startup.

- This consumes roughly ~9GB of RAM/VRAM just to hold redundant models.
- It makes server startup painfully slow.
- It guarantees a CUDA "Out of Memory" (OOM) crash for anyone trying to run or host the backend on a smaller GPU (e.g., 8GB VRAM).

### Proposed Solution

I propose implementing a Singleton ModelManager to act as a single source of truth for the heavy NLP models.

**Implementation Steps:**

- Create a ModelManager class that uses the __new__ method to ensure it is only ever instantiated once.
- Load t5-large, spacy, Sense2Vec, and FreqDist exactly once inside this manager.
- Refactor the __init__ methods of the generator classes so they request the existing ModelManager instance and simply hold lightweight pointers to the shared models in memory.

Here is the exact implementation I have tested:

Image

(The __init__ methods of the generator classes are then refactored to simply call manager = ModelManager() and assign self.model = manager.qg_model).

### Impact & Testing

This architectural change is purely internal plumbing and does not change the external API behavior or endpoint logic.

- Memory footprint: Drops from ~9GB down to ~3GB.
- Server startup: Drastically faster since models are read from disk only once.
- Status: I have already built and tested this locally. All endpoints return 200 OK and pass the assertions in test_server.py with identical output logic.

### Code of Conduct

- [x] I have joined the [Discord server](https://discord.gg/hjUhu33uAn) and will post updates there
- [x] I have searched existing issues to avoid duplicates

Guide de contribution

Aucun guide de contribution indexé pour ce dépôt

Évaluation

Cette issue n'a pas encore été évaluée.

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.