AOSSIE-Org / AOSSIE-Org/EduAid

[Enhancement]: Fix redundant model loading with Singleton ModelManager

Abierto
#520 0 comentarios 0 reacciones 0 asignados Ver en GitHub
enhancement
Lenguaje dominante
JavaScript
Estrellas
171
Forks
423
Métricas de merge de PR
Sin PR fusionados en 30 d

Descripción

### Feature and its Use Cases

### The Problem Context

Currently in backend/Generator/main.py, massive machine learning models and NLP tools are being initialized independently inside the __init__ methods of MCQGenerator, ShortQGenerator, and ParaphraseGenerator.

Because server.py creates instances of these classes globally, it forces Python to load t5-large (which is ~3GB) and en_core_web_sm into memory three separate times at startup.

- This consumes roughly ~9GB of RAM/VRAM just to hold redundant models.
- It makes server startup painfully slow.
- It guarantees a CUDA "Out of Memory" (OOM) crash for anyone trying to run or host the backend on a smaller GPU (e.g., 8GB VRAM).

### Proposed Solution

I propose implementing a Singleton ModelManager to act as a single source of truth for the heavy NLP models.

**Implementation Steps:**

- Create a ModelManager class that uses the __new__ method to ensure it is only ever instantiated once.
- Load t5-large, spacy, Sense2Vec, and FreqDist exactly once inside this manager.
- Refactor the __init__ methods of the generator classes so they request the existing ModelManager instance and simply hold lightweight pointers to the shared models in memory.

Here is the exact implementation I have tested:

Image

(The __init__ methods of the generator classes are then refactored to simply call manager = ModelManager() and assign self.model = manager.qg_model).

### Impact & Testing

This architectural change is purely internal plumbing and does not change the external API behavior or endpoint logic.

- Memory footprint: Drops from ~9GB down to ~3GB.
- Server startup: Drastically faster since models are read from disk only once.
- Status: I have already built and tested this locally. All endpoints return 200 OK and pass the assertions in test_server.py with identical output logic.

### Code of Conduct

- [x] I have joined the [Discord server](https://discord.gg/hjUhu33uAn) and will post updates there
- [x] I have searched existing issues to avoid duplicates

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.