AOSSIE-Org / AOSSIE-Org/EduAid

[Enhancement]: Fix redundant model loading with Singleton ModelManager

Offen
#520 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
enhancement
Vorherrschende Sprache
JavaScript
Sterne
171
Forks
423
PR-Merge-Kennzahlen
Keine gemergten PRs in 30 T.

Beschreibung

### Feature and its Use Cases

### The Problem Context

Currently in backend/Generator/main.py, massive machine learning models and NLP tools are being initialized independently inside the __init__ methods of MCQGenerator, ShortQGenerator, and ParaphraseGenerator.

Because server.py creates instances of these classes globally, it forces Python to load t5-large (which is ~3GB) and en_core_web_sm into memory three separate times at startup.

- This consumes roughly ~9GB of RAM/VRAM just to hold redundant models.
- It makes server startup painfully slow.
- It guarantees a CUDA "Out of Memory" (OOM) crash for anyone trying to run or host the backend on a smaller GPU (e.g., 8GB VRAM).

### Proposed Solution

I propose implementing a Singleton ModelManager to act as a single source of truth for the heavy NLP models.

**Implementation Steps:**

- Create a ModelManager class that uses the __new__ method to ensure it is only ever instantiated once.
- Load t5-large, spacy, Sense2Vec, and FreqDist exactly once inside this manager.
- Refactor the __init__ methods of the generator classes so they request the existing ModelManager instance and simply hold lightweight pointers to the shared models in memory.

Here is the exact implementation I have tested:

Image

(The __init__ methods of the generator classes are then refactored to simply call manager = ModelManager() and assign self.model = manager.qg_model).

### Impact & Testing

This architectural change is purely internal plumbing and does not change the external API behavior or endpoint logic.

- Memory footprint: Drops from ~9GB down to ~3GB.
- Server startup: Drastically faster since models are read from disk only once.
- Status: I have already built and tested this locally. All endpoints return 200 OK and pass the assertions in test_server.py with identical output logic.

### Code of Conduct

- [x] I have joined the [Discord server](https://discord.gg/hjUhu33uAn) and will post updates there
- [x] I have searched existing issues to avoid duplicates

Beitragsleitfaden

Für dieses Repository ist kein Beitragsleitfaden indexiert

Bewertung

Dieses Issue wurde noch nicht bewertet.

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.