hoangsonww / hoangsonww/Customizable-AI-Chatbot

Prod Hardening + RAG Quality: Auth, Rate Limits, Telemetry, and Evaluations

Open
#13 0 comments 0 reactions 1 assignee Claimed by @hoangsonww View on GitHub
bug documentation enhancement good first issue help wanted question
Dominant language
TypeScript
Stars
25
Forks
15
PR merge metrics
No merged PRs in 30d

Description

## Summary

Before wider adoption, let’s add a minimal “production hardening” layer and a small RAG quality harness. Goals:

* 🔐 **Auth & abuse-prevention** for API routes (LLM + Pinecone upsert/query).
* ⏳ **Rate limiting + quotas** per user / IP to protect budgets.
* 📊 **Telemetry** (request metrics, token usage, latency) + structured logs.
* 🧪 **RAG evals** (golden Q\&A set + retriever checks) to avoid regressions.
* 🧰 **Operational UX**: health checks, provider fallback, feature flags.

---

## Scope

### 1) API Security & Usage Controls

* Add optional **JWT session** (NextAuth or simple signed tokens) → gate `/api/chat` and `/api/rag/*`.
* Implement **IP + user rate limiting** (sliding window; e.g., Upstash Redis or in-memory fallback).
* Add **per-user daily quota** (messages, tokens, RAG queries). Return `429` with friendly guidance.
* CORS tightened to deployed origins.

### 2) Telemetry & Structured Logs

* Middleware to record: route, status, duration, model, **input/output token counts**, provider, cache hit/miss.
* Emit JSON logs (one line per request). Include correlation `requestId`.
* Optional exporters:

* Console (default).
* OTLP endpoint env-flag (Grafana/Datadog compatible).
* Add a minimal `/api/health` returning versions + provider readiness.

### 3) Provider Resilience

* **Provider router**: ordered fallback across OpenAI → Anthropic → Fireworks with **capability map** (streaming, JSON mode).
* Circuit-breaker: if N consecutive failures for a provider, **backoff** for T minutes.
* Configurable **model map** in `configuration/intention.ts` or `lib/providers.ts`.

### 4) RAG Quality Harness

* `evals/` folder with small **golden set** (YAML/JSON):

```yaml
- id: faq-001
query: "Who is the owner and what is their role?"
must_include:
- "Son"
- "Nguyen"
must_cite: true
```
* Scripts:

* `npm run eval:retriever` — sanity on **top-k recall** (BM25+embed optional).
* `npm run eval:rag` — ask model with retrieval → check **must\_include**, **must\_cite** (regex), **max\_tokens**.
* Output a **scorecard** (pass %, per-item failures) and save raw runs under `.evals/`.

### 5) Pinecone Hygiene & Upsert Tooling

* Validate upsert payloads (size, mime) and **chunking** with consistent metadata (`source`, `chunk_id`, `doc_id`).
* Add **namespace per environment** (`dev`, `staging`, `prod`) + env check to prevent cross-pollution.
* CLI `make upsert:path docs/` to bypass Ringel when desired; prints counts & index stats.

### 6) Feature Flags & Config

* New envs:

* `AUTH_REQUIRED=true|false`
* `RATE_LIMIT_WINDOW=60`, `RATE_LIMIT_MAX=30`
* `QUOTA_DAILY_REQUESTS=200`, `QUOTA_DAILY_TOKENS=150000`
* `PROVIDERS_ORDER=openai,anthropic,fireworks`
* `OTEL_EXPORTER_ENABLED=false`, `OTEL_EXPORTER_OTLP_ENDPOINT=...`
* `RAG_NAMESPACE=my-ai` (override per env)
* Sensible defaults so a local dev runs without extra services.

---

## Acceptance Criteria

* ✅ Protected routes return **401** when `AUTH_REQUIRED=true` and no session/token present.
* ✅ Hitting `/api/chat` > limit returns **429** with JSON: `{retryAfterSeconds, quotaHint}`.
* ✅ Logs include `{requestId, route, model, provider, ms, tokens_in, tokens_out}`.
* ✅ If primary provider fails (>=3 in 60s), router **falls back** and logs a circuit event.
* ✅ `npm run eval:rag` prints scorecard and writes `.evals//*.json`.
* ✅ Upsert path enforces chunk size & metadata; Pinecone namespace respects `NODE_ENV`.
* ✅ `/api/health` returns `{ok:true, providers:{openai:true,...}}`.

---

## Suggested Implementation Plan

**Phase 1 – Controls & Flags**

* [ ] Add `middleware.ts` for auth (NextAuth optional) + request `requestId`.
* [ ] Rate limiter util (Upstash Redis preferred; memory fallback for dev).
* [ ] Env flags + config guard.

**Phase 2 – Telemetry & Health**

* [ ] JSON logger + request timing wrapper for API routes.
* [ ] `/api/health` endpoint + provider ping.

**Phase 3 – Provider Router**

* [ ] `lib/providers.ts`: capability map, fallback chain, circuit-breaker.

**Phase 4 – RAG Evals**

* [ ] `evals/` golden set, runner scripts, scorecard printer.
* [ ] CI job: run evals on PR (smoke subset) and post summary.

**Phase 5 – Pinecone Hygiene**

* [ ] Upsert validator + CLI path; namespace by env; index stats print.

**Phase 6 – Docs & DX**

* [ ] README “Production Hardening & RAG Evals” section.
* [ ] Example `.env.local` with new flags.

---

## Notes / Nice-to-haves (later)

* Prompt-caching (Redis) for identical recent queries.
* Local embedding fallback (e.g., `text-embedding-3-small` → `all-MiniLM` dev).
* Simple **admin page** for metrics & recent 429s.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.