[Guidance] Add workload-based sizing guidance for GPT-RAG deployments
- Dominant language
- Python
- Stars
- 1.2k
- Forks
- 321
- Avg merge
- 6h 22m
- Merged PRs (30d)
- 26
Description
### Problem
Organizations need a repeatable way to select GPT-RAG service SKUs and configurations based on expected workload. Pilot, medium-usage, and larger deployments should not require teams to independently derive every sizing decision.
### Proposed improvement
Document **Low, Medium, and High usage profiles**, with:
- Explicit assumptions for request volume, concurrency, token usage, indexed data volume, and ingestion frequency.
- Recommended service SKUs, capacity settings, and scaling options.
- Main bottlenecks and observable signals indicating when to scale up or down.
- Illustrative monthly cost estimates, including region, pricing date, and usage assumptions.
- Links to the corresponding deployment parameters and pricing calculations.
Profiles should be starting points for validation, not guaranteed performance levels. Security and availability requirements should remain explicit across profiles.
Keep component requirements and lifecycle guidance in the separate component reference issue; link to it rather than duplicating that material here.
### Expected outcome
Teams can choose an initial configuration, estimate its cost, and adjust capacity using measured workload behavior.
### Related
Azure/GPT-RAG#692 covers component requirements, optional resources, and lifecycle guidance.
Contributor guide
Research direction
No files or tests are named; begin by locating the deployment-parameter and pricing-calculation references, then review the separate component reference issue, Azure/GPT-RAG#692. Done means documenting Low, Medium, and High usage profiles with assumptions, SKU and scaling guidance, bottlenecks and signals, dated regional cost estimates, links, and explicit security and availability requirements.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- azure
- Domain
- cloud, documentation
- Issue type
- Documentation
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 58/100