Azure / Azure/GPT-RAG

[Guidance] Add workload-based sizing guidance for GPT-RAG deployments

Open
#693 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.2k
Forks
321
Avg merge
6h 22m
Merged PRs (30d)
26

Description

### Problem

Organizations need a repeatable way to select GPT-RAG service SKUs and configurations based on expected workload. Pilot, medium-usage, and larger deployments should not require teams to independently derive every sizing decision.

### Proposed improvement

Document **Low, Medium, and High usage profiles**, with:

- Explicit assumptions for request volume, concurrency, token usage, indexed data volume, and ingestion frequency.
- Recommended service SKUs, capacity settings, and scaling options.
- Main bottlenecks and observable signals indicating when to scale up or down.
- Illustrative monthly cost estimates, including region, pricing date, and usage assumptions.
- Links to the corresponding deployment parameters and pricing calculations.

Profiles should be starting points for validation, not guaranteed performance levels. Security and availability requirements should remain explicit across profiles.

Keep component requirements and lifecycle guidance in the separate component reference issue; link to it rather than duplicating that material here.

### Expected outcome

Teams can choose an initial configuration, estimate its cost, and adjust capacity using measured workload behavior.

### Related

Azure/GPT-RAG#692 covers component requirements, optional resources, and lifecycle guidance.

Contributor guide

Open the contributing guide

Research direction

No files or tests are named; begin by locating the deployment-parameter and pricing-calculation references, then review the separate component reference issue, Azure/GPT-RAG#692. Done means documenting Low, Medium, and High usage profiles with assumptions, SKU and scaling guidance, bottlenecks and signals, dated regional cost estimates, links, and explicit security and availability requirements.

Written by the indexing model from the issue text.

Assessment

Tech stack
azure
Domain
cloud, documentation
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
58/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.