Show an estimated AI‑credit cost for the next message on the usage gauge
- Lenguaje dominante
- Sin datos de lenguaje
- Estrellas
- 2.1k
- Forks
- 153
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Descripción
### Feature summary
_No response_
### What problem are you trying to solve?
Cost per turn varies depending on the selected model, the reasoning effort, how much context is loaded, and how the assistant is being run (e.g. a single interactive reply vs. a longer autonomous run). Today users only learn the cost *after* spending it, which makes it hard to make informed choices (switch to a cheaper model, lower reasoning effort, trim context, etc.) before sending.
### Proposed solution
Add a lightweight **"Next message" estimate**: a single, compact prediction of what the upcoming turn will roughly cost, learned only from the user's own recent usage (on‑device, no server calls or extra data collection). Changing the model, reasoning effort, or run mode should visibly move the number.
**UX**
- Visualized in a usage popover, e.g. **"Next message — ~X credits (est.)"**, on a single line alongside "Session" spend.
- Framing: it's an estimate of the *typical* next turn, not a guarantee.
**How the estimate is built**
1. **Learn a typical cost ("anchor") at several granularities.** Maintain a smoothed, geometric (log‑space) moving average of realized per‑turn cost so a few unusually large or small turns don't dominate. Track and blend it at a few levels:
- **Per‑configuration** — keyed by the cost‑relevant choices the user controls: model, reasoning effort, context size tier, and run mode. This is what makes the estimate react when the user switches any of those.
- **Per‑session** — captures the "weight" of the current conversation (a heavy session tends to keep being heavy), ramped in as the session accumulates turns.
- **Global** — a cross‑session fallback used before a given configuration has any history.
2. **Cold start.** Before any history exists, fall back to a context‑proportional approach.
3. **Self‑calibrate.** After each turn, compare what actually happened to what was predicted and fold the realized cost back into the averages, so the estimate improves over time and adapts to the user's habits.
### Workflow impact
_No response_
### Installation context
_No response_
### Additional context
_No response_
Guía de contribución
Línea de trabajo
Comienza localizando el popover de uso y los datos de coste existentes por turno. Rastrea cómo se representan el modelo, el esfuerzo de razonamiento, el contexto y el modo de ejecución; después, determina dónde se puede leer y actualizar el historial de uso de la sesión y el historial de uso global. La tarea estará terminada cuando el popover muestre una estimación claramente delimitada para el siguiente mensaje que responda a esas opciones y se calibre a partir de turnos posteriores sin realizar llamadas al servidor.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Área
- ai, desktop
- Tipo de issue
- Nueva funcionalidad
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Estado de actividad
- Tranquilo
- Claridad
- Bastante claro
- Aptitud para principiantes
- 38/100