Jordan-Hall / Jordan-Hall/browser
[P3][VOICE-03] Read-aloud, multilingual and optional wake word
- Dominant language
- No language data
- Stars
- 0
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Description
Programme: #1
Epic: #22
## Objective
Extend the speech layer with opt-in local output, multilingual packs and optional wake-word interaction without weakening privacy or command safety.
## Scope
- Local TTS/read-aloud engine abstraction and voice/language packs.
- Pronunciation dictionary and user technical vocabulary reuse.
- Per-workspace/content sensitivity checks before reading aloud.
- Optional local wake word with explicit on-device state and enable/disable UX.
- Multilingual STT/TTS pack selection and mixed technical vocabulary.
- Audio-device routing, interruption/barge-in and mute behavior.
- Evaluation corpus across target accents, languages, noise and sensitive-content scenarios.
## Privacy/safety rules
- Wake word is opt-in and visibly active.
- Sensitive content is never read aloud merely because TTS is available.
- Wake detection/voice identity is not authorization for consequential actions.
## Acceptance criteria
- [ ] Target language/accent/noise suites meet declared support thresholds.
- [ ] Sensitive content requires context-appropriate user choice before read-aloud.
- [ ] Wake-word mode remains local for supported configuration and exposes clear recording/listening state.
- [ ] User can interrupt output immediately.
- [ ] Language-pack installation exposes size/license/privacy metadata.
- [ ] No retained audio is created without the configured opt-in.
## Dependencies
- VOICE-01
- SEC-03
**First phase:** P3
**Maturity target:** P5
**Owner:** local-ai-speech
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reading VOICE-01, SEC-03, and Epic #22 to establish the speech and privacy boundaries. Break the scope into local TTS/STT packs, wake-word state, sensitivity checks, audio interruption, and evaluation suites. Done means meeting the declared support thresholds while satisfying every privacy, consent, metadata, and interruption criterion.
Written by the indexing model from the issue text.
Assessment
- Domain
- ai, audio-video-rtc, security
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 25/100