Feature request: Live Voice as a first-class Codex orchestration surface
Nobody has claimed this yet.
- Dominant language
- Rust
- Stars
- 125k
- Forks
- 19.5k
- PR merge metrics
- PR metrics pending
Description
The request
I want live voice in Codex to become a serious operating surface for real work—not a separate, lightweight way to talk to Codex.
Right now, voice is useful for conversation. The bigger opportunity is to let it run the same kind of focused work a good Codex task can run: keep context, follow a working style, use the skills I have approved, and coordinate work through my own setup.
This is not a request to expose every model knob to everyone. OpenAI can keep choosing the sensible default. What I am asking for is deliberate control when it matters.
Make orchestration the centre of live voice
The hero feature should be native orchestration.
When I start a live-voice session, I should be able to say: “Work as my operations lead for this task,” select a saved style, give it the right task context, and let it use the skills and workflows I have approved. It should be able to move between voice and text without becoming a different, forgetful assistant.
For me, that means voice can become the front door to Codex: I speak the goal, Codex keeps the thread, coordinates the work, asks only for meaningful decisions, and brings the result back. That would be a huge jump in usefulness.
What I would like to see
-
A real voice-session setup screen
- Start fresh or deliberately resume a specific task.
- Pick a saved working style / instruction profile for that session.
- See what context, skills, tools, and connections the session will use.
- Avoid accidentally reopening an old, oversized task just because it is the most recent one.
-
First-class orchestration
- Voice sessions should use approved Codex skills, workflows, and task capabilities—not stop at chat.
- The session should preserve task continuity across voice and text.
- It should be able to coordinate bounded work and report back clearly, with the same permission model users expect from Codex.
-
Secure connection to the user's own harness
- Let users connect a voice session to their personal context and local execution setup—for example, a knowledge system such as My Brain, or a local harness.
- Make every connection explicit and permissioned. I want visible control, not hidden access.
-
Model and voice controls where supported
- Keep “Automatic” as the default.
- Where OpenAI allows it, offer an advanced model choice in the voice-session picker.
- Let users select the voice and set a practical voice working style without having to repeat it every session.
Why this matters
Live voice is the most natural way to start work when I am away from the desk or on my phone. If it can launch a properly configured Codex task—with my context, my working style, and my approved capabilities—it stops being a side feature and becomes the fastest way into serious work.
I would be happy to test this and provide direct product feedback.
— Hammad
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
The issue does not name files, tests, or entry points. Start by locating any existing live-voice, session, and orchestration surfaces, then narrow the broad request to one explicitly scoped behavior with a clear permission and continuity path. Done would require an agreed implementation scope and tests for that selected behavior.
Written by the indexing model from the issue text.
Assessment
- Domain
- ai, developer-experience
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100