This morning, a voice appeared in Codex.

Two hours later, I realised I had stopped using it as a voice feature and started using it as a chief of staff.

I had come downstairs to continue work on the Chief Agentic Officer Briefing. I am redesigning parts of it so the business can work with sovereign and open models as well as American ones. Kimi K3 had just appeared, and I wanted the architecture to be ready for that kind of choice.

That matters to me for capability, but also for operational resilience. If a supplier, model or jurisdiction becomes unavailable, I want a route that keeps the work moving.

Then Codex told me Voice was available.

I love voice. I use it while walking, thinking and working through ideas. But voice inside the agentic environment had been the missing part.

And, honestly, wow.

This was not a chatbot with a microphone

Most voice AI has felt like a chatbot that happens to speak.

You ask a question. It answers. You ask another question. It answers again. The voice may be pleasant, but it is still sitting outside the work.

This felt different because the voice sat over an environment that could do things.

In my setup, Codex already has access to the projects and machines I have deliberately connected to it. I use it heavily because I can coordinate work across my Macs and Mac minis from my laptop, wherever I happen to be. It knows the current task, can work inside the right repository, and can follow the permissions and tools available to that task.

So I started talking.

What is happening across my projects? Is anything stuck? What have I forgotten? Can you check this date? Can you look at the calendar? Can you prepare the note I need to send?

The important shift was not that it spoke back.

It was that the conversation could become work.

OpenAI has joined the voice to the operating environment

OpenAI now describes Voice in Work and Codex as a live interface that can start tasks, check progress and coordinate multiple agents. It uses the tools and permissions available to the selected experience.

That last sentence is the important one.

Voice is no longer only a way to ask the model a question. It is becoming a way to direct an agentic working environment.

The voice itself has changed too. GPT-Live uses what OpenAI calls a full-duplex architecture, which means it can listen and speak at the same time. It can decide whether to respond, continue listening, pause, interrupt or invoke a tool. For deeper work, it can delegate to a frontier model in the background and keep the conversation moving.

That explains the small things I noticed. It could wait while I thought. It could make those little conversational acknowledgements. It did not feel as if I had to package every thought into a perfect command before speaking.

It felt much closer to thinking alongside somebody.

Ultra is not the voice, but it changes what the voice can conduct

I also think this becomes more interesting when you put it beside Ultra.

Ultra is not a separate voice system. It is OpenAI's highest-capability setting for GPT-5.6. OpenAI says it coordinates four agents in parallel by default for difficult work.

In plain English, the voice can become the conductor while the agents do the work.

I can describe the outcome, send one part to research, another to inspect a system, another to check an assumption, and then ask what has come back. I do not have to sit silently waiting for one long task to finish before I can think about the next thing.

This is the role I keep describing as the shepherd rather than the controller.

I am not personally completing every step. I am setting the goal, checking the direction, making the difficult judgement calls and deciding what is allowed to happen next.

That is a profound change in the human interface.

The approval problem becomes more important in voice

I trusted it with a couple of emails this morning because I understood the context and the guardrails. But the experience immediately made me think about how approval should work when the interface is spoken.

In a written interface, I can see a proposed action and press an approval button. In voice, the conversation can feel so natural that it becomes easy to forget that one sentence may carry real authority.

So I would want a few things to remain explicit:

  • a visible transcript of what I asked;
  • a clear distinction between a draft and an external action;
  • a short spoken and visible confirmation before sending, publishing, spending, deleting or disclosing;
  • the recipient, content and consequence repeated back when the action matters;
  • an audit trail or receipt after the action;
  • clear stop-lines for privacy, money, reputation and safety;
  • a fast way to pause or stop the work.

I would also like stronger assurance that it is me speaking. But voice recognition should not become a magical substitute for authentication. The machine, session, account, permissions and approval boundary all still matter.

OpenAI's own material on running Codex safely makes the same architectural point: low-risk work should move inside a bounded environment, while higher-risk actions stop for review. Voice should make that boundary easier to use, not make it disappear.

Voice is the interface.

It is not the authority.

Was this built for a robot?

My mind immediately jumped to robots.

If a natural voice can understand a person, maintain a conversation, delegate work and use tools, it is easy to imagine it becoming the top layer of a physical agent too.

But I need to separate my excitement from the evidence.

There is now a much more direct reason for making that connection. Bloomberg reports that OpenAI is developing a movable, screenless AI companion for the home. The reported device has mechanical elements that can move on their own and is intended to feel like a physical manifestation of ChatGPT.

This is reporting about a product still in development, not an official OpenAI launch. OpenAI has not said that the Voice experience I used this morning will be its interface. That connection remains my speculation, but it no longer feels quite so far-fetched.

It is still a useful thought experiment.

A robot does not only need motors and hands. It needs a way for a person to express intent, interrupt, correct and understand what it is doing. The same voice layer that helps me coordinate digital agents could eventually help a person coordinate physical ones.

The governance questions would then become even more important.

The strange future of working on a train

There is a very practical problem with all of this.

Voice works brilliantly in my quiet workspace. It works far less well on a noisy train, in an airport or in a busy office. Even when modern headphones stop me hearing the surrounding noise, the microphone can still hear voices around me. Everybody around me can also hear what I am saying.

I have written before that agentic work needs quiet. Voice makes that question even more urgent.

There are already some rather unusual attempts to solve it. Shiftall's mutalk 2 is a mouth-covering soundproof microphone. The manufacturer says it can reduce the user's audible voice by more than 20 decibels and is designed to avoid picking up ambient sound. Research prototypes such as WhisperMask are exploring mask-type microphones for whispered speech in noisy environments.

I am not endorsing either as the finished answer. Comfort, hygiene, microphone quality, compatibility and the willingness to wear one in public all need testing.

But I can see the shape of my future travel setup: display glasses, a private microphone, and a voice that can help me coordinate the work without making me open six machines and stare at six screens.

I may look ridiculous.

I may also get a great deal done.

A proper Her moment

The film Her is often remembered because the AI voice feels emotionally natural.

What struck me this morning was something slightly different.

The voice knew enough about the work to be useful.

It could hold the thread while other things happened. It could help me empty the little unfinished thoughts out of my head. It could turn some of them into action and bring the result back. The friction between remembering something and doing something became much smaller.

That was the moment.

Not a voice pretending to be human.

A voice becoming a natural control layer for an agentic working environment.

It is still early. Some of my MCP connections do not yet work as I want. Only one voice conversation can run at a time. Voice use has its own metering, and the approval language needs to mature. The whole thing will need careful governance as it becomes more capable.

But I had a massive smile on my face.

This morning, Codex felt like Her.

And for the first time, that did not feel like science fiction.

Sources and notes

This is a personal observation based on an early Voice rollout in the ChatGPT desktop app. Product availability, limits, model routing and permissions may vary by plan and workspace. Bloomberg and TechCrunch have reported that OpenAI is developing a movable, screenless home AI companion; OpenAI has not formally launched that device or said the Voice experience described here will be its interface. That connection is Tony's speculation. Kimi says K3 is available through its products and API, with full model weights due for release on July 27, 2026.