Video

OpenAI Called It Voice Mode. It's Actually an Orchestrator.

I saw it on X, thought "cool, another transcriber," and almost never opened it. That was the wrong category — this thing starts, checks, and steers work in threads I'm not even looking at, and that changed my day more than any model release did this week. I also get into Opus 5, which in my own work has been fast enough and capable enough to change what I hand to which model.

What this video covers

  • I misread ChatGPT Voice as another transcriber, and I think the name is going to cause most people to make the same mistake.
  • Dictation puts my words into the window I am already looking at, while dispatch starts, checks, and steers work in threads I am not looking at.
  • Being stuck in one window at a time was the real constraint on my multi-project work, and it was never the models that were slow.
  • In my own work Opus 5 has performed as well as Fable and sometimes better, it is noticeably faster, and on my subscription it leaves far more of the allowance intact.
  • This does not work in any room where you cannot talk out loud, it needs a paid plan, and I expect adoption to lag well behind the capability.

Watch on YouTube →

Read the companion article →

Full episode transcript

OpenAI shipped an orchestration layer last week and called it voice mode. I almost skipped it. I saw it on X, thought "cool, another transcriber," and opened it out of habit. Once I opened it, I realized I had filed it in the wrong category. And I want to be honest about that, because the way I misread it is the way I think most people are about to misread it. Voice mode. Okay. So it listens better, maybe the transcription's cleaner. That was genuinely the whole thought. Open it, poke at it for a minute, get on with my day. The name undersells it. Badly. If you call something voice mode, everybody files it next to the microphone button they already ignore. Nobody hears control layer. And the thing it actually does — start work somewhere I'm not even looking — doesn't live anywhere in those two words. That gap is the only reason I nearly walked past it. Here's why it hit me the way it did. Everything I do is a dynamic workflow. I'm never sitting in one window. I'm across Claude, ChatGPT, Codex, Cowork, Slack, a pile of different projects, all day, for my work and for client work. And moving through all of that means typing. Typing is slow, and I know how that sounds, but it's a real throughput problem. The thought is finished, and then I go hunting for the right window and getting it into the right place, and by the time it lands I've trimmed it down just to save the time. So I'd already gone looking for a fix before any of this. I've been using Wispr Flow for a while and it's good. It turns what I say into text right in whatever app or field I'm working in. That helps more than people expect, because when you talk, the whole thought gets out — the caveat, the second half, the part you'd have deleted for brevity if you were typing it. Talking is just a wider pipe than my hands are. But that's still dictation. It's putting my words where I'm already pointed. I still have to be the one standing there pointed at it. That's the ceiling — every window is still waiting on me to show up in it. ChatGPT Voice isn't doing that. It isn't typing for me. It's dispatching for me. What OpenAI documented is that you can start a task in voice, and then ask it to start, check, or steer work in other threads. Other threads — the ones I'm not sitting in. I can talk about something I'm not currently looking at and work gets moving over there. The way I think about it now — if I'm the main node, this is the second in command. I'm still the one deciding what the instruction is, and it carries that instruction out to the places where the work actually lives. And I don't have to be linear about it anymore. Because I was linear. And I was linear because the interface was — one window, one thing, finish it, close it, go to the next one. Every project I wasn't personally looking at was effectively paused, sitting there waiting on me to get around to it. So now it looks like this. I say: build the packet for this company, and while that's running, check that my systems are still up. Two different pieces of work, one spoken instruction, and neither one is waiting on me to open a window and start it. I'm still the one doing the checking, which I want to be — I just show up to work that's already moving. Now the honest part, because I don't want to oversell this. It's useless anywhere you can't talk out loud. If you're in an office with people around you, this does not work. If it's late and you're sitting next to your wife narrating tasks into a laptop, that gets old fast, and I would know. Then you're back on the keyboard like everybody else. And one thing you should hear from me directly: I use these tools every day in my own work, and ChatGPT Voice needs a paid plan — it is not on the free tier. Which brings me to Opus 5, because Anthropic put that out the same week and I've been working in it since. Here's what it's actually been like for me. For the work I do, Opus 5 has been performing as well as Fable, and in places better. That's my read from my own use. I didn't run a benchmark, and I can't tell you what it'll do on your work. It's also fast, and speed is the most underrated thing about a model. When something comes back quick, you hand it the next thing. When you're waiting, you start doing the work by hand. Then there's what it costs me. Fable is expensive for me to run. I'm on a subscription like most people are, and on my plan Fable chews through the allowance a lot faster than Opus does. Same day's work, and one of them has me rationing by the afternoon. The framing Anthropic published with it is that Opus 5 comes close to Fable 5's frontier intelligence at half the price. To me that announcement reads like a price story. I think it's a routing story. Once dispatch is this cheap — once I can push work somewhere just by talking — what I'm deciding all day is what goes where. Smartest available stopped being the whole question. Near-frontier capability at half the cost is a different answer than I had a week ago. And picking a model was never only about intelligence anyway. For me it comes down to how fast it works, how accurate it is, whether it hallucinates on me, whether it does the thing I actually asked instead of the thing next to it, and how it holds up on unattended work when I'm not sitting there watching. Cost sits on top of all of that. Something a hair smarter that burns twice the budget getting there is a bad trade on an ordinary Tuesday. That split is already how I work. Codex is the rapid-fire one, the hard worker, my daily driver for most of the day-to-day. Claude I go to for the bigger, more complex thinking, new projects, the stuff where I want more room. That separation existed long before last week. What changed is what it costs me to act on it — it used to cost a window switch and a paragraph of typing, and now it costs a sentence. And my bet, and I'll call it a bet, is that having real options at this level keeps pushing price down and speed up. That's good for everybody trying to get actual work done with this. I'll say the bigger thing carefully. I think AGI is a spectrum. Nobody walks through a door and it's suddenly here. Fifty years from now people are going to point at completely different moments and argue about which one it was, and I doubt any of them will be exactly right. But this is one of the moments that shifts how it feels. Handing something off used to feel like operating a tool. Now it feels closer to handing it to a partner. I didn't measure that. I'd rather tell you plainly what my days feel like than dress it up as something I tested. I don't know how fast this gets adopted. Honestly, probably slow. The payoff shows up when you're already running a lot of agents across a lot of projects, and most people aren't there yet — they'll poke at it, decide it's kind of neat, and go back to typing. That's fine. But if you work with AI at all, go try it. I think it was genius. It's one of the most impressive tools I've ever been handed, and I'm saying that as the guy who opened it expecting a transcriber.