ChatGPT Live and the Revival of the Chat Empire
TBAFor years I have closed the day the same way: sitting back down at the computer I had just escaped, opening Obsidian, typing an entry into the daily note. I hate doing it, because by evening the last thing I want is more keyboard. But I also hate losing track of the work I did and the ideas I had that day.
ChatGPT Live (or other similar products in the future) will change this.
In one sense talking to AI is not new in itself. It has been a long time since OpenAI first astonished everyone with its advanced voice mode demo, the one with the Johansson voice, and the market has produced no shortage of AI you can hold a conversation with since. Having the conversation is one thing. Making the conversation do work, even the most ordinary and domestic work, is another. Could I just talk about the day, away from the desk, lying on the bed, and have the conversation end up in the journal on its own?
Once the idea came, the setup was easy. Using GPT-Live itself, I built a skill so ChatGPT would know where my journals live: an Obsidian vault in iCloud, Journal/Daily, one file per date. I showed it yesterday’s note on screen so it could read the format off the real thing instead of a description of it, and it picked up the conventions I actually use: date-led Markdown, first-person prose, H4 headings separating one subject from the next. Along the way it noticed that my Daily Notes settings pointed at a folder I had reorganized months ago and a template I had since renamed. It fixed both, then cleaned the empty placeholder headings my template had been dragging into every new note.
Then I said: put this in my journal.
Boom! The whole process of creating this skill, the hiccups, the solutions, showed up in my journal. Beautifully done.
What has happened?
Maya Has the Voice. GPT-Live Has the Context and Agency.
Two things made that feel like magic, and they are worth separating, because they are different matters entirely.
The first is that I could talk to AI in a genuinely natural way. The second is that the talking got things done.
I have written about the first before. Last July, in AI Companionship on the Rise, I argued that we had been thinking about voice AI wrong by treating it as a faster way to issue commands.
It is wrong to assume that we can reduce our conversational lives into transactions. In truth, it’s all about connection rather than simply checking tasks off a list. Think about it: very few of us surround ourselves with minions that we command; instead we seek companionship from friends and family: you share and respond to sharing. It’s not just about answering questions or executing commands; it’s about creating a bond that feels genuine and meaningful.
That still holds. I still believe there is value in a conversation for its own sake, as opposed to assigning work and reviewing the result.
I ventured that argument on the strength of Sesame, the product that made companionship feel imminent. When Maya arrived in February 2025 she did what no text-to-speech layer had managed: she had presence. Micro-pauses, breath, a laugh in the right place, mock impatience. Sesame’s researchers called it crossing the uncanny valley of voice, and for once the phrase was accurate rather than promotional. People described conversations indistinguishable from real ones. Some described falling in love.
That enthusiasm has a shadow, and OpenAI’s own research has since measured it: heavy, extended voice use correlates with greater loneliness and less real-world social interaction. It is close to the worry I raised in the earlier essay. The companionable voice and the isolating one are the same voice, and GPT-Live is built to hold you in conversation longer than anything before it could.
Eighteen months on, Maya is still the better conversationalist: more emotionally present, quicker to catch a tone, more like a person who is with you. GPT-Live beside her is noticeably flatter: natural enough that you never think about the interface, but comparatively monotone, and rarely conveying much about how it feels to be in the conversation. If what you want is the experience the companionship essay described, Sesame remains the only place to get it. GPT-Live is at best a robotic secretary.
The complaint that has trailed GPT-Live since launch makes the gap concrete. It back-channels constantly: mhmm, got it, right, a stream of affirmative noise laid over whatever you are still saying. Presumably this was built in to sound attentive. It achieves the opposite, because sound intrudes on your attention in a way text never does, and an assistant that keeps jumping in does not read as engaged. It reads as impatient.
I do not remember Maya doing this. She knew when to stay quiet, which is harder than knowing when to speak and is most of what conversational intelligence actually consists of. So the back-channeling is not a cosmetic flaw to be tuned down in the next release. It is evidence that GPT-Live has not yet worked out how a conversation runs, and that Sesame’s lead is real rather than a matter of nicer audio. There is a lot of ground between here and there.
But Maya has a big problem.
She has no context of my work. Every session begins in a vague recollection of what we talked about before, if not a complete void. I cannot share with her what I am writing, nor can she help with any of the work in my hands. I cannot give her my notes, my drafts, or the vault where many years of thinking lives. She is an emotional helpline, so to speak, that soothes your spirit but leaves the work outside the door. There is no transcript either, so nothing that happens survives the conversation that produced it. She can search the web, but that is beside the point. What she is missing is not the world. It is the person talking to her.
That is a strange gap to leave in a companion, because companionship is partly made of context and continuity. What makes a friend a friend is not the quality of any one conversation; it is that they remember the things you did. Warmth plus amnesia does not produce a friend. It produces a very good stranger, over and over.

Which brings me to the second part of the magic, the part I underestimated.
I had always filed personal assistant under luxury. It is what heads of state have, and CEOs, and the sort of executive whose calendar is a full-time job for somebody else. Not because ordinary people would not benefit, but because a human assistant costs a salary and you need enough going on to justify one. So the whole category drifted into the language of status rather than the language of utility, and I never thought about it as something that might apply to me.
A voice session with proper context turns out to be genuinely game-changing. When I asked about the progress of my current book project, it found a note I had written in June and told me something true about my own working habits. It was a revealing moment: 1+1 > 2. What makes voice assistants so valuable is not that they can talk, or that they know your stuff, but a curious combinatory effect of the two.
Companion and assistant are not opposite categories. Sesame currently offers the more natural, more emotionally persuasive conversation. GPT-Live offers the stronger practical assistant: it can be given the context I choose, reach my files when I permit it, and set real work in motion, writing a file, running a search, sending an agent off to do something.
It tells you where things are, what you decided, what is still pending, what you said last time. That is not a rare or elevated need. It is the most ordinary need there is, and almost nobody has ever had it met.
Nothing prevents Sesame from acquiring these agentic abilities, and I would welcome it. The obstacle looks like product direction rather than capability. The demo that had everyone talking in early 2025 took a long time to become an app you could carry in your pocket, and in the interval the company has not moved decisively toward context, memory, or action.
GPT-Live clears a lower bar than Sesame did, natural enough that you keep talking, and it turns out that bar is high enough to be transformative once the conversation can also remember, retrieve, and act. While Sesame struggles to get Maya and Miles into a wearable, GPT-Live has shown everyone what the selling point of those wearables actually is. Imagine walking in the countryside, or riding a bike, or cooking a meal and conducting research at the same time. Or writing poems in the shower. Voice unlocks your productivity in all the situations where typing is structurally impossible.
It could also be live translation, without requiring the other party to be wearing a Meta glass too. It could be rehearsing a high stake speech, conducting interviews.
The use cases are endless.
The Language Tutor
Everyone is currently poking at GPT-Live to find its edges, and most of the demos going around are the obvious ones: interrupt it mid-sentence, make it sing, hold something up to the camera. I wanted an answer to a narrower question. What can it actually hear?
There is an architectural reason to ask. As I understand it, Sesame runs the traditional pipeline: speech in, transcribed to text, text to a language model, the reply spoken back out. Sesame’s real contribution sits at the far end of that chain, generating speech with acoustic texture rather than flat synthesis, which is why Maya sounds the way she does. But the part that thinks never hears you. It reads a transcript. Every hesitation, every vowel, every trace of where you learned the language is gone before the intelligence gets involved.
GPT-Live takes the audio itself. Both systems are full-duplex, as I understand it, so both can interrupt and be interrupted. The difference is upstream of that. One is listening to your words. The other is listening to your voice.
If that is true, it should be audible. So I tested it.
I asked GPT-Live to identify my English accent, and it correctly identified a Chinese Mandarin inflection. A transcript could not have told it that. Encouraged, I asked it to perform various English accents: a Southerner, a New Yorker, a Bostonian, and so on. Although it explained the differences well, it did not demonstrate them with sufficient difference. This points at the limits of using it as a language tutor.
I collect languages the way other people collect kitchen appliances. Chinese, English, French, Japanese, Italian, Spanish, a little German, some of them in various states of disrepair. It makes me a useful test rig, and this is the run I put every voice model through.
My native tongue is Mandarin, so I asked it to place me regionally, refusing to let it guess from one phrase or ask where I grew up. After a long unrelated story it decided I was northern, urban, somewhere around Hebei, Shandong, or Henan.
This is not terribly wrong. I grew up in a large city on the Yangtze, southern in the way that actually matters there: no winter heating, and a childhood spent cold indoors. But years in the north had leveled my Mandarin enough to make the guess defensible. Its own spoken Mandarin, meanwhile, was poor enough to sound distinctly like a foreigner’s.
French is my third language. I studied it intensively for a few years. I still understand it easily but speak it badly, for lack of practice. If GPT-Live could help me brush that up, it would be greatly appreciated. Here things started to work really well.
The sentences it gave me were the right ones: not textbook constructions but the phrases you actually reach for in conversation, short enough to hold in my head and repeat. Within five minutes I was producing whole paragraphs of French. Not perfectly, but continuously, which was the point. What I had been missing was never vocabulary. It was the loop: hear something, respond to it, get the proper way to say it, and try again before the sound has faded. All of this is possible with a human tutor, of course. Not everyone can afford a good one.
I ran the same experiment in Japanese, a language I spent a summer with many years ago and had almost completely lost, and it went the same way. GPT-Live built individual phrases for me, and once I had those, assembled them into a small comic scene, so the pieces sounded like a person saying something rather than a list.
GPT-Live has real potential as a language tutor. Even without a lesson plan it was already impressively useful; the model has in-depth knowledge of the language and was inventing pedagogy on the spot. Hand the same voice to somebody building on the realtime voice API, give it an actual curriculum, and that live language tutor becomes not only affordable but convenient.
The Personal Assistant Everyone Can Afford

I said earlier that GPT-Live can get things done by sending off agents. That is not an ability of the voice model itself. It is an architectural connection.
The reason GPT-Live can do this is that it is connected to Codex (now a part of ChatGPT) rather than sealed off as a voice toy. But the connection is still loose. GPT-Live runs as a layer alongside your sessions rather than inside them. Start a voice conversation from within a project and it arrives knowing nothing about that project; you have to tell it to go find the project and create the task there. It works, and it is plainly a workaround. I expect it to be fixed, because the fix is obvious.
To find out whether the voice could do work, I had a long conversation about a book I have been struggling to move forward, then asked GPT-Live to go find what I had written about the project in my own journal. It searched the daily notes and came back with a timeline. Dates, entries, project ideas, all accurate, all beside the point. I interrupted: it is not the dates, it is the gaps between the dates.
That single correction reorganized everything downstream. Reread with the silences as the evidence, the record showed a fast burst of drafting in June, then four weeks of nothing, then two more. Those were not ordinary quiet stretches. They were the same wall each time. I had not found an honest way to make the manuscript more substantial while I was still testing the thing it argues for rather than actually living it. I asked how this compared with Ideas, Not Words, the book I did finish, and the record showed pauses there too, attached to revision and production rather than to a conceptual blockage.
So we made a decision out loud: not to abandon the book, but to stop forcing it toward a word count, keep running the experiments, and let the evidence decide whether it ends up a prescriptive book or a candid account of an unfinished search. Then I asked for that retrospective to go into the day’s journal, and it did.
Every step was possible by typing, or even by dictation. But dictation is still not a conversation. When you dictate a prompt you still have to read what comes back. There is nothing wrong with that, except that not having to read feels entirely different. I can close my eyes while doing this, or keep looking at some other document. Either way I choose what my eyes are doing, which keeps me on the ideas. Reading interrupts that flow. It distracts from the thinking.
That is the second difference from voice dictation, which I already use constantly. The GPT-Live voice stays engaged throughout, because it runs in the background. When you say nothing it waits patiently, and sometimes you forget it is there at all. Dictation makes you mark the beginning and the end of every utterance. Dictation is turn-based; GPT-Live is real-time strategy. The work becomes a continuous flow.
When the work is search, retrieval, or writing a file, it dispatches an agent to GPT 5.6 Sol, the frontier model behind it, and keeps talking to me while that runs. You can set several things going at once and it reports back as each finishes. So it is not only that you have a personal assistant. This assistant has an army of valets waiting outside the door to run the errands.
What Goes Under the Hood

My language sessions were pure conversation; nothing outside the exchange was touched. But when a spoken sentence found my vault, read my past entries, and wrote a paragraph back into a file, something else was going on, and it is worth knowing what, because these capabilities are not evenly distributed across everything called ChatGPT.
Here is the practical map.
ChatGPT now has three components merged into one product: Chat, Work, and Codex. When the merge first landed, OpenAI seemed to have terrible ideas on how to arrange the three. Work and Codex sat in a dropdown menu, and chat had been reduced to a small window inside ChatGPT Work. I wrote about that mess three weeks ago. Since then OpenAI has moved quickly to remedy it. The dropdown now offers ChatGPT and Codex. So where is Work?
It turns out to be almost a hidden feature.
A live voice conversation starts in chat, consistently across desktop, web, and mobile. I see no way to start one as a Work or Codex session, or to put it under a project. As implemented today it is a special kind of chat, attached to nothing.
The interesting part is that you can start a conversation in the browser or on mobile, resume it in the desktop app, and convert Chat into Work there. That conversion happens the moment you ask for access to the local file system.
The word “desktop” in that sentence is carrying more weight than it looks. A voice session on the phone gets none of this: no connectors, no reach into the vault on my disk or in iCloud, no way to dispatch an agent or open a task in Work. On mobile it is a conversation and nothing more.
This is a different implementation from Anthropic’s. In Claude, Chat and Cowork sessions are separate and stay that way. An entry on the sidebar is either a chat session, showing the chat icon, or a task, showing the to-do list icon.
For ChatGPT desktop, the sidebar still tells you which is which: all the sessions on the sidebar are labeled as either chat or work. The categories did not disappear; the moment you have to care about them moved, from the start of the conversation to the point where the conversation reaches for the world.
Decline or Revival?

In a recent essay, The Decline of the Chat Empire, I argued that both leading labs had restructured their consumer surfaces around the same conviction, that conversation was being demoted from destination to doorway, and that the race had moved to trust with delegation. I ended on a line: stop asking AI how good its answers are, and start asking what it can finish for you.
Much of that holds. But the word “chat” was doing two jobs in that essay, and I only killed one of them. What I had in mind was the instant-messaging kind: you type, you wait, you read, you type again. That is a dead end, and I stand by the obituary.
Meanwhile people were trying to rescue chat from inside the text box. The OpenClaw movement early this year was the clearest case: take the messaging channels people already live in and wire real agents into them. That it caught on at all shows how much appetite was still there for the text exchanges.
But chat also means a voice conversation, and that meaning is far older and far larger. Think about how much of what people actually do runs through talking. You teach in a voice session. You hold a meeting in a voice session. You hang out with your friends, and that is a voice session too.
Look closely at what a meeting is. Nobody builds anything in a meeting. You find out where things stand, you get your understanding lined up with everybody else’s, and now and then an idea surfaces and the room works out on the spot whether it is feasible. None of that is the work. All of it is part of the work, and it is the part that decides what the other part will be. Then everyone leaves the room and goes and does the thing.
That meeting is the gap GPT-Live fills, and it is why the architecture matters more than the voice. The session is the talking half, and it is only worth having because the voice arrives already knowing where my things are, the way a colleague walks in already knowing the state of the project. Go read that guideline. Go revise that report. The moment I say it, the doing leaves the conversation. It happens somewhere else, by an agent, in files I am not looking at, while the talking continues without waiting.
We are not quite there yet. The context and the agents live on the computer, and the phone gets the conversation by itself, which means the walk and the bike and the shower remain places where you can think out loud but not set anything in motion. In my tests the current remote pairing does not work for voice chat. Think of this as a feature request, if someone is reading this. If this turns out to be an architectural limit rather than a shipping order, then you are still tied to your desk whenever you need the context.
Remember Her, the film that still sets the terms thirteen years on? Samantha was never just a chatbot. She was the Operating System. She had a voice Theodore can talk to at any time, walking, lying on bed, and she also read his mail, sorted his correspondence, and quietly got his book published. Companion and assistant were never separate things in her. She knew his life and could act inside it, sometimes even in physical forms with the use of a surrogate body. Thanks to OpenAI, we are visibly one step closer to her.
Voice is the universal code of human interaction. And now AI is cracking this code.

Comments