When AI Should Speak
15 min read
In April 2024, a company called Humane shipped the purest version of a very old dream: an AI you wear, with no screen, that you simply talk to. You'd ask, it would answer, all by voice, out in the world. The people behind it were ex-Apple, the money behind it was serious, and the idea behind it was the one we'd all absorbed from movies. The assistant that's just there, in your life, all day.
It failed almost immediately. The device was widely panned. Sales never came. Within a year the company sold itself to HP for less than half of what it had raised, and the pins people had already bought stopped working a few months after that. Hundreds of millions of dollars, some of the best-pedigreed founders in the industry, and the screenless talk-to-it assistant turned out to be something people returned.
Here's what makes that interesting instead of just sad. At the very same time, a different version of the same dream was quietly working. Ray-Ban's glasses with Meta's AI inside sold more than seven million pairs in a single year, triple the two years before put together. Demand was high enough that the company held back its overseas launch to keep up at home. Apple, Google, and Samsung are now all pushing into the category. The assistant that lives with you all day didn't die with Humane. It just moved onto your face and started selling.
So one version got rejected and one version is taking off, and the gap between them is the whole point. It would be easy to say the glasses won because the technology got better. That's not what the reviews actually say. Read them and the praise is almost never about how smart the assistant is. It's about how little it gets in the way. The phrase people keep using for the good version is some flavor of "there when you want it, gone when you don't." The glasses that stick are the ones that hand you a translation or a reminder and then disappear. The ones sold on being calm and staying out of your day.
None of that is a capability story. All of it is a behavior story. The thing the market is rewarding isn't a cleverer answer. It's restraint.
That matters more than it sounds, because of where these assistants now live. When AI sits on a screen you choose to open, getting the timing wrong is a notification you flick away and forget. When it sits on your face or in your ear from morning to night, getting the timing wrong is a different kind of problem. It's inside your attention whether you invited it in or not. The same assistant that feels like magic when it says the right thing at the right second feels like a person talking over you the moment it misjudges. And people do not tolerate that for long. They take the glasses off. They mute it down to nothing. They do exactly what Humane's customers did, except now there's a working, capable product underneath, dying anyway because the conduct around it wore out its welcome.
This is the wall the whole industry is walking toward, and putting a better model on it doesn't help. A smarter assistant that still interrupts at the wrong moment just interrupts more confidently. More features give it more ways to speak up when it shouldn't. The all-day, always-present form doesn't fix the timing problem. It removes the escape hatch. There's no app to close. The assistant is simply there, and if it doesn't know when to be quiet, it doesn't get to be useful at all.
So the real question underneath this entire race is not how capable these things can get. It's a smaller, harder one that almost nobody is set up to answer: who decides when the assistant talks?
Not what it says. Models are already good at what to say and getting better every month. The decision that comes first. Should this surface right now, or wait, or stay silent? And if it does speak, what does it lead with, and what can hold? In most companies that decision has no owner. It isn't anyone's job, it shows up in no review, and it ends up being whatever a default was set to once and a notification rule copied from some older product. The behavior isn't designed. It just accumulates. And then everyone acts surprised when people switch the thing off.
The good news is that this is a design problem, and a solvable one. The decision has real inputs you can actually name. How much does this matter. How urgent is it, honestly. What is the person in the middle of, and how much attention do they have to spare. What happens if the assistant is wrong. Feed those in and a system can do something far better than say everything the instant it's confident. It can wait for a better moment. It can offer the one thing that counts and hold the rest. It can stay silent and be able to tell you, afterward, exactly why. There is already a worked-out framework for this decision, with its inputs, its rules, and how they fit together. It's called Behavioral Orchestration.
The assistant in your ear isn't a prediction anymore. It's shipping by the millions and every major platform is in the race. The question that decides who wins it isn't who has the smartest model. Humane proved the smart, always-there assistant gets returned when it doesn't respect the moment. The glasses are proving, almost by accident, that the ones people keep are the ones that know when to disappear. That sense of when to speak and when to stay out of the way won't come from a bigger model. Someone has to sit down and design it, on purpose, as the actual product.
That someone is going to be the difference between the assistants people keep and the ones they quietly turn off.
Want to dig into the framework behind this?
- The whitepaper — the decision logic with no product attached, then shown working inside a moving car.
- Watch it decide in real time — an interactive that shows the system reading a moment and choosing to surface, wait, ask, or take over.
- A worked in-vehicle example — how this played out, from the problem to the system.