Voice dictation needs somewhere to land

Danielle Morrill
Groupthink remembers the people in your life.Start free →

Wispr Flow published its master plan this month. It is worth reading. The short version is three phases: make voice input reliable enough that you never edit what it wrote, then turn voice into action and a second brain, then put it on wearables so you can use it anywhere.

We agree with almost all of it, including the parts that are uncomfortable for us. Eight billion people type every day and most of them are slower at it than they are at talking. Their target of a zero edit rate is the right bar, because a dictation tool that is wrong one time in twenty is a tool you stop trusting. And they are right that most people do not speak in clean commands. They speak in half-formed thoughts, which is the thing assistants have always handled worst.

We use their gesture. Hold a key, talk, let go. It is the correct design and we did not improve on it.

The part we would add

Their second phase is where we see it differently. A second brain with no subject is a notes app, and we have all tried a notes app.

The words have to go somewhere that gives them meaning after you say them. For a developer that place already exists. You dictate into the editor, or into the prompt you are writing for your coding agent, and the destination is obvious because the work and the words live in the same file. That is also who these tools reach first. Wispr Flow’s own settings have a tab called Vibe coding.

Everyone else is dictating about people. Where the deal stands. What you promised someone and when. Who you should introduce to whom. Why the call went sideways. That is what non-developers say out loud all day, and for that content the destination does not exist yet, so it goes into a Slack message and disappears.

This is the gap between their plan and their product today, and we do not think it closes by making transcription better. You can hit a zero edit rate and still lose everything you said, because the problem was never accuracy. It was that a perfect transcript in a text box is still a text box.

Where we put it

Groupthink is a personal CRM. It keeps a record of the people you deal with regularly, so you walk into a conversation knowing where you left off. Not only colleagues. Clients, candidates, the people you met once at a conference and should not have lost.

So when you hold the key and say “Adam is pushing the renewal to Q1 and wants pricing before the board meeting,” two things happen. The sentence is typed wherever your cursor was, which is the Wispr behavior and the reason you would use it at all. Then the same sentence is saved to your Captures and filed against Adam, so it is waiting the next time his name is on your calendar.

When the match is not obvious, you get told. The capture is kept and marked for you to confirm rather than filed somewhere confident and wrong. We would rather ask than guess, because a record you cannot trust is worse than no record.

Nothing about that requires you to file anything, or to open Groupthink, or to have decided in advance that the thought was worth keeping. That last part is the whole point. The thoughts worth keeping never announce themselves at the time.

Where we agree again

They argue that the AI landscape stays fragmented, that people will use different models for different jobs, and that something has to sit across all of them. We think that is correct, and it is the same reason we are building what we are building from the other side.

They want to be the input layer across whatever AI you use. We want the record to be the thing that survives whichever AI you use, and whichever employer, and whichever tools your next job makes you adopt. Those are not competing ideas. An input layer with nowhere durable to write is a faster keyboard, and a durable record with nothing feeding it is an empty database.

We are not trying to beat them at transcription. They are better at it and they have said out loud that they intend to stay better at it. We are betting that the more interesting question is the one after the words are correct, which is whether anyone will ever see them again.

The consumer part

There is a version of this that only ever serves people who already live in a terminal, and that version is a smaller business than either of us wants.

The reason voice matters for everyone else is that the alternative is not typing. It is not capturing anything at all. Nobody opens a CRM after a hallway conversation. Nobody writes a contact note in a parking lot. They tell themselves they will remember, and then they do not, and six months later the relationship has gone quiet and nobody can say exactly when.

Holding a key for four seconds is short enough to survive a real day. That is the bet. Voice is not a faster way to do the thing you were already doing. It is the only input cheap enough to catch what you were otherwise going to lose.

Voice dictation is live now in the Groupthink desktop app for macOS.