The bot got normalized, not accepted

A meeting bot is an eavesdropper we taught to say hello. Somewhere in the last few years a third participant appeared in everyone's calls: 'So-and-so's Notetaker has joined the meeting.' We all learned to ignore it. The idea of taking meeting notes without a bot came to sound exotic, when it should be the plain default.

Normalization is not acceptance. It is the quiet agreement to stop objecting to something you never actually chose. Clients didn't stop noticing the bot; they stopped mentioning it. Familiar and welcome are different states, and only one of them is worth designing for.

None of this arrived by decree. No team ever sat a customer down and argued that a robot belonged in their private conversations. It arrived the way most defaults do, one quiet increment at a time: a setting on by design, a calendar plugin that invited itself, a colleague whose notetaker showed up first so that yours felt ordinary by comparison. Each step was small enough to wave through, and none of them ever had to make the case out loud. Normalization is simply what a hundred unexamined defaults look like once they have finished settling into place, and the absence of an argument is not the same thing as a good one.

The bot exists for a mundane engineering reason: joining the call is the easiest place to intercept audio. A program that dials into the meeting as a guest can hear whatever the meeting sends it, without touching your machine or the meeting vendor's internals. It was never a product decision anyone would make on its own merits; it is a plumbing decision that leaked into the room. And the people in that room absorb its cost so the software can take the easy path, which is worth looking at directly, because it lands hardest on the person you least want to unsettle.

The cost lands on the client

Every bot that joins a call makes a small announcement: this conversation is being processed by a system you did not choose, run by a company you did not vet. On an internal standup, nobody cares. On a first sales call, a legal intake, a coaching session, or a candidate interview, that announcement is doing real work, and none of it is in your favor.

Ask a sales team how many deals started with two minutes of 'is it okay if my bot records this?' and you'll hear the real cost of normalized. Those two minutes are not neutral. They move the client out of your pitch and into their own risk model, at the exact moment you wanted their attention on the problem you solve.

There is also the question the bot can't answer from the lobby: where does the audio go, who can request it later, how long does it live. A guest that phones home to a vendor's cloud has already answered that with 'somewhere else, for a while.' A tool that never sends the audio anywhere can answer differently, and the difference stops being aesthetic the moment you have to put it in writing. If you are weighing that against a specific incumbent, the comparison with a cloud notetaker like Fireflies lays it out side by side.

Look at who benefits and who pays, and the asymmetry is hard to miss. The vendor gets a copy of the conversation and a sticky account; the person running the tool gets a little convenience; the client, who had no vote, carries the exposure, their words on someone else's servers under a policy they never read and retained for a span they were never told. None of this asks the client to be paranoid. It only asks them to be a professional with obligations of their own, and the polite thing is to give them less to carry, not a smoother way to agree to more.

Capture at the source instead

Your Mac already has everything live capture needs. The microphone is your side. The system's audio output is their side. Savory captures both directly, through the same sanctioned macOS frameworks screen recorders use, and transcribes the result on the device itself.

This is not a clever hack around the meeting. It is a more honest description of what a meeting already is: sound coming out of your speakers and sound going into your microphone. Capture at that layer and you are recording the thing you were already hearing, with nothing standing between you and it.

There is a durability argument here as well. A bot lives entirely at the mercy of the thing it joins. When a meeting platform tightens a guest policy, changes a permission, or ships a new client build, the bot can quietly stop working, and you find out from a blank transcript after the call that mattered. Capture at the source carries no such dependency. The operating system's audio path is stable ground that does not renegotiate its terms every quarter, which is the unglamorous reason this approach keeps working while other tools maintain ever-lengthening lists of what they no longer support.

Two sides is a macOS capability, and it is worth being exact about that. On macOS, Savory takes your microphone and the system audio output and keeps them as distinct sides of one conversation. On iPhone and iPad the sanctioned path is your microphone alone, so iOS captures your side of the call. We would rather say that plainly than blur it into a promise the platform will not keep.

Meeting notes without a bot, on any call

Because nothing joins the meeting, nothing depends on the meeting app. Zoom, Meet, Teams, FaceTime, a phone call on speaker, a conversation across a desk: if your Mac can hear it, Savory can transcribe it. There's no integration list because there's no integration.

This is the part that sounds too simple to be a strategy. Every bot-based tool ships a support page listing which platforms its bot can join, which meeting types it can't, and which enterprise settings quietly block it. Meeting notes without a bot delete that page. There is nothing to stay compatible with, because there is nothing to connect.

Inside that shared audio, speaker separation tells voices apart, and one click turns 'Speaker 2' into a named person across the whole transcript. You can see how that works on the live recording page. The result is structured, attributed notes from a call that no software ever entered.

The unglamorous cases are where this quietly pays off. A quick call taken on speaker while you walk to your desk. A client who insists on a platform your bot-based tool never supported. A conversation at a table with no meeting software running at all. None of these has a lobby for a bot to wait in, and every one of them is simply sound your Mac can hear. The absence is the feature: a tool that joins nothing is the only one that reliably shows up for all of them, which is a strange advantage to win by subtraction.

The echo problem, solved where it lives

One honest complication comes with capturing both sides over speakers: your microphone also hears the other person, and a naive recorder would transcribe that echo as you. Savory runs echo cancellation between capture and transcription, so the far side's words stay on the far side, and your transcript doesn't credit you with things the other person said.

That is as far as we will take the mechanism here, because it deserves its own room. Why removing the echo is harder than muting a channel, and how the two sides come out clean, is the subject of a separate piece on capturing both sides of a call without echo. The short version is that the complication is real, that it is handled before anything is transcribed, and that it is handled without sending either side off your machine to do it.

A bot outsources the awkward question to a lobby notification. Savory keeps it where it belongs, with you. The polite thing, telling the people on the call that you're taking notes, stays yours to say, because it should be a sentence from a person, not a banner from a program.

Where the software helps is in remembering what you decided. When you record from a calendar event, an explicit confirmation stands between you and the microphone, and per-person rules can mark people who should never be captured. That enforcement lives on the calendar path; ad-hoc capture is more manual by nature. The full shape of it, and how a person's preference travels with them, is its own post on per-person recording consent.

There is a social reason to keep this human, not only a legal one. When a person says, I am going to take notes, is that alright, they are taking responsibility in front of the other party. A banner takes responsibility for nothing; it simply appears and waits to be ignored. The sentence costs you a small moment of awkwardness, and that moment is precisely the point. It is the part of consent that a lobby notification was quietly engineered to route around, and routing around it is exactly what makes the notification feel off.

Software shouldn't launder that conversation. A bot that announces itself is not asking permission; it is informing you of a decision already made by whoever added it to the invite. Moving the decision back to a human is not a limitation. It is the correct owner holding it.

The polite bot is still a bot

The industry's answer to the discomfort has been to make the bot polite. Give it a friendly name, a soft chime, a line in the chat, maybe a small avatar. The reasoning is that a well-mannered intrusion stops being an intrusion. It does not.

A polite bot is still a third party on the line. The chime does not change where the audio goes or who ends up holding it; it changes how the room feels for the first ten seconds and nothing after. Politeness is a coat of paint on the same architecture: something joined the call that did not have to.

There is a name for this pattern outside our field: disclosure theater. You perform the ritual of telling, you collect the reassurance of having told, and you change nothing about the exposure underneath. Ask who the courtesy is really for and the answer is the person running the bot, so they feel they disclosed; it rarely reduces what the client actually has to trust. A named, chiming bot confidently answers the question, did we say something, and steps around the one that matters, which is whether this thing should be on the call at all. The client is only ever asking the second one, and a friendlier answer to the first does not reach it.

So 'no bot' is not a smaller, quieter version of 'polite bot.' It is a different design with a different bill of materials. Nothing to announce, because nothing is there. Nothing to trust with the audio, because the audio never leaves your machine.

The quiet part

Here's what the no-bot architecture makes possible. Since capture happens on your Mac, the audio never has to leave it. Savory goes one step further: the audio is never even written to disk. It is held in memory, transcribed locally, and gone. No cloud copy exists because no cloud was ever in the call.

Notice that none of this was a privacy feature bolted on at the end. It falls out of the first decision. Once you capture at the source and refuse to send a guest into the meeting, there is simply no moment where the audio would need to travel somewhere else. The strongest version of a promise is the one the design cannot break even if it wanted to.

Consider what stops existing alongside that cloud copy. There is no retention window to configure, no breach that can someday spill a call you had months ago, no request that reaches audio a server never kept, no data-processing addendum to negotiate over notes that never left your desk. Each of those is a thing other tools must manage, document, and ask you to trust, and each is a promise with a way to fail. The version here is duller and stronger at once: the risky object was simply never created, so there is nothing to secure, expire, or explain.

The best meeting bot is no bot. Savory is in early access for macOS and iOS, and your next call can be the first one with nothing extra in it.