Ambient scribe: what the term means and how it works during a consultation
An ambient scribe is a program that records a conversation between a doctor and a patient in the background and turns it into a structured documentation draft, without anyone dictating or typing anything during the conversation. The name makes two claims about the workflow: "scribe" claims that someone is writing down what is said, the way a person taking notes during an appointment would; "ambient" claims that the recording runs alongside the conversation rather than something being dictated into it.

Editorially checked against product behaviour and the stated primary sources; not individual medical or legal advice.
What an ambient scribe is
An ambient scribe is a program that records a conversation between a doctor and a patient and turns it into structured documentation, without anyone announcing or typing anything during the conversation. The name makes two claims about how it works: "scribe" claims that someone is writing down what is said, the way a person taking notes during an appointment would; "ambient" claims that the recording runs alongside the conversation rather than something being dictated into it.
The difference from classic dictation software sits in that one word: dictation software transcribes what a person actively speaks into it. An ambient scribe records the natural conversation and only orders it afterwards. For the practice, that means nobody interrupts the conversation to announce a finding, but the recording only works meaningfully if it has already been established that it is running, and for how long.
What happens during the conversation
Five steps take place, even when a vendor presents them as a single experience: documented consent, the recording, transcription of what was said, attribution of the contributions to the two speakers, and the creation of a structured draft. Each of these steps is its own technical process with its own source of error.
This separation is more than technical bookkeeping: a tool can handle transcription reliably and still fail at speaker attribution, for example when both voices sound similar or more than two people speak. Anyone who treats the individual steps as one closed unit cannot spot a weak point before it shows up in the finished draft.
- Consent: the recording may only begin once the patient has agreed.
- Recording: the microphone records the conversation without interrupting it.
- Transcription: the recording is converted into text.
- Speaker attribution: the text is assigned to the two speakers.
- Structured draft: the attributed text becomes a structured documentation template.
What an ambient scribe does not do
An ambient scribe does not evaluate the conversation. It does not weigh any statement and does not decide what matters — that stays with the person who conducted the conversation. Nor does the draft replace the doctor's own wording: it is a starting point that is read, corrected, and only then adopted.
And it knows only what was said in the room. Information from earlier records, values that were not read aloud, or observations nobody put into words do not appear in the draft, because they were never part of the recording.
Where vendors differ
Within the five steps there is room for variation, and that is exactly where vendors differ from one another. The points below are questions to ask a product, not a rating of any particular vendor — a dated comparison with full vendor names is on the linked comparison page.
- Where the recording is processed: on the practice computer or on a vendor's server.
- Where the transcript text is processed and who has access to it.
- Whether the contributions are attributed to the two speakers or the text stays unstructured.
- Whether the draft's structure can be adapted to the practice's own documentation.
- How the finished draft reaches the practice management system.
- What happens to the recording after the consultation and who can verify that.
What a German practice settles before using one
Consent comes first, not as an afterthought: recording without the patient's prior agreement is a criminal offence under § 201 StGB (German Criminal Code), regardless of how well the tool works afterwards. The practice decides in advance how consent is obtained and documented, and the tool must only allow the recording to start technically after that.
As soon as transcript text leaves the practice, the recipient is a data processor for health data and needs a contract under Art. 28 GDPR. That applies regardless of whether the vendor keeps the text only briefly or stores it for longer.
Within the practice, one person should be named who can account for how the recording, text and draft are retained and deleted — not as an extra formality, but because in case of doubt this is exactly the person who will be asked.
And regardless of which tool is used, the content of the record remains the practice's responsibility. An automatically generated draft changes nothing about that; it must be read before it is adopted.
How KiPT Voice implements this workflow
KiPT Voice runs the recording, transcription and speaker attribution on the practice computer. The recording does not leave the computer. The transcript text goes to the language model configured in the settings, unless the practice sets a local model; in that case the text also stays on the computer.
The recording only begins once the documented consent step is complete. The resulting draft is checked and edited on screen before it is copied into the practice management system; the template's structure can be adapted to the practice's own documentation in the settings.
The review screen also flags when a value in the note does not appear in the transcript in that form, and asks for it to be checked. The flag says the figure could not be found in what was said — it does not say the figure is wrong, and it cannot tell whether a figure that was mentioned ended up in the right place.
Frequently asked questions
What does the term "ambient scribe" mean?
There is no single fixed definition. In essence, the term describes a recording that runs alongside the conversation, and a person taking notes without anything being dictated to them.
Is an ambient scribe the same as dictation software?
No. Dictation software transcribes what someone actively speaks into it; an ambient scribe records the natural conversation and only orders it afterwards. The difference and its consequences for the practice are covered in the linked comparison of both approaches.
Does the program listen continuously?
No. The recording is deliberately started and stopped, and only after documented consent. There is no continuous background operation.
What happens to the recording after the consultation?
That is set by the practice's deletion policy. With KiPT Voice, the recording stays on the practice computer and is not transmitted; retention and deletion follow the rules the practice set before using it.