KI Praxis Tools

Cloud or local: where AI documentation processes audio and text, and which questions a practice should ask

AI documentation creates three data paths: the audio recording, the transcript, and the draft. "Local" is therefore not a yes/no question but three questions: where is the recording processed, where is the text processed, and who has access? A practice should have a concrete answer for each path before the first consultation is recorded.

Abstract microphone inside a protected frame, next to a dashed path leading to a cloud
Redaktion KiPT Voice · J Medical GmbH

Editorially checked against product behaviour and the stated primary sources; not individual medical or legal advice.

Three data paths, not one answer

The recording contains everything said in the room, including what does not belong in the record. The transcript contains the wording as text, including names, values and asides. The draft contains the ordered version of what was said. Each of these three states can be processed, stored and deleted in a different place.

Vendor statements such as "GDPR-compliant" or "hosted in the EU" say nothing about which of the three paths is meant. The practice should therefore ask the question for each of the three paths separately and get a concrete answer for each.

What cloud processing means

With a cloud tool, the recording or a data stream derived from it is transmitted to the vendor, transcribed there, and afterwards deleted according to the vendor's statements. The vendor thereby becomes a data processor for health data; the practice remains the controller and needs a contract under Art. 28 GDPR, a review of the technical and organisational measures, and an answer to who at the vendor has access.

For medical confidentiality duties, the vendor is a person assisting within the meaning of § 203 StGB (German Criminal Code). This is permissible, but requires that the disclosure stays limited to what is necessary and that the vendor is bound to secrecy.

  • Is the recording transmitted, or only a data stream derived from it?
  • In which region are the servers located, and which sub-processors are involved?
  • When is the recording deleted, and who can verify that?
  • Is content used to train models? The answer belongs in the contract, not in an FAQ.

What local processing means, and what it costs

With local processing, the recording stays on the practice computer; transcription and speaker attribution run there. There is then no audio transmission to assess. The price for this is computing power: an up-to-date computer is needed, the language models sit as files on the device, and updates concern the installation at the practice rather than a server at the vendor.

If the text is also to stay local, a language model for the draft must run at the practice. This is possible but places higher demands on memory and graphics performance, and the quality of the drafts depends on the model used. A practice choosing this path should test it with real conversations, not with demonstrations.

The hybrid case: audio local, text external

Many tools, including KiPT Voice, process the recording locally and send only the transcript text to a language-model vendor, which generates the draft from it. This is a genuine difference from a pure cloud approach, because the recording never leaves the practice. But it is not fully local processing, and must not be described as such: the transcript text is health data, and the same contract questions above apply to its recipient.

Where this text vendor processes the data is a contractual commitment from the vendor. The application itself cannot technically enforce a region; it can only use the configured vendor. Anyone wanting to avoid this configures a local model and accepts the requirements described above.

Questions for the vendor, in this order

The answers belong in writing in the practice's data protection advisory records.

  • Does the audio recording leave the practice computer? If so: to where, for how long, and who deletes it?
  • Does the transcript text leave the practice computer? If so: to whom, in which region, under which contract?
  • Is the draft stored at the vendor, or only at the practice?
  • Which sub-processors are involved, and are they named in the contract?
  • Is content used for training, and is that contractually excluded?
  • What happens to an ongoing recording if the vendor fails?
  • Which person at the practice can verify access, retention and deletion?

How KiPT Voice answers the three questions

The recording, its transcription and the speaker attribution run on the practice computer; the recording is not transmitted and is stored encrypted. The transcript text goes to the provider configured in settings, unless the practice sets a local model. The draft is reviewed at the practice and copied into the practice software. Recording only starts after the documented consent step.

Frequently asked questions

Is "servers in Germany" the same as local processing?

No. A server in Germany is a server outside the practice. The data leaves the practice computer, and the operator is a data processor. Local means: the processing happens on the computer at the practice.

Is a data processing agreement enough?

It is necessary as soon as a vendor processes health data, but not sufficient. The practice must also check whether the processing is necessary, what measures the vendor takes, and whether a data protection impact assessment is needed.

Can a practice run AI documentation entirely without the cloud?

Yes, with a locally running language model for the draft. That requires a capable practice computer and a test with real conversations, because draft quality depends on the model.

What happens to the recording after the consultation?

That is set by the practice's deletion policy. With KiPT Voice, the recording stays encrypted on the practice computer and is not transmitted; retention and deletion follow the rules the practice set before using it.

Read on