# What cloud voice assistants keep from your voice, and what on-device speech recognition changes > Alexa, Gemini and Siri differ in where speech is processed and how long recordings and transcripts are kept. A source-based review, and what on-device AI changes. FireAI Security & Research Team (HisnLabs) · Published 2026-10-08 Canonical: https://hisnlabs.com/en/blog/cloud-voice-assistants-recordings-on-device-speech A voice assistant turns sound into text, text into an answer, and the answer back into sound. Each of those steps can run on a remote server or on the device itself, and the choice decides who holds a recording of the user’s voice, a transcript of what was said, and for how long. This note compares, from the vendors’ own privacy pages, a regulator’s complaint and a peer-reviewed measurement study, what three widely used assistants send off the device and retain. It then sets out what on-device speech recognition and on-device language models change, and what they leave unchanged, with AISir, a HisnLabs planner and voice assistant for the Mac released on 8 October 2026, as the worked example. ## Background A spoken request passes through four stages: voice activity or wake-word detection, speech recognition (speech to text), a language model or rules that decide on an answer, and speech synthesis (text to speech). In most consumer assistants, only the first stage runs locally: the device listens for its wake word and then streams audio to the vendor’s servers, where the other stages run. Two artefacts can then be retained by the vendor: the audio recording and the transcript. Retention matters for three reasons. A stored recording is biometric-adjacent data that identifies a speaker; a stored transcript can reveal names, addresses, health details and work matters; and either can be read by people, either human reviewers employed to improve the service or anyone who later gains access to the account or to the vendor’s systems. Open speech-recognition models made a different design practical. In December 2022, OpenAI researchers described Whisper, trained on “680,000 hours of multilingual and multitask supervision”, and stated: “We are releasing models and inference code to serve as a foundation for further work on robust speech processing” [[8]](https://arxiv.org/abs/2212.04356). In July 2025, the authors of WhisperKit presented “an optimized on-device inference system for real-time ASR that significantly outperforms leading cloud-based systems”, a claim they make for their own system and their own test set [[9]](https://arxiv.org/abs/2507.10860). ## Findings ### Amazon Alexa: cloud processing, recordings kept unless the user opts out Amazon’s privacy page offers a setting not to save recordings: “If you choose not to save your voice recordings, they will be automatically deleted after Alexa processes your request.” Transcripts follow a separate rule: “You will still be able to review the transcripts of your Alexa requests for 30 days before we begin automatically deleting them.” Users who keep recordings can choose to “automatically delete your voice recordings on an ongoing three- or 18-month basis”, and Amazon explains why saving is offered: “By choosing to save your voice recordings, you have access to more personalized features, Alexa can better understand requests, and we can continue to improve the service” [[1]](https://www.aboutamazon.com/news/devices/alexa-makes-privacy-even-easier). Local processing was briefly available on a few devices and then withdrawn. TechCrunch reported on 15 March 2025 that Amazon had emailed owners of the fourth-generation Echo Dot, the Echo Show 10 and the Echo Show 15 who had enabled “Do Not Send Voice Recordings”, saying it would stop supporting the feature on 28 March. The email, as quoted, read: “As we continue to expand Alexa’s capabilities with generative AI features that rely on the processing power of Amazon’s secure cloud, we have decided to no longer support this feature” [[2]](https://techcrunch.com/2025/03/15/amazons-echo-will-send-all-voice-recordings-to-the-cloud-starting-march-28/). Retention has also been the subject of enforcement. On 31 May 2023, the US Federal Trade Commission announced a complaint, filed with the Department of Justice, alleging that Amazon “kept sensitive voice and geolocation data for years, and used it for its own purposes, while putting data at risk of harm from unnecessary access”. The proposed order required Amazon to pay 25 million dollars and to delete children’s data. Samuel Levine, Director of the FTC’s Bureau of Consumer Protection, was quoted: “COPPA does not allow companies to keep children’s data forever for any reason, and certainly not to train their algorithms” [[6]](https://www.ftc.gov/news-events/news/press-releases/2023/05/ftc-doj-charge-amazon-violating-childrens-privacy-law-keeping-kids-alexa-voice-recordings-forever). ### Google Gemini: chats retained, and some reviewed by people for up to three years Google’s Gemini Apps Privacy Hub, last updated on 24 September 2026, states that even with the activity setting off, chats are kept: “Temporary chats and chats you have when Keep Activity is off are retained with your account for 72 hours.” Human review creates a longer tier: “Chats reviewed by human reviewers (and related data like your language, device type, location info, or feedback) are not deleted when you delete your activity. Instead, they are retained for up to three years.” On voice, the page notes that “Audio recordings may begin a few seconds before activation” and that “Gemini can activate accidentally, like if it detects a noise like ‘Hey Google’” [[3]](https://support.google.com/gemini/answer/13594961). ### Apple Siri: more on-device processing, transcripts still used Apple changed its practice in 2019, after suspending a review programme in which, it said, it had reviewed “a small sample of audio from Siri requests — less than 0.2 percent — and their computer-generated transcripts”. Its stated new default: “By default, we will no longer retain audio recordings of Siri interactions. We will continue to use computer-generated transcripts to help Siri improve” [[4]](https://www.apple.com/newsroom/2019/08/improving-siris-privacy-protections/). In a statement of 8 January 2025, Apple said that “for capable devices, the audio of user requests is processed entirely on device using the Neural Engine, unless a user chooses to share it with Apple”, and that it “does not retain audio recordings of Siri interactions unless users explicitly opt in to help improve Siri” [[5]](https://www.apple.com/newsroom/2025/01/our-longstanding-privacy-commitment-with-siri/). The same statement adds that “certain features require real-time input from Apple servers”, so on-device recognition does not mean that no request leaves the device. ### Activations nobody asked for Retention policies apply to whatever the device captures, including audio captured by mistake. Dubois and colleagues played “two rounds of 134 hours of content from 12 TV shows” near popular smart speakers in the US and the UK and observed “0.95 misactivations per hour, or 1.43 times for every 10,000 words spoken”; for some devices, 10 per cent of misactivation durations lasted at least 10 seconds [[7]](https://petsymposium.org/popets/2020/popets-2020-0072.php). Google’s own page acknowledges the same phenomenon for Gemini [[3]](https://support.google.com/gemini/answer/13594961). | Assistant | Speech recognition | Audio kept by default | Transcripts and review | | --- | --- | --- | --- | | Amazon Alexa (Echo) | Amazon’s cloud; local option ended 28 March 2025 on the three devices that had it | Yes, unless the user chooses not to save; auto-delete after 3 or 18 months available | Transcripts reviewable for 30 days when recordings are not saved | | Google Gemini | Processed by Google; the page does not describe the speech step separately | Recordings may begin a few seconds before activation | Kept 72 hours with activity off; reviewed chats up to three years | | Apple Siri | On capable devices, entirely on device; some features need Apple servers | No, unless the user opts in | Computer-generated transcripts used to improve Siri | | AISir (Mac) | On the Mac (Whisper via WhisperKit) | No: audio stays in memory and is never saved | Conversation forgotten when AISir quits; only call transcripts and notes are kept, on the Mac | *Where speech is processed and what is kept, as each source describes it (checked 8 October 2026)* > AISir hears you through Whisper on your Mac: the audio stays in memory, is never saved and is not sent to a server. Gemma answers and the Chatterbox voice speaks, both on the Mac too. [Download FireAI for Mac](https://hisnlabs.com/fireai/en/download) ## Implications for Mac users Three distinctions follow from the sources. First, “we don’t keep the audio” and “we don’t keep what you said” are different promises: several policies drop recordings but keep transcripts, and a transcript carries most of the sensitive content. Second, deletion controls apply to the vendor’s primary store; Google states that reviewed chats survive the deletion of activity, so a user’s deletion and the vendor’s retention can diverge. Third, local processing is a property of a specific device and software version, and it can be withdrawn, as the Echo case shows. A setting is only as durable as the vendor’s decision to keep it. On-device recognition removes the server from the speech step entirely. It does not by itself make an assistant private: the language model that interprets the text may still run in the cloud, the answer may require a web search, and the transcript may be stored locally where other software or other users of the Mac could read it. The relevant question is therefore not “is speech recognition local?” but “which of the four stages leave the device, and what is stored, where, for how long?”. ## Recommendations 1. For each assistant you use, find its privacy page and note separately what happens to audio, to transcripts and to chats reviewed by people. 2. Turn on the shortest automatic deletion the service offers, and check whether deleting your history also deletes reviewed copies. 3. Mute or unplug smart speakers in rooms where confidential conversations take place; misactivations are documented. 4. Prefer assistants that state which stages run on the device, and re-check after major updates, since local options can be withdrawn. 5. Watch the network: an app that says it works offline should open no connections while you speak. ## Relevance to AISir AISir 1.0.1 is a planner, quick bar and voice assistant for macOS 15 on Apple silicon. According to its Help, every stage of a spoken conversation runs on the Mac: “Silero detects speech, Whisper writes it down, Gemma answers and Chatterbox speaks. Audio is kept in memory only and never saved.” The language model is Gemma 4 E2B, run with MLX. A spoken conversation supports barge-in (talking over AISir stops it at once) and ends when the user says goodbye. The conversation is shown on screen and forgotten when AISir quits. A voice note (hold ⌥⇧Space) is transcribed by Whisper on the Mac, and the recording is dropped right after it is written down. The only times AISir goes online by itself are the one-time model downloads, each file checked against its checksum. Web lookups are off by default; when turned on, they send short, neutral search terms to Wikipedia and Wikidata, or to Brave Search with the user’s own key, directly from the Mac. Details are on the [AISir page](https://hisnlabs.com/aisir/en). AISir does not remove every limit described above. Its answers take about four seconds, and the small model makes mistakes, which is why tasks it prepares wait for Accept, Edit or Dismiss. Arabic support is weaker than French and English in places. It runs only on Apple silicon. Call transcripts and notes, when the user chooses to take them, are kept on the Mac, so the Mac’s own security, such as FileVault and the user account, still matters. And AISir cannot change what other assistants on the same Mac or in the same room record. > AISir 1.0.1 is out: talk to it, interrupt it, say goodbye, and nothing you say is recorded to a server. Every feature is open for 17 days, with nothing to sign up for. [Download FireAI for Mac](https://hisnlabs.com/fireai/en/download) ## Limitations This note compares published policies, not measured behaviour; none of the vendors’ statements were verified by traffic analysis here. Policies change: the Amazon page shows no date, and the Google page was updated twelve days before this note. The OpenAI help page on voice-mode retention could not be opened and is not described. The misactivation study dates from 2020 and used television audio, not live households. AISir’s description is taken from its own Help and product page, and HisnLabs makes AISir, which readers should weigh. Related reading: [local LLMs on Apple silicon](https://hisnlabs.com/en/blog/local-llm-apple-silicon-security) and [taking back camera and microphone access](https://hisnlabs.com/en/blog/mac-camera-microphone-permissions-take-back-access). ## How FireAI and HisnLabs fit in Whatever an assistant’s privacy page says, the network shows what it does. FireAI lists, per app, every connection your Mac opens, so you can see whether a voice or AI app talks to a server while you speak, and block it if you want. FireAI is HisnLabs’ own product: an on-device AI firewall for Mac. It shows every connection your apps make, in plain language, and lets you decide what leaves your Mac — its AI runs locally, so your traffic is never sent to us or anyone else. HisnLabs’ security research team is the group that keeps that decision-making accurate: cataloguing which domains are ordinary telemetry versus a real product, tracking the country and network behind a connection, and training the on-device model (its FireAI Pilot feature) on real traffic patterns, all without any of it leaving your Mac. You can read the technical decisions behind it, or try FireAI for 17 days, at [FireAI, by HisnLabs](https://hisnlabs.com/fireai/en/download). ## Sources - [Amazon: “Alexa makes privacy even easier”, About Amazon (no date shown on the page)](https://www.aboutamazon.com/news/devices/alexa-makes-privacy-even-easier) - [TechCrunch: “Amazon’s Echo is ending its ‘Do Not Send Voice Recordings’ feature, starting March 28”, 15 March 2025](https://techcrunch.com/2025/03/15/amazons-echo-will-send-all-voice-recordings-to-the-cloud-starting-march-28/) - [Google: Gemini Apps Privacy Hub, last updated 24 September 2026](https://support.google.com/gemini/answer/13594961) - [Apple Newsroom: “Improving Siri’s privacy protections”, 28 August 2019](https://www.apple.com/newsroom/2019/08/improving-siris-privacy-protections/) - [Apple Newsroom: “Our longstanding privacy commitment with Siri”, 8 January 2025](https://www.apple.com/newsroom/2025/01/our-longstanding-privacy-commitment-with-siri/) - [US Federal Trade Commission: FTC and DOJ charge Amazon with violating children’s privacy law by keeping kids’ Alexa voice recordings forever, 31 May 2023](https://www.ftc.gov/news-events/news/press-releases/2023/05/ftc-doj-charge-amazon-violating-childrens-privacy-law-keeping-kids-alexa-voice-recordings-forever) - [Dubois, Kolcun, Mandalari, Paracha, Choffnes and Haddadi, “When Speakers Are All Ears: Characterizing Misactivations of IoT Smart Speakers”, PoPETs 2020 (4)](https://petsymposium.org/popets/2020/popets-2020-0072.php) - [Radford et al., “Robust Speech Recognition via Large-Scale Weak Supervision” (Whisper), arXiv:2212.04356, December 2022](https://arxiv.org/abs/2212.04356) - [Orhon et al., “WhisperKit: On-device Real-time ASR with Billion-Scale Transformers”, arXiv:2507.10860, July 2025](https://arxiv.org/abs/2507.10860) - [HisnLabs: AISir product page](https://hisnlabs.com/aisir/en)