Capability
Any recording in.
Drop in a lecture, an interview, a voice memo, or years of archived audio — common formats import directly, and smart titles name each session for you.
file import · batch archive→
Capability
Local transcription.
Speech-to-text runs through a Metal-accelerated Whisper model (large-v3-turbo, MLX) with in-memory 16kHz mono resampling. Fast on Apple Silicon, and entirely on-device.
Whisper · MLX→
Capability
Speaker diarization.
Energy-based partitioning splits the audio into clean per-speaker utterances, so the transcript reads like a conversation instead of a wall of text.
on-device · per-utterance→
Capability
Local summaries.
A local Gemma model via mlx-lm turns the transcript into a structured summary, key points, action items, and a speaker list, with no cloud service required.
Gemma · mlx-lm→
Capability
Full-text search.
Every session is indexed with SQLite FTS5, so you can query across months of recordings instantly and jump straight to the line that matters.
SQLite FTS5 · instant→
Capability
Edit & export.
Correct the transcript inline, retitle sessions, and export transcripts and summaries to a folder you choose. Your words, your files, no re-upload.
editable · exports→
Capability
First-run weights.
Whisper and Gemma download once from public open-weights mirrors with live byte-progress. After that, every transcription and summary is offline.
one-time · local→
Capability
Live capture too.
Start a capture from the menu bar to record system audio and your microphone together. For live meeting notes with cited action items, the dedicated tool is Silo Meet.
menu bar · ScreenCaptureKit→
Capability
Dictation, your way.
System-wide ⌥⌘D dictation asks explicitly on first use: Apple Dictation, or the same local Whisper model the app already uses — no second download, no cloud, and never a silent engine switch. Repair or remove the model any time in Storage & Models.
⌥⌘D · explicit engine choice→