A voice, not a mix.
Every file, link and recording is run through vocal isolation as it arrives, using a Core ML conversion of HTDemucs v4. The result is levelled to a consistent loudness, so a reference sounds like the speaker rather than the room they were in. Turn it off for material that is already a dry voice.
Hear it, scrub it, trim it.
The isolated take opens as a scrubbable waveform. Drag either edge to set the in and out points, play just the selection, and watch the clone-readiness score follow your trim. Five to ten seconds of clean speech clones better than a long take, and the app tells you when a clip is too quiet, too short, or mostly silence.
Attested, recorded, carried.
The app cannot be used until you affirm you will only clone voices that are your own or that you have explicit permission to use, and every voice you create asks again. That attestation is stored with the voice, and provenance metadata travels inside every voice you export.
Import an audio or video file, paste a direct media file URL, or record straight into the app through any input device you choose.
High fidelity continues from the reference audio and its transcript. Fast uses a compact speaker embedding. Switch per take and keep whichever you prefer.
Pick input and output devices the way you would in Audio MIDI Setup, choose where multi-gigabyte weights are stored, and point the app at compatible models you already have.
Publish a voice to the shared catalog and the Silo and Ohms apps installed on your Mac can speak with it. Export a voice as a portable bundle at any time.