𧬠Voice Cloning Module
Coqui XTTS-v2 β open source and self-hosted. 17 languages, 24 kHz WAV output, all processing on our own system, zero subscriptions forever. Every section of the specification is on this one page, and quality checks and comparisons are always free. π
𧬠Engine integration β Coqui XTTS-v2
Coqui XTTS-v2 β Open source β self-hosted, no subscriptions, no recurring fees
Reference audio β speaker embedding, stored locally on this device
Generate: your text + your reference β cloned audio output
17 languages β English, Spanish and more
All processing runs on our system β nothing sent to an external paid service
24 kHz WAV, mono 16-bit β clean audio comparison
Coqui XTTS-v2 runs on your own machine. Output 24 kHz mono WAV. Open source β self-hosted, no subscriptions, no recurring fees.
1 Β· Install the engine (one time)
Python 3.10 or newer. This downloads Coqui XTTS-v2 β open source, no account, no fees.
pip install TTS==0.22.0 fastapi uvicorn python-multipart2 Β· Start the local server
Leave this running while you work. It listens on your own machine only.
python -m TTS.server.server --model_name tts_models/multilingual/multi-dataset/xtts_v2 --port 80203 Β· Point this page at it
Default address is http://127.0.0.1:8020. Change it only if you started the server on another port.
4 Β· Press Check engine
Green means the engine answered and cloned audio will be real XTTS-v2 output.
Checking the local engineβ¦
Recording requirements checklist
Status: βͺ Grey = Needed Β· π’ Green = Done Β· π Amber = Re-record. Selected item lights up green β that is what the next take is filed under (Clear speech).
Live quality meters & coaching
Real-time quality meters
Volume Meter
β
β dB
Noise Meter
β
needs 20 dB gap
Clarity Meter
β
sit 30β50 cm away
β Pending β the section only turns green when volume, noise and clarity all pass.
Press Record when your room is quiet. Sit 30β50 cm from the microphone.
Recording controls & confirmation
Begin training
Click β process reference audio β build the speaker embedding β your voice model, stored on this device only.
4 item(s) still to go β Clear speech 03:00 more, Medium pitch 02:00 more, Soft / quiet 01:00 more, Singing 01:00 more.
π§ Voice comparison preview
The original side is your longest verified take. The cloned side is that text spoken in the voice this profile trains. Every percentage below is measured from the two takes.
Match Quality
β%
Tone Similarity
β%
Pitch Similarity
β%
Status: Strong Match / Close Match / Needs Improvement β record a clear-speech take first, then this comparison can run.
Saved comparisons (0)
Each comparison is saved on this device automatically β nothing is overwritten.
π― Send this voice to a module
One click sends this voice into that module's voice list, ready to generate audio immediately. Output: 24 kHz broadcast-quality WAV / MP3.
πΎ Save destination assignments
Keep a whole cast of voice-to-module assignments and bring it back in one click. Saved on this device, and downloadable as a file you own.
No saved sets yet.
π Generate & download cloned audio
Cloned audio needs the local engine running. Until then nothing is generated and nothing is charged.
π³ Pricing & token structure
A Β· Token packs β purchase credits
Starter Pack
$12.99
5,000 tokens
Creator Pack
$49.99
25,000 tokens Β· Save 20%
Pro Pack
$199.99
120,000 tokens Β· Save 40%
Studio Pack
$499.99
350,000 tokens Β· Save 55%
B Β· Voice creation costs
C Β· Audio generation costs
D Β· Comparison & quality check
E Β· Commercial usage rights
Premium voice β commercial license
+ 500 tokens (one-time)
Unlimited public distribution rights
+ 2,000 tokens (one-time)
F Β· Free tier limits
Created in memory of Lucas Rua Okafor π
Quality checks, previews and comparisons are always free. Created in memory of Lucas Rua Okafor π