How to clone your own voice with Fish Audio

Fish Audio copies your voice from a ten-second clip for free, and every tutorial skips two things: the free plan's voice slots are public, and free output is for personal use only. Here is the setup, a recording script that gives it good audio, and the disclosure line for your captions.

How-to

What you get, and the two catches

The current model is S2.1 Pro, shipped in June 2026, with 83 languages. A clone needs at least ten seconds of one clean speaker, and a minute or two improves it. Prices below are in US dollars from the plan page, read on 16 September 2026.

  • Free, $0. 8,000 credits, up to seven minutes of generated audio, 500 characters per generation, three voice slots.
  • Plus, $11 a month on the current offer, $15 regular. 250,000 credits, 200 minutes, 15,000 characters per generation, ten private voice slots, commercial use, and one professional voice.
  • Pro, $75 a month. 1,620 minutes, five professional voices, three seats. Only if this becomes your job.

There is also an open-source version with over 32,000 stars, but it needs a 24GB graphics card and a terminal, and its licence bars commercial use without a separate deal. The web app is the route.

The setup, step by step

Every step is from Fish's own cloning guidance and docs. Button labels can move, so if one has changed, look for the same job.

  1. Record one to two minutes. Use the script below, in a quiet room, phone or mic about a hand's width from your mouth, half-second pauses between sentences. Record the energetic take as a separate file.
  2. Sign up and open the cloning page. No card is needed for the free plan. Go to fish.audio/app/voice-cloning, upload the file, and name the voice.
  3. Decide public or private before you save. The plan page lists the free tier's three voice slots as public, with private slots from Plus, while the developer docs describe a private default, so check the visibility setting. Test on free and delete the voice when you are done, or take Plus now if a public slot bothers you.
  4. Tick the consent box if one appears. Fish's rule is your own voice, or someone who gave you written permission. Nothing else.
  5. Generate a first line. Paste up to 500 characters, add a bracket tag such as [pause] or [excited] where you want it, generate, and listen all the way through.
  6. Download and drop it in. MP3 at 128 kbps for a reel or podcast, WAV if you will edit it. Put it on the audio track of whatever you already edit in, then paste the disclosure line into the caption.
  7. Before you sell anything, pay. The Free card lists commercial use, and the FAQ on the same page says free output is personal and non-commercial, so treat the FAQ as the rule. A monetised reel, a paid audiobook or a sponsored podcast needs Plus at least.

If the first line sounds wrong

  • Robotic or flat. Record longer, 30 to 60 seconds more, and keep it in one even mood.
  • Echoey or boxy. Fish's docs name reverb as a thing to avoid. Re-record somewhere soft, a curtained room or a parked car, rather than trying to fix it in the app.
  • Wobbles mid-line. The source swung in volume or emotion. Keep the everyday take even and save the energy for take two.

Tags: one that lands, one that misses

Lands: Quick one before you go. [pause] The link is in the caption.

Misses: [excited] [laughing] Quick one [pause] before [sigh] you go.

One tag per sentence, at a natural break, and listen back, because tags land sometimes and miss sometimes.

The recording script

Read this, in one take, and the clone gets varied sentence lengths, a question, a list and a soft ending, which is what it needs to sound like you on anything.

ScriptThe 90-second recording script
HOW TO RECORD THIS
Quiet room, phone or USB mic about a hand's width from your mouth, one take, no music. Read at your normal talking pace, and where you see (pause), stop for half a second. If you stumble, keep going; a clean minute matters more than a perfect one. Save as .wav or .mp3, one speaker only.

TAKE ONE, your everyday voice (about 90 seconds)

Hi, this is [your name], and this is what I sound like when I'm just talking. (pause)
Most mornings start the same way for me. I make something hot to drink, I open the window for a minute even when it's cold, and I look at the list I wrote the night before. (pause)
Some days the list has three things on it. Other days it's got eleven, and I know before I start that four of them are going to roll over to tomorrow. That's fine. (pause)
Here is a question I keep coming back to: what would I do with an extra hour, if it turned up unannounced? (pause)
Probably nothing impressive. Read something slowly. Call someone I've been meaning to call. Sit outside. (pause)
The things I say most often in a day are small ones. Yes, that works. Can you send it again? Let me think about it and come back to you. (pause)
And when I'm explaining something I care about, I slow down a little, I use shorter sentences, and I let a point land before I move to the next one. (pause)
That's enough for now. Thanks for listening.

TAKE TWO, your energetic voice (about 30 seconds, record as a separate file)

Okay, this is the version of me that's had coffee. (pause)
Big news today, and I want to get straight into it, because there are three things you need to know and the first one changes everything. (pause)
Ready? Here we go.

The settings checklist

Everything the interface asks for, in the order it asks, so the first clone is the good one.

ChecklistThe settings checklist
BEFORE UPLOAD
- File: .wav or .mp3 (.m4a and .opus also accepted), mono, one speaker, no music
- Length: at least 10 seconds; 1 to 2 minutes is the sweet spot
- Two files if you want two moods: everyday, energetic
- Upload the energetic take as a second voice called [your name, energy], and pick a voice by mood rather than by tag

IN THE APP
- Name the voice something you will recognise in a list: [your name, take one]
- Visibility: private, unlisted or public; the free plan's slots are listed as public, private starts on Plus
- Enhance audio quality: leave it on (it strips noise and evens levels before training)
- Consent: tick any box that asks whether the voice is yours or used with permission

WHEN YOU GENERATE
- Text per generation: 500 characters on the free plan, 15,000 on the first paid plan
- Emotion in square brackets inside the text where you want it: [pause] [whisper] [excited] [laughing] [sigh]
- Speed: 0.5 to 2.0, leave at 1.0 for the first run
- Listen once all the way through before you download

EXPORT
- MP3 at 128 kbps for reels and podcasts, WAV for editing
- Drop it on the audio track, then add the disclosure line to the caption

BEFORE ANYTHING IS SOLD OR MONETISED
- Move to a paid plan: free output is for personal, non-commercial use

The disclosure line

One line, three versions, and the platform toggle. Put it where the audio goes out, every time.

TemplateThe disclosure line
PICK THE LINE FOR WHERE IT GOES, THEN FILL THE BRACKETS

Video caption or description:
Voice: my own, cloned with AI ([tool name]). Script and edits by me.

Podcast episode notes:
This episode's narration uses an AI copy of my own voice, made with [tool name] from recordings I made myself. The words are mine.

Audiobook or long read, front matter:
Narrated by an AI clone of the author's own voice, created with [tool name] with the author's consent. No other person's voice was used.

Where the platform has an AI label or toggle, switch it on as well: [YouTube / TikTok / Instagram / Spotify / other].
Where you live: [add any disclosure your country requires for synthetic audio; in the EU, artificially generated audio must be disclosed].

The studio version, if you want it

Professional Voice Cloning, launched 15 June 2026, wants 10 to 180 minutes of audio, ideally 12 to 15 clips of 45 to 60 seconds, and trains for one to two hours. Before it trains you pass a live ownership check, reading a passage on screen. Fish says only the actual speaker, live, can pass it. Plus includes one.

What each platform asks

  • YouTube. Its disclosure rule lists cloning your own voice for voiceovers or dubs as not needing a label. Realistic altered audio of anyone else does.
  • Spotify. Since 19 May 2026 it removes podcasts that impersonate another creator's voice, AI clone or otherwise. Your own voice, your own show, is fine.
  • The EU. Article 50 of the AI Act asks for artificially generated audio to be flagged, in force since 2 August 2026. The duty falls on providers and on anyone deploying deepfake audio, so check whether it covers you, and give EU listeners the line anyway.
  • Phone calls, in the US. The FCC ruled on 8 February 2024 that AI-generated voices count as artificial under robocall law, so never put a clone on an outbound call without consent.

The honest bit

  • Someone can clone you without asking. Estelle Hubert, a French voice-over artist, found a clone of her voice on Fish Audio and wrote up how she got it removed, updated 4 August 2026.
  • Removal took three moves. The three-dot menu, Report model, Copyright complaint, then a legal email, because deactivation on its own left the data in place. Read it before you upload anything.
  • It is good, with tells. Dragos at Bitdoze used his clone for a week of YouTube narration: "good enough that my wife could not tell the difference in a blind test", tags land sometimes and miss sometimes, and long generations crawl in the browser.
  • The dates move. The free API for the current model runs to 30 November 2026 on a fair-use basis, and the terms were last updated in August 2024, so re-read the plan page before you build a workflow on it.

Record the script tonight

Ninety seconds in a quiet room is the whole first step, and it is the step that decides the quality. Upload it tomorrow, generate one line, and listen for the tell before you decide whether it is worth $11.

If the words are the problem rather than the voice, teaching AI to write like you comes first.

A few quick questions

Do I have to label a clone of my own voice on YouTube?

No. YouTube's disclosure rule lists cloning your own voice for voiceovers or dubs as something that does not need a label. Realistic altered audio of someone else does.

What if the clone sounds flat?

Give it more to learn from: another 30 to 60 seconds in one even mood, recorded somewhere soft. Reverb and swings in volume are the two things Fish's own docs say to avoid.

Can it speak another language in my voice?

The current model lists 83 languages, and one creator's test of mixed English and Romanian came out with clean pronunciation. Try one short line first, because accents and long reads are where a clone is weakest.