If you've ever fed a script into a text-to-speech tool and cringed at the playback, you know the problem: it reads the words, but it doesn't feel them. ElevenLabs v4 and v4 Turbo are built to close that gap. The new voice models let you steer emotion, pace and delivery straight from your script, so one block of text can shift from a warm explainer tone to a hushed aside — no microphone, no second take, no voice actor on standby.
What happened
ElevenLabs announced two new versions of its speech generation models: Eleven v4 and v4 Turbo. The headline feature isn't a prettier voice — it's control.
The mechanism is refreshingly low-tech: tags. You drop short instructions into your script, and the model treats them as stage directions. A tag can shift the tone of a line, push a sentence into a whisper, or add a different intonation entirely. Think of the difference between handing an actor a wall of dialogue and handing them the same dialogue with notes scribbled in the margin.
The update also holds up on long texts, which matters more than it sounds. Emotion used to collapse somewhere past the third paragraph, leaving you with a robotic drone for the rest of the piece. And the models cover 90+ languages, including Russian, with the ability to switch between them without the awkward seams you normally hear when a synthetic voice changes language mid-sentence.
If you want to hear it before you read another word, ElevenLabs has published a demo, and the models are free to test.
What it means for you
This isn't a tool for engineers tuning parameters. It's a tool for anyone who has ever needed a voice and didn't have one. Here's what that looks like in practice.
At home
Bedtime stories are the obvious one. You can turn a favourite book chapter into narration that actually sounds like storytelling — slowing down at the tense part, dropping to a whisper at the punchline. The same trick works for accessibility: long articles, PDFs and newsletters become listenable on a commute or a walk, and pacing tags stop the audio from droning you to sleep before the third section.
At work
Training videos and internal onboarding decks are where this pays off fastest. You write the script once, generate the voice, and when a process changes you edit one line and regenerate it. No rebooking a voice artist for a 40-second fix, no studio time, no waiting three weeks for a revised file.
For your business
Product explainers, ads, phone menus, e-learning modules — all of it can come from one script. The 90+ language support is the real unlock here: a single campaign can go out across multiple markets without hiring a separate native voice for each one. Seamless language switching means a bilingual video doesn't sound like two different recordings stapled together.
For studying
Students can turn dense reading lists into audio and actually get through them. Language learners can hear a phrase delivered calmly, then excitedly, then whispered, which is far closer to how people really talk than a flat textbook recording. Teachers can produce listening material without recording it themselves at 11pm.
For creative projects
Podcast intros, game characters, indie film scratch tracks, dialogue scenes — the emotional range is suddenly wide enough for actual performance. A single script can carry a confident narrator and a nervous character without switching tools.
As a side income
This is the one to watch. Narration for YouTube channels, audiobooks for self-published authors, ad reads for small local businesses — the work is real and the barrier just dropped. The money isn't in pressing generate; it's in the editing, the pacing and the ear for what sounds human. If you're building that kind of content workflow, the free tools on mykreatool.com can cover everything around the voice work.
How to try it right now
You can do this in under ten minutes, and you can do it for free.
1. Start with the free option. Open ElevenLabs Text to Speech and test the models at no cost before you commit to anything.
2. Paste a short script. Three to five sentences is plenty for a first pass. Don't start with a whole ebook.
3. Pick a voice. Browse the library and choose one that fits the mood you're after — warm, crisp, dramatic, whatever the piece needs.
4. Add your tags. Write your delivery directions directly into the text where you want them: a tone change, a whisper, a different intonation. The exact tag syntax lives in the model documentation, and a two-minute skim will save you a lot of guessing.
5. Choose your model. Eleven v4 or v4 Turbo — the details are laid out on the v4 page.
6. Set the language. With 90+ supported, you can run the same script in a second language and switch mid-flow.
7. Generate, listen, adjust. Regenerate as many times as you like. This is the part that used to cost money and now costs seconds.
8. Hear the target first. If you want to know what good sounds like before you start tweaking, watch the official demo.
Upsides and what changes
Studio-grade emotion used to require a studio. Now it requires a well-written script and a bit of patience with tags. Iteration is nearly free — changing one line no longer means re-recording the whole thing. One script scales across dozens of languages. Long-form content finally holds its emotional shape from start to finish.
The bigger shift is where the skill goes. When anyone can generate a decent voice, the advantage moves to whoever writes better copy and directs it better. That's good news if you're a writer or a marketer, and less good news if your only selling point was access to a booth.
Limitations
Be honest with yourself about the ceiling. The tags are only as good as the script feeding them — vague writing produces vague delivery, no matter how many instructions you stack on top. Genuine subtext is still hard: sarcasm, grief, the pause that means more than the sentence, the breath before a lie. You'll also spend real time on trial and error, since syntax and results vary between v4 and v4 Turbo and quality across 90+ languages won't be perfectly even — some will sound native, others will carry a slight accent. Names, jargon and acronyms still need proof-listening or they'll come out wrong. And none of this touches the legal and ethical side: if you're cloning someone's voice, you need their permission, and audiences increasingly expect to know when audio is synthetic. Always listen to the full track before you publish — one misplaced emphasis can turn a friendly line into a threatening one.
Conclusion
ElevenLabs v4 and v4 Turbo take AI voice from "read this text out loud" to "perform this text the way I want it." Emotion, pace and style are now editable fields, not fixed settings. You get 90+ languages, long-form consistency and free testing, and the only real bottleneck left is how well you write and direct.
Here's your one action for today: open the free Text to Speech page, paste five sentences of something you've already written, and add a single whisper tag to one line. If that one line lands, you've just found a new way to make things.



Comments 0