Our own speech engine

Every voice your work needs. In seconds.

Generate lifelike speech, clone voices from a single clip, convert one performance into another, transcribe, dub and publish — from script to finished, mixed audio.

10,000
free characters a month
1 clip
to clone a voice
44.1 kHz
WAV masters, no watermark
The Working Parts

What it does, and what it refuses to fake

We run the speech engine ourselves, so a clone works the moment it uploads and no vendor takes a cut of every render after. These cards name the numbers — and the limits that come with them.

Bulletins at 4:20am

Draw your coverage area, set your dayparts, and traffic bulletins write, voice and deliver themselves. An empty incident list gets the clear-roads line you wrote — it never invents one.

Warnings lead first

Weather bulletins lead on the active warning rather than the weekend outlook, in your market's own units and clock.

Five ways into playout

Bulletins land by pull URL, webhook, FTP, SFTP or Dropbox — overwrite one filename for a watch folder, or stamp each one to archive it.

Cuts through the bed

Voice and bed through a real broadcast chain — 75 Hz shelf −6, 400 Hz −3, 6 kHz +4, then compressor and limiter — ported from one already going to air. Mixing costs no characters.

Fix one word

Click the word that came out wrong, retype it, and it re-renders in the same voice and splices back with a crossfade. The rest of the take is untouched.

Undo never re-bills

Cuts, fades and gain are free and endlessly reversible, and undoing a re-rendered word re-splices rather than paying twice. Saving writes a new clip, so the master survives forty edits.

Names said right

Write the respelling once — Decatur becomes Deekayder — and studio takes, API calls, project blocks and the 4am bulletin all inherit it.

Voice a fix

Flag a mangled word, say it properly into your mic, and it comes back in the voice that got it wrong. The engine's mistake, so it is on us rather than your quota.

Written to the slot

Copy cut to a real word budget — 72 to 84 words for a :30 — and measured again on the way back, from an email, a web page or a document.

Forty tags, one script

Put a tag in the copy, paste the rows out of your sheet, and get one take each — up to 500, priced before you commit and delivered as a project you can retry row by row.

Fix one paragraph

A book is chapters and paragraphs, each rendered on its own, so a typo in chapter forty re-reads one paragraph. It renders on a queue, so you can close the tab.

Rescue a bad recording

Four purpose-built chains rather than one Enhance button — the interview chain keeps the pauses, because in an interview the pause is often the answer. Costs no characters.

The tags we refuse

A pause marker gives you exactly the silence you ask for, to a tenth of a second. We do not offer laugh or sigh tags — we measured them on this engine, and they do not work.

Pin the take

Every take shows the seed it came from. Pin the one the client approved and it comes back identically on any later render.

No vendor per take

Zero-shot cloning on an engine we run ourselves: one clean clip, usable the moment it uploads, with nobody taking a cut of every render after. 200 voices ship with it.

Failed takes cost nothing

Quota is claimed before a render starts and refunded the moment one fails, so you pay for finished work. Transcription bills at a tenth of the synthesis rate.

It asks first

Describe the job and the agent drives the same API you can, at your own role, proposing anything destructive for you to confirm. Every call is logged, including the refused ones.

One spec, no drift

The reference, the code samples, the playground and the OpenAPI document all generate from one spec, so they cannot disagree. One key drives both REST and MCP.

…and a REST API for all of itSee the API
How it works

Three steps, and the first one is optional

  1. 1

    Choose or clone a voice

    Start from the premade library, or upload a clip and have your own voice ready before the kettle boils.

  2. 2

    Write, upload or point at a file

    Type a script, drop in a recording to re-voice, or push text through the REST API from your own pipeline.

  3. 3

    Ship it

    Download the master, export a whole project, or let a webhook tell your system the render has landed.

For creators

The publishing work that isn’t performing

Sponsor reads, pickups, dubs, back-catalogue narration — the audio work between recordings. Clone your own voice once and the workspace handles the takes you’d rather not re-record, for a podcast feed, a channel or an audiobook.

1

Clone yourself once

One clean clip and your voice lives in the workspace — sponsor reads, intros and corrections without setting the mic up for a two-line fix.

2

Fix the flub, keep the take

Retype the word that came out wrong and it re-renders in the same voice. The episode keeps its energy; you keep your afternoon.

3

Publish past your language

Send an English episode, audio or video, and dubbing returns it re-voiced in another language. Same cut, new audience.

4

Ship the back catalogue

Projects split a manuscript into chapters, render while the tab is closed, and export one clean master.

Start free
For broadcasters

Traffic on the fives, with nobody in the building

A live incident feed for the roads your listeners drive, ranked down to the few that are worth airing, written in your station’s wording and read in your station’s voice — on the schedule you set, whether or not anyone is on shift.

1

Select your coverage area

Draw the circle your signal covers and name the roads in and out of it. Those are the names a bulletin leads on.

2

Choose a voice, or clone your talent

Any voice in the workspace reads it — including a clone of the presenter who would have read it live.

3

Customize it for your station

Your intro, outro and sponsor line, how many incidents make the cut, and what to say when the roads are clear. Paste in old bulletins as house style.

4

Set your schedule

Dayparts rather than one cadence — every ten minutes through drive, hourly overnight — each built from what the feed says at that moment.

And the weather, on the same rails

Pin a market and the forecast gets the same treatment: warnings first, your units, your dayparts. Traffic and weather share the delivery machinery, so setting up one sets up the other.

Bulletins are metered like everything else: the finished script's characters come off the same monthly allowance as the studio and the API, and a failed render is never billed. No per-minute rate, nothing extra per coverage area.

Set up a coverage area
For publishers

The manuscript is the input

An ebook, a course, a back catalogue nobody ever recorded — long-form projects turn a manuscript into a narrated master without booking a booth, and every step of it is also reachable from the API.

1

Split it into chapters

Paste the manuscript and it becomes chapters and paragraph blocks, each rendered on its own.

2

Cast the narrator

Any voice reads it — a premade narrator, a designed one, or a clone of the author — and blocks can switch voices.

3

Teach it the names

The pronunciation dictionary applies to every block, so the protagonist's name is right in chapter one and chapter forty alike.

4

Render while the tab is closed

Batches run on the server and a webhook says when the book is done. Export one master, or pull chapters through the API.

Narrate a chapter free
Pricing

Pay for characters, not for seats

Every plan includes the whole studio. What changes is how much you can generate, how many voices you can keep, and — for batches, projects, mixes and dubs — where that work sits in the queue when the platform is busy. Invite the whole team on any of them.

Free

Free

Try the whole platform on a small monthly allowance.

10,000
characters / month
  • 10,000 characters per month
  • 3 custom voices
  • Instant voice cloning
  • Speech to text
  • Non-commercial use
Start free

Starter

$5/month

For solo creators shipping regular work.

30,000
characters / month
  • 30,000 characters per month
  • 10 custom voices
  • Commercial licence
  • Dubbing
  • API access
Choose Starter

Creator

Popular
$22/month

Higher quality settings and professional voice cloning.

100,000
characters / month
  • 100,000 characters per month
  • 30 custom voices
  • Professional voice cloning
  • Long-form projects
  • 44.1 kHz WAV downloads
Choose Creator

Pro

$99/month

For studios and production teams with a queue to clear.

500,000
characters / month
  • 500,000 characters per month
  • 160 custom voices
  • Everything in Creator
  • Priority in the render queue
Choose Pro

Scale

$330/month

Volume production with room to grow.

2,000,000
characters / month
  • 2,000,000 characters per month
  • 660 custom voices
  • Everything in Pro
  • Top priority in the render queue
Choose Scale

Prices in USD. Characters are counted on the text you send; failed renders are never billed.

Queue priority is a head start, not a separate lane, and it applies to work that goes on the background queue — batches, long-form projects, mixes, dubs, multivoice documents and voice previews. A render you start in the studio, and a call to the text-to-speech API, run inside their own request and never queue, so those are the same speed on every plan.

Where it does apply it is bounded on purpose: within a class of work, no job is ever overtaken by work queued more than 8 minutes after it, so free-plan work still drains rather than waiting behind an endless paid queue. Class comes first and plan only breaks ties inside one, so nothing anybody pays lets a background sweep outrank somebody else’s live render. It orders work and nothing else: no plan on this page carries an uptime commitment, a service-level credit or a support response time. (Account credit, the dollar balance a promotional code grants, is a different thing entirely — it offsets a charge and promises nothing about availability.) The Terms of Service are what govern.

For developers

One key, one endpoint, an MP3 back

The API is the same engine the studio uses — same voices, same settings, same quota. Issue a key from the workspace, keep it in your environment, and render from anywhere.

  • Keys are hashed at rest and scoped per surface — tts, sts, stt, voices.
  • Every request is metered against the workspace, whoever made it.
  • Webhooks fire when a long render finishes, so nothing has to poll.
Get an API key
POST /api/v1/text-to-speech/{voice_id}
curl -X POST https://voicebuddy.radioworkflow.com/api/v1/text-to-speech/{voice_id} \
  -H "xi-api-key: $VOICEBUDDY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "text": "Episode three is live. [pause:0.8] Here is what changed.",
    "voice_settings": { "stability": 0.5, "similarity_boost": 0.75 },
    "output_format": "mp3"
  }' \
  --output take-01.mp3
Voice Buddy

Hear your first line in about a minute

Ten thousand characters a month on the free plan. No card, no sales call, no watermark on the audio.