Everything that has shipped, and what it cannot do
Dated to the day the code landed, written from the code rather than from a commit message, and carrying the limits — a release note that only lists wins is an advertisement.
- Updates
- 36
- Areas
- 7/ 7
- Shipping days
- 4
3 updates
The console in thirteen languages
HighlightThe interface reads in thirteen languages, chosen from a picker in the sidebar that names each in its own script — Deutsch, 日本語, العربية — with Arabic laying the whole shell out right to left. Make no choice and it follows what your browser asks for.
It changes the console only, not the voices you render with or the language your audio is spoken in.
Platform#Weather bulletins
Drop a pin on a market and weather bulletins run on the same dayparts, voices, house wording and delivery targets as traffic. An active warning always leads, so a listener who hears only the first sentence hears the warning, and the bulletin reads the agency's own wording rather than improvising safety advice of its own.
Today plus three days ahead, and no further.
Broadcast#Pull URLs and webhooks deliver
Two ways of getting a finished bulletin out to playout did not work. Every pull URL a station installed answered a 404, because nothing served the address the panel printed, and webhook targets failed outright because the bulletin events were never offered for subscription.
Fixes#
11 updates
Non-destructive voice editor
Open any finished read as a waveform with its words underneath, then cut, trim, fade, ride the level or drop a pause in. Nothing is written until you save and saving mints a new clip, and a gap you silence is filled with the clip's own room tone rather than with digital silence.
Up to 80 edits in a single save; past that a read wants re-recording rather than editing.
Studio#Audio cleanup
Drop in a rough recording — a phone tag, a field interview, a client's voice memo — pick one of four presets and get back a conditioned file: rumble out, noise down, sibilance tamed, dead air trimmed and the loudness landed on -14, -16, -19 or -23 LUFS. It prints the loudness it measured going in and coming out, and spends no characters at all.
This is not source separation: music under a voice comes back quieter, cleaner and still there. 50 MB and twenty minutes per run.
Studio#Languages honestly offered
Choosing a language for a read now tells you whether it will actually be performed in that language. Where the running speech model would ignore the choice and hand back an English read of the words — billed in full, with nothing in the audio to say why — the option is switched off carrying its own reason, and the public text-to-speech endpoint refuses the same request rather than charging for it.
Where nothing can confirm which model is loaded the picker warns instead of blocking, so render one take and listen before committing a batch.
Studio#Batches survive a closed tab
Generate all runs on the server, so you can close the tab and come back to a finished project. A re-run renders only what changed: each finished block stores a fingerprint of its text, voice and settings, so fixing one paragraph re-renders that paragraph and nothing already correct is paid for twice.
A run that reaches the end of the workspace's allowance stops there and resumes from the same block once there is more.
Long-form#Bulk Copy
Forty dealer tags off one script: paste the spreadsheet, write the line once with {{column}} where the variables go, and get a take per row. A copy straight out of Excel works, since tabs, semicolons and pipes are read as readily as commas, and a voice column casts each row by name.
500 rows to a batch; rows past that are reported back rather than quietly dropped.
Long-form#Scheduled traffic bulletins
HighlightDraw a coverage area, pick a voice, set the dayparts, and traffic bulletins go to air with nobody in the building. Each run polls the live incident feed, writes the copy, renders it and hands it to your playout system, and the bulletin is stamped with the moment it was due rather than the moment it finished rendering.
Every bulletin spends speech characters against the workspace, scheduled runs included.
Broadcast#Incident ranking, with receipts
A metro feed returns thirty to a hundred live incidents and a bulletin has room for five. Name the roads your audience actually drives and they are lifted enough that a major hold-up on one leads over a severe one nobody in the market drives, and every published bulletin keeps the incident list and the scores it was ordered by.
Roads are matched on the words you type, so an arterial has to be entered the way your audience says it.
Broadcast#Five ways out to playout
A finished bulletin can be pulled from a stable URL, announced as a signed webhook, or pushed over FTP, SFTP or Dropbox, each target with its own filename pattern — one fixed name to overwrite for a watch folder, a timestamped one to build an archive. Every configured target gets a row on the bulletin, so a bulletin that aired and went nowhere is visible rather than silent.
One bulletin is pushed at most three times per target — the recovery for a missed 07:20 is the 07:30, not a retry hours later.
Broadcast#Remote MCP server
Point Claude, or any other MCP client, at your workspace with an API key and it can list voices, render speech, convert a recording, transcribe a file and manage clones. The tools are the published REST endpoints described from the same catalogue as the documentation, and the list a client sees is filtered to the scopes the key carries, so a read-only key reads as read-only in the host's own tool picker.
The protocol has no confirmation step, so a destructive tool runs when it is called — withhold the scope rather than relying on a warning.
API & agent#Conversational agent
Describe what you want and the agent drives the same API you could drive yourself, with your permissions and nobody else's. Reads and renders happen on their own; deleting, publishing or anything that changes how a voice sounds becomes a card someone has to confirm at their own role, and every tool call is recorded for the workspace to read, including the ones that were refused.
Eight tool steps a turn, and a conversation may spend 20,000 characters or whatever the workspace has left, whichever is smaller.
API & agent#Cleanup stopped refusing files
Audio Cleanup rejected every upload and told people to re-export the file, including files this product had just produced. The fault was in how it measured a recording's level rather than in anyone's audio, and when something genuinely does fail the message now names what failed instead of guessing.
Fixes#
4 updates
Professional voice cloning
Turn an instant clone into a fine-tuned one: the platform books a GPU, trains an adapter on that voice's own finished renders and attaches it when it lands. The panel shows how many of your renders qualify before you commit — it needs at least twenty, each still paired with the script that produced it — and you can close the tab, because the run carries on without you.
Creator plan and up, three runs a billing period, one at a time. Training learns from render history rather than from the clips you uploaded, and a run that fails leaves the voice serving its instant clone.
Voices#Broadcast WAV delivery
Every finished spot comes out as a 44.1 kHz stereo PCM WAV you can hand straight to a traffic department, with an optional delivery loudness of as-mixed, -14, -16 or -19 LUFS. The loudness stage is one measured gain, so it cannot ride the voice and the bed around the way an automatic pass does.
The loudness trim only ever works downward; the master limiter already ships the mix hot.
Broadcast#Retry-safe renders
Send an Idempotency-Key with a render and a retry stops being a second charge: the same key with the same payload hands back the first call's result and marks the response a replay. It covers speech, voice conversion, transcription and starting a dub, and a request that failed releases its key at once, because a failed render costs nothing and has to stay retryable.
Keys are remembered for 24 hours, and changing the payload under the same key is a different request that renders and bills.
API & agent#Long recordings transcribe whole
Anything past roughly four and a half minutes came back cut off mid-sentence, after the whole file had been processed and charged for in full. Transcripts now return complete with per-word timings, which fixes it in every place it was wrong: Speech to Text, the reading stage of a dub, "From an existing ad" in the Script Writer, and the public transcription endpoint.
Word timings add a little decode time on a long file.
Fixes#
18 updates
Text to speech studio
HighlightPaste a script, pick a voice, and shape the read with four dials — Stability, Similarity, Style and Speed. Leave all four alone and you get the house sound the engine was tuned for; pin a seed and the same script comes back as exactly the same performance, take after take.
10,000 characters per render, so a long script has to be split.
Studio#Performance pass
Turns a written script into a spoken one — breath where a person breathes, a held beat before the payoff, spoken word order in place of written — at three depths from Subtle to Characterful. The prepared read is previewed with its marker count and character difference, and one button puts your original back.
Silence and punctuation only: this engine performs no laugh or sigh tags, so any that get invented are stripped and listed rather than read out on air.
Studio#Workspace pronunciation dictionary
Teach the workspace how to say the names it keeps getting wrong — sponsors, suburbs, hosts, call signs — and every render afterwards says them right. Matching is whole-word, so an entry for Vic never turns Victoria into Vickstoria, and you can hear a respelling before you save it: the same line rendered twice in one voice, as written and as respelled, pinned to the same seed.
Respellings rather than phonetic symbols — you spell it the way it sounds, Decatur to Deekayder — and a render applies the 500 most recently added words.
Studio#Voice changer
Record or upload a performance and hear it back in another voice. Your timing, phrasing, emphasis, pauses and breaths all survive and only the timbre changes, and the result sits beside the source so you can compare the two before keeping it.
One speaker only — overlapping voices come out as one, and room reverb or music under the read is re-voiced along with the words. Conversion bills roughly 1,000 characters per minute of source audio.
Studio#Multi-voice from a script
Paste a script and the speakers are picked out of it — NAME:, [NAME], (NAME) and standalone screenplay cues all parse — then give each character a voice and render the lot as one file, gaps and all. Fix one line of a twenty-line script and only that line is re-rendered: the rest are reused and the file is re-stitched around it.
A colon label is only read as a character when the rest of the document supports it, so a lone "Note:" in prose stays in the narration.
Studio#Replace a word, keep the voice
Retype a word in the transcript and it is re-synthesised in the clip's own voice and spliced into the original recording. It is rendered inside the words that surround it, so it arrives mid-phrase rather than as its own small sentence with a full stop on the end, then level-matched to its neighbours so the fixed word does not jump.
Those surrounding words are billed too, so the button states the exact character count before you press it.
Studio#Licence-aware effects library
Search freesound.org's Creative Commons recordings from inside the console, filter by licence, length or rating, and drop what you keep under, before or after any finished render. Every result carries its licence and a copy-ready credit line, so a non-commercial clap is flagged as unusable in a paid spot before you fall in love with it.
What you keep is the site's mp3 preview rather than the master file, copied into your workspace so it survives the original being taken down.
Studio#Instant voice cloning
HighlightDrop in a recording and the voice is ready to use straight away — there is no training step to wait for, and the upload returns as soon as the audio lands rather than holding open while a preview renders. Add up to ten clips at once, and every cloned voice gets a preview of the same fixed sentence, so the whole library is comparable like for like.
The engine clones from one clip at render time, so it wants a clean single-speaker take with nothing underneath it. 50 MB a file.
Voices#Two hundred premade voices
The library opens on the platform's own catalogue rather than an empty shelf: two hundred voices on a deliberate grid — sixteen accents, two genders, three age bands, with the delivery rotated so no two voices in the same bucket read alike — plus eight one-offs the grid would never produce, among them a late-night blues host and a hoarse sports commentator.
Premade voices are read only: usable in every studio, but you cannot edit their samples, relabel them or fine-tune them.
Voices#The broadcast voice chain
The Mixer runs the read through the chain a station would — rumble off below 75 Hz, chest lifted at 110, boxiness cut at 400, presence at 6 kHz, then a compressor and a limiter before the bus — and lands it in a slot from :10 to :120. An over-running read is time-compressed to fit exactly, or the spot extends and the bed runs under the whole read; nothing is cut either way.
The EQ curve is fixed, and time compression only speeds an over-running read up — a short read is placed inside the slot, never slowed to fill it.
Broadcast#Long-form projects
HighlightPaste or upload a manuscript and it arrives as a project: headings and horizontal rules become chapters, paragraphs become blocks, and a paragraph too long for one render is divided at a sentence boundary rather than cut short. Every block can carry its own voice, and one button joins every rendered block into a single MP3 in reading order.
Import reads .txt and .md, and a project past 500 MB has to go out a chapter at a time.
Long-form#Copy cut to the slot
The Script Writer writes to a slot rather than to a vague length: a :30 carries a 72–84 word budget, the model is held to it, and what comes back is measured again rather than taken on trust. The meter re-counts as you edit, and pause markers are lifted out of the word count and added back as seconds.
A broadcast read lands between 2.4 and 2.8 words a second, so the figure is a band and says "about" — render the take for the real duration.
Long-form#Dubbing, line by line
Move a recording into another language: the audio is read, split into lines, each line translated and re-read in the voice you picked, then joined into one master. Every line is listed with its source text, its translation, where it sits in the original and its own audio, and any single line can be re-rendered on its own.
One voice reads the whole thing, the source has to be English, and the master is a sequential read — lines are not stretched to land on the original's timings.
Long-form#Speech to Text
Drop in a recording — audio or video, up to 50 MB — and get text you can read, search and paste. Video is demuxed and only the audio track is read, and the upload is kept beside its transcript so you can check one against the other.
One continuous block with no speaker labels, and recognition is English only — other languages come back as English-shaped nonsense rather than as an error.
Long-form#Public REST API
HighlightSpeech, voice conversion, transcription, voices and render history over a keyed REST API running the same engine the studio uses, with audio returned as bytes on the response rather than a job to poll; payloads are shaped so client code written against ElevenLabs ports with a change of base URL. Issue a key with only the scopes it needs, and the API Keys page lists the last forty calls with their status, how long each took and the characters each one billed.
Text is capped at 10,000 characters a request and uploaded audio at 50 MB, and scopes are fixed when a key is issued — widening one means issuing a new key.
API & agent#Signed webhook deliveries
Point a URL at your workspace and we post to it when a render finishes or fails, a transcript lands, a dub completes or a project finishes. Every delivery carries an HMAC-SHA256 signature over the timestamp and the body together, so a receiver can prove the message came from us and reject one captured and replayed days later.
A receiver has eight seconds to answer.
API & agent#The metering ledger
Every piece of work writes one row you can read back: which surface it came from, the voice, the seconds of audio, the engine time and the characters it cost. The rates sit above the table — a character of script costs a character, a minute of voice-changer source costs a thousand, transcription is a tenth of the synthesis rate — and a failed render is never billed.
Fixed 7, 30 and 90-day windows, and no export yet.
Platform#Plan changes and failed cards
Move up a plan mid-month and you are charged the difference only, scaled to the days left in the period, so you are never billed twice for the overlap. A declined renewal switches nothing off — the charge is retried after one day, then three, then five, and the workspace keeps its full allowance until a fourth attempt fails.
A move down is not refunded; the higher allowance simply runs to the end of the period.
Platform#