A September launch, checked in October
ElevenLabs published its v4 and v4 Turbo announcement on 28 September and updated it on 5 October. It describes more expressive text-to-speech, improved direction controls and a low-latency Turbo variant, available through its creative, agents and API products. This article is new coverage on 10 October, not a claim that the models launched today.
Read the timing footnote with the headline
The same official announcement reports approximately 100 ms median inference latency and approximately 150 ms median time from request to audible speech. Its benchmark footnote says network latency was measured and removed, using identical scripts and default settings; Turbo used WebSocket streaming. These are vendor-reported measurements, not Fewertools test results.
A reader’s connection and application pipeline are therefore outside the quoted network-adjusted number. It should not be relabelled as the total time from a person finishing a question to hearing a useful answer.
Give each timestamp a name
For a proposed voice-agent trial, draw a timeline before comparing products. Record when the user finishes speaking, when the application submits the text-to-speech request, when the first audio data arrives and when speech becomes audible. Record completion separately too. These timestamps answer different questions; combining them into a single unexplained “latency” column hides the source of delay.
Use a short synthetic script with a number, a proper noun and a sentence requiring a pause. Keep the voice selection, language, input text and output format fixed for the comparison. Write down the connection type and application build so the run can be repeated. If a model needs different direction instructions, retain those instructions with the output.
After any authorized native trial, preserve the audio and timing record, count retries, and log pronunciation corrections and billable usage. Repeat the same case enough times to report the observed spread rather than selecting the fastest clip. State the number of attempts and the conditions alongside any median.
Expression needs a separate result
A timing chart cannot answer whether a line sounds right. A proposed listening worksheet would ask whether the number was spoken correctly, whether the name remained consistent, whether the pause changed the intended meaning and whether a listener could understand the answer at the chosen pace. Keep those observations separate from the clock measurements.
We did not generate audio, clone a voice, run a listening panel or use paid provider credits for this report. It establishes the announcement and the benchmark’s stated limits. The ElevenLabs profile and test register distinguish the broader product record from published executions.
Sources and scope
Official sources read 10 October 2026. Facts above are attributed to the provider; the evaluation worksheets are proposed checks. This article does not report a native product test or assign a tool score.