Pika Speech

Powered by Pika

For narration and character reads, Pika Speech is the fastest, most cost-efficient text-to-speech model on the market. Use it to turn any script into a spoken performance, read by one of over 50 voice presets.

Everything Pika Speech can do

76 voices, named for the read

Narrators, anchors, hosts, coaches and characters. The names describe delivery rather than a person — Smoky Noir Monotone, Bright Peppy Vlogger, Booming Ring Announcer — so you pick by the sound you are after.

Whole scripts in one pass

15,000 characters reach the model as a single request, so a full narration is one job rather than a stack of lines to render and stitch.

48 kHz out

Full-band audio at the rate a video timeline already runs, so the read lands next to production dialogue without a resample on the way in.

Fast enough to keep rewriting

A minute of speech generates in about a second. Change a word, re-run it, and compare takes in the time a slower model spends on the first one.

Examples

Frequently Asked Questions

How long can the script be?

15,000 characters in one request. Past that the panel holds Generate and says how far over you are.

How many voices are there?

76 presets, each named for its delivery. The panel opens on Calm Documentary Narrator.

Can I clone my own voice?

Not from this panel — it reads your script in one of the 76 presets. Cloning from a recording of your own is not part of this composer.

What quality does it render at?

48 kHz audio, generated at a real-time factor of 0.02 — roughly a second of compute for a minute of speech.

Can it sing?

No. This one speaks a script; melody, lyrics and a sung performance come from Pika Music, one tab across on the Mode row.

Does the same voice come back every time?

Yes. A preset is a fixed voice, so one choice carries the same identity across every line and every project.