narrate
Reference for the narrate DVM, which does text-to-speech with per-character pricing, quality tiers, voice personas, dialogue, and captions.
narrate converts text into audio with per-character pricing. A single quality knob routes across TTS vendors: the caller picks the trade-off (speed, realism, expressiveness), and narrate picks the vendor. Outputs are content-hash cached and stored in blob storage when large.
Canonical endpoint: https://narrate.dvmkit.ai. Two capabilities: synthesise (the paid audio job) and list-voices (free catalog discovery).
Published reference. This page is the source of truth for narrate's documented inputs, outputs, prices, and errors. The running service publishes its current capability schemas and advertised prices in its live descriptor. Full output envelope shapes are in the dvm CLI reference.
Quality tiers
quality is required on every synthesise call. It determines the vendor, feature availability, and the per-character price.
| Tier | Vendor | Best for | USD/char | $/M chars |
|---|---|---|---|---|
fast | Cartesia Sonic | Notifications, chat replies, real-time agent voices | $0.0001 | $100 |
balanced | ElevenLabs Turbo | Long-form narration, articles, audiobooks | $0.00015 | $150 |
expressive | ElevenLabs v3 | Emotion, audio tags, native multi-speaker dialogue | $0.00025 | $250 |
Quality does not cap input length. Long-form works on every tier. The tier choice is about voice realism, feature availability, and cost, not input size.
Feature gating
Several features are honoured only on expressive; requesting them on fast/balanced rejects up-front with an invalid_input envelope before payment.
| Feature | fast | balanced | expressive |
|---|---|---|---|
| Text + dialogue synthesis | ✓ | ✓ | ✓ |
speaking_rate | ✓ (clamped 0.6–1.5) | ✓ | ✓ |
| Multilingual content | ✓ | ✓ | ✓ |
Markdown structure (input.format) | ✓ | ✓ | ✓ |
output.normalize / output.captions / output.chapters | ✓ | ✓ | ✓ |
style parameter | rejected | rejected | ✓ |
Inline audio tags ([whispering]) | rejected | rejected | ✓ |
| Native multi-speaker dialogue API | per-turn stitch | per-turn stitch | ✓ (≤ 2k chars) |
Raw Cartesia voice IDs are also rejected on expressive: that tier's price covers EL v3 features Cartesia can't deliver.
synthesise
Input schema
{
"quality": "balanced",
"input": {
"text": "Hello, world.",
"format": "plain",
"language": "en"
},
"voice": "narrator",
"speaking_rate": 1.0,
"output": { "format": "mp3", "normalize": false, "captions": false, "chapters": false }
}
input: exactly one oftext(monologue) ordialogue(array of{ voice?, text, style?, language? }turns; up to 100 turns, ≤ 100,000 chars total).formatdefaults to"plain";"markdown"adds structure-aware prosodic pauses.languageis auto-detected when omitted.voice: persona name (narrator,host,formal,casual,audiobook), a raw ElevenLabs 20-char ID, or a Cartesia UUID. Omitted falls back to the operator default, else thenarratorpersona.style: expressive-only free-form performance direction (≤ 500 chars).output:format(mp3/wav/ogg/opus),normalize(EBU R128),captions(VTT or{ format: "srt" }),chapters(ID3v2 CHAP frames; mp3 only).
Voice personas
Each persona is a stable voice character; the underlying vendor voice varies per tier.
| Persona | Character | fast (Cartesia) | balanced / expressive (EL) |
|---|---|---|---|
narrator | Warm narrative female, mid-range | Caroline | Dorothy |
host | Energetic casual male | Blake | Patrick |
formal | Composed authoritative female | Gemma | Rachel |
casual | Warm expressive female | Katie | Jessica |
audiobook | Deep weighty male, long-form | Ronald | Brian |
CLI example
dvm request -d dvmkit--narrate/synthesise --data '{"quality":"balanced","input":{"text":"Hello, world."},"voice":"narrator"}'
dvm request -d dvmkit--narrate/synthesise --data '{"quality":"expressive","input":{"text":"Hello [excited] friend!"},"style":"warm and reassuring"}'
dvm quote -d dvmkit--narrate/synthesise --data '{"quality":"fast","input":{"text":"Ping."}}'
list-voices
Free. Returns the persona catalog with per-tier voice identity and sample URLs. Optional filters: language (ISO 639-1), quality, query (free-text).
dvm request -d dvmkit--narrate/list-voices --data '{"language":"fr","quality":"expressive"}'
Each entry carries persona, character, a quality_tiers map ({ provider, voice_id, model, label, sample_url } per tier — provider here names the upstream TTS vendor behind that tier, Cartesia or ElevenLabs, not narrate itself), and language_coverage. The per-tier sample_url lets a caller audition the exact voice they'd get on that tier.
Auth
Signed requests are required. Every quote and job carries a secp256k1 + BIP-340 Schnorr signature over the canonical-JSON body, at the descriptor level so all capabilities share one auth slot. Drift window ±5 minutes; replay window 10 minutes. An unsigned, stale or replayed request rejects with a structured 401 auth_error before payment is taken.
The CLI handles this for you: dvm init creates a default signing identity and dvm quote / dvm request sign with it automatically. Pass --as <identity> only when you want to sign as a different one. If you see auth_error, run dvm identity list to check you have one.
Prepaid credit
narrate offers a prepaid balance, so an agent making repeated calls funds once instead of paying per job. The funding menu rides on every quote and every 402.
| Field | Value |
|---|---|
| Minimum funding | $0.10 |
| Maximum residual balance | $5.00 |
| Credit lifetime | 30 days from the most recent funding |
dvm credit fund dvmkit--narrate --amount 2.00
dvm credit balance dvmkit--narrate
dvm credit drain dvmkit--narrate # reclaim it, works after expiry too
A job priced above the maximum still clears: the ceiling binds how much balance you may hold, not what you may spend. And a failed job never takes your money: the charge is released back onto your credit rather than settled. See Prepaid credit for the full model.
Pricing
Per-character, per-tier: see the Quality tiers table. Cost = character_count × tier_rate, converted to sats at the 402 handshake. Content-hash caching (identical inputs from any caller) means cache hits return the same audio at no charge; feature-gate rejections happen before payment, so a bad-tier request never bills.
Error codes
Failures surface via the CLI error envelope with a stable code, kept in lockstep with narrate's error taxonomy. No failure charges you, whatever the code says. A failed job releases its draw back onto your credit. The attribution column says who to chase, never what the failure cost: Caller means change the request and retry; anything else means the fault is ours and the retry needs time or an operator. The provider_* codes below name the upstream TTS vendor (Cartesia, ElevenLabs), not narrate itself; "provider" elsewhere in these docs means narrate.
| Code | Trigger | Attribution |
|---|---|---|
invalid_input | Validation or tier gating: a style param, inline audio tags, or per-turn style on a non-expressive tier; chapters without mp3 + markdown; a raw Cartesia UUID at expressive. Rejected up-front, before payment, with a hint pointing at the fix | Caller |
content_rejected | The vendor's moderation refused the caller's text (a recognized content-policy signature) | Caller |
provider_rejected | The vendor rejected our request shape (an unrecognized 4xx carrying no content-policy signature), so never blamed on the caller's content | Builder/operator |
provider_allowance_exhausted | The operator's own vendor plan is out of credit, whatever status the vendor bills it with. Held distinct from auth_error: a dry plan needs a top-up, a dead key needs rotating | Operator |
rate_limited | Vendor 429. Transient, safe to retry | Operator |
provider_error | Vendor 5xx or network error. Transient, safe to retry | Operator |
auth_error | Either the request signature, clock, or nonce is invalid (rejected before payment; see Auth), or the operator's vendor key is dead or misconfigured | Caller or operator |
See also
- Explore the catalog · dvm CLI reference · Caller quickstart
- Other first-party DVMs: scrape · scribe · cast · discover