dvmkitdocs
DVM Reference

narrate

Reference for the narrate DVM, which does text-to-speech with per-character pricing, quality tiers, voice personas, dialogue, and captions.

narrate converts text into audio with per-character pricing. A single quality knob routes across TTS vendors: the caller picks the trade-off (speed, realism, expressiveness), and narrate picks the vendor. Outputs are content-hash cached and stored in blob storage when large.

Canonical endpoint: https://narrate.dvmkit.ai. Two capabilities: synthesise (the paid audio job) and list-voices (free catalog discovery).

Published reference. This page is the source of truth for narrate's documented inputs, outputs, prices, and errors. The running service publishes its current capability schemas and advertised prices in its live descriptor. Full output envelope shapes are in the dvm CLI reference.

Quality tiers

quality is required on every synthesise call. It determines the vendor, feature availability, and the per-character price.

TierVendorBest forUSD/char$/M chars
fastCartesia SonicNotifications, chat replies, real-time agent voices$0.0001$100
balancedElevenLabs TurboLong-form narration, articles, audiobooks$0.00015$150
expressiveElevenLabs v3Emotion, audio tags, native multi-speaker dialogue$0.00025$250

Quality does not cap input length. Long-form works on every tier. The tier choice is about voice realism, feature availability, and cost, not input size.

Feature gating

Several features are honoured only on expressive; requesting them on fast/balanced rejects up-front with an invalid_input envelope before payment.

Featurefastbalancedexpressive
Text + dialogue synthesis✓✓✓
speaking_rate✓ (clamped 0.6–1.5)✓✓
Multilingual content✓✓✓
Markdown structure (input.format)✓✓✓
output.normalize / output.captions / output.chapters✓✓✓
style parameterrejectedrejected✓
Inline audio tags ([whispering])rejectedrejected✓
Native multi-speaker dialogue APIper-turn stitchper-turn stitch✓ (≤ 2k chars)

Raw Cartesia voice IDs are also rejected on expressive: that tier's price covers EL v3 features Cartesia can't deliver.

synthesise

Input schema

{
  "quality": "balanced",
  "input": {
    "text": "Hello, world.",
    "format": "plain",
    "language": "en"
  },
  "voice": "narrator",
  "speaking_rate": 1.0,
  "output": { "format": "mp3", "normalize": false, "captions": false, "chapters": false }
}
  • input: exactly one of text (monologue) or dialogue (array of { voice?, text, style?, language? } turns; up to 100 turns, ≤ 100,000 chars total). format defaults to "plain"; "markdown" adds structure-aware prosodic pauses. language is auto-detected when omitted.
  • voice: persona name (narrator, host, formal, casual, audiobook), a raw ElevenLabs 20-char ID, or a Cartesia UUID. Omitted falls back to the operator default, else the narrator persona.
  • style: expressive-only free-form performance direction (≤ 500 chars).
  • output: format (mp3 / wav / ogg / opus), normalize (EBU R128), captions (VTT or { format: "srt" }), chapters (ID3v2 CHAP frames; mp3 only).

Voice personas

Each persona is a stable voice character; the underlying vendor voice varies per tier.

PersonaCharacterfast (Cartesia)balanced / expressive (EL)
narratorWarm narrative female, mid-rangeCarolineDorothy
hostEnergetic casual maleBlakePatrick
formalComposed authoritative femaleGemmaRachel
casualWarm expressive femaleKatieJessica
audiobookDeep weighty male, long-formRonaldBrian

CLI example

dvm request -d dvmkit--narrate/synthesise --data '{"quality":"balanced","input":{"text":"Hello, world."},"voice":"narrator"}'
dvm request -d dvmkit--narrate/synthesise --data '{"quality":"expressive","input":{"text":"Hello [excited] friend!"},"style":"warm and reassuring"}'
dvm quote   -d dvmkit--narrate/synthesise --data '{"quality":"fast","input":{"text":"Ping."}}'

list-voices

Free. Returns the persona catalog with per-tier voice identity and sample URLs. Optional filters: language (ISO 639-1), quality, query (free-text).

dvm request -d dvmkit--narrate/list-voices --data '{"language":"fr","quality":"expressive"}'

Each entry carries persona, character, a quality_tiers map ({ provider, voice_id, model, label, sample_url } per tier — provider here names the upstream TTS vendor behind that tier, Cartesia or ElevenLabs, not narrate itself), and language_coverage. The per-tier sample_url lets a caller audition the exact voice they'd get on that tier.

Auth

Signed requests are required. Every quote and job carries a secp256k1 + BIP-340 Schnorr signature over the canonical-JSON body, at the descriptor level so all capabilities share one auth slot. Drift window ±5 minutes; replay window 10 minutes. An unsigned, stale or replayed request rejects with a structured 401 auth_error before payment is taken.

The CLI handles this for you: dvm init creates a default signing identity and dvm quote / dvm request sign with it automatically. Pass --as <identity> only when you want to sign as a different one. If you see auth_error, run dvm identity list to check you have one.

Prepaid credit

narrate offers a prepaid balance, so an agent making repeated calls funds once instead of paying per job. The funding menu rides on every quote and every 402.

FieldValue
Minimum funding$0.10
Maximum residual balance$5.00
Credit lifetime30 days from the most recent funding
dvm credit fund dvmkit--narrate --amount 2.00
dvm credit balance dvmkit--narrate
dvm credit drain dvmkit--narrate     # reclaim it, works after expiry too

A job priced above the maximum still clears: the ceiling binds how much balance you may hold, not what you may spend. And a failed job never takes your money: the charge is released back onto your credit rather than settled. See Prepaid credit for the full model.

Pricing

Per-character, per-tier: see the Quality tiers table. Cost = character_count × tier_rate, converted to sats at the 402 handshake. Content-hash caching (identical inputs from any caller) means cache hits return the same audio at no charge; feature-gate rejections happen before payment, so a bad-tier request never bills.

Error codes

Failures surface via the CLI error envelope with a stable code, kept in lockstep with narrate's error taxonomy. No failure charges you, whatever the code says. A failed job releases its draw back onto your credit. The attribution column says who to chase, never what the failure cost: Caller means change the request and retry; anything else means the fault is ours and the retry needs time or an operator. The provider_* codes below name the upstream TTS vendor (Cartesia, ElevenLabs), not narrate itself; "provider" elsewhere in these docs means narrate.

CodeTriggerAttribution
invalid_inputValidation or tier gating: a style param, inline audio tags, or per-turn style on a non-expressive tier; chapters without mp3 + markdown; a raw Cartesia UUID at expressive. Rejected up-front, before payment, with a hint pointing at the fixCaller
content_rejectedThe vendor's moderation refused the caller's text (a recognized content-policy signature)Caller
provider_rejectedThe vendor rejected our request shape (an unrecognized 4xx carrying no content-policy signature), so never blamed on the caller's contentBuilder/operator
provider_allowance_exhaustedThe operator's own vendor plan is out of credit, whatever status the vendor bills it with. Held distinct from auth_error: a dry plan needs a top-up, a dead key needs rotatingOperator
rate_limitedVendor 429. Transient, safe to retryOperator
provider_errorVendor 5xx or network error. Transient, safe to retryOperator
auth_errorEither the request signature, clock, or nonce is invalid (rejected before payment; see Auth), or the operator's vendor key is dead or misconfiguredCaller or operator

See also

On this page