# Luminautas Voice Canon & QA v0.1 Status: active voice-production policy. Project: `PROJ-LUMINAUTAS-v1` ## Core rule **One canonical voice identity per mascot.** The mascot voice is a versioned production asset, not an incidental TTS setting. For every speaking mascot persist: - `character_id` - `voice_identity_version` - provider / engine - exact `voice_type` - exact `voice_id` - language / locale - speech-rate baseline - pitch baseline - loudness baseline - pronunciation lexicon version - emotion profile version - sample IDs / reference clips - lipsync QA version - approval state Never invent voice IDs. Never silently swap a mascot to a different preset or engine. ## Voice identity vs engine Voice identity is the character. The engine is only the synthesis implementation. Example: `Lumi Voice v1` may initially run through one engine, but any future engine migration must reproduce Lumi Voice v1 within the defined similarity and QA tolerance before it can replace the active binding. Changing provider without re-validation is a voice change. ## Active mascot voice directions ### Lumi — franchise protagonist Voice qualities: - bright - intelligent - adventurous - empathetic - naturally curious - confident without sounding adult - emotionally expressive Target pacing: approximately 135–155 wpm for normal dialogue, slower for wonder/Clarity Island moments. Avoid: - lecturer tone - exaggerated baby voice - adult presenter voice - constant excitement ### Tiko Voice qualities: - warm - playful - curious - slightly cautious - funny without becoming noisy Target pacing: approximately 125–145 wpm. Signature delivery works well with short `¿Y si…?` questions and nervous-curious reactions. ### Nova Voice qualities: - calm - precise - friendly - helpful - slightly synthetic only if it remains emotionally warm Target pacing: approximately 115–135 wpm. Avoid cold assistant/robot stereotypes. ### Marina Voice qualities: - calm - warm - wise - patient - grounded Use marine/ocean storytelling cadence. Avoid generic mystical-old-mentor exaggeration. ### Rin Voice qualities: - gentle - observant - caring - alert when ecosystems/animals need attention Target pacing: approximately 110–130 wpm. ### Arqueo Voice qualities: - curious investigator - careful with evidence - enthusiastic when a clue becomes meaningful - respectful, never treasure-hunter caricature ### Pepe Voice qualities: - fast - playful - rhythmic - clever - energetic but controlled Target pacing: approximately 145–165 wpm. Pepe must never destroy Clarity Island comprehension through excessive chatter. ## Canonical audition pack Before binding a mascot voice, generate the exact same audition script across candidate voices/engines. Audition must include: 1. neutral introduction 2. curiosity question 3. surprise 4. worried/uncertain line 5. thoughtful reasoning 6. determined line 7. gentle emotional moment 8. laughter / playful beat where relevant 9. character signature phrase 10. difficult Spanish-LATAM pronunciation set ## Spanish-LATAM pronunciation lexicon Create and version a shared ES-419 pronunciation dictionary plus mascot-specific overrides. Mandatory classes: - mascot names - Luminautas terms - scientific vocabulary - indigenous / Nahuatl-derived names - Xochimilco - axolotl / ajolote terminology - archaeology/history names - marine species - place names - foreign proper nouns - Curiosity Graph terms Every pronunciation correction becomes reusable structured knowledge rather than a one-off prompt edit. ## Voice QA dimensions Score each candidate / generation: - identity similarity - age fit - emotional readability - naturalness - pronunciation accuracy - ES-419 fit - pacing consistency - pitch consistency - loudness consistency - sentence-boundary naturalness - laughter / reaction quality - long-session fatigue risk - lipsync compatibility - audio artifact rate ## Hard gates Production voice approval requires: - identity similarity >=95 across approved samples - pronunciation QA >=98 for canonical test lexicon - emotional readability >=90 - no severe audio artifacts - child-safe delivery - correct locale - lipsync test pass for speaking mascots ## Episode/Short voice regression Before release compare generated dialogue against the active canonical voice profile. Fail/review when: - wrong `voice_id` / `voice_type` - unexpected engine change - large pitch/rate drift - identity similarity below threshold - pronunciation regression - emotional state contradicts scene contract ## Engine benchmark policy Evaluate TTS engines separately from voice identity. Current live Higgsfield audio engines include: - Seed Audio 1.0 - Qwen Audio 3.0 TTS Flash - Text to Speech V2 with ElevenLabs / MiniMax / Seed Speech / Vibe Voice / Cozy Voice variants - other available engines where language/voice requirements fit Benchmark per canonical voice using: - identity preservation - pronunciation - emotion controllability - latency - cost per finished dialogue minute - retry / failure rate - lipsync compatibility Do not declare one global TTS winner. The chosen engine must preserve the specific mascot voice identity. ## Voice migration A new engine or voice version can replace the active binding only after: 1. matched audition set 2. blind similarity QA 3. pronunciation regression suite 4. emotion suite 5. lipsync suite 6. cost/latency comparison 7. explicit approval 8. manifest version update Historical episodes retain their exact original bindings for provenance. ## Analytics Track per mascot and voice version: - dialogue seconds generated - generation latency - credits/cost - retries - accepted dialogue seconds - cost per accepted dialogue minute - pronunciation defects - lipsync defects - identity-drift defects - human QA time - audience retention around dialogue beats when available ## Production principle **The audience should recognize a mascot by voice before seeing the screen.**