# Higgsfield Image / Video / Voice Model Selection v0.1 Status: active benchmark policy for CartoonOS / Luminautas. Date: 2026-09-08. ## Principle **Do not select one global best model. Select the best model for the exact production task.** CartoonOS benchmarks three independent modalities: - image - video - voice Each uses different units, hard gates and economics. ## 1. Image model track Current live Higgsfield image catalog includes, among others: ### Google / Gemini Nano Banana family - `nano_banana` — Nano Banana - `nano_banana_pro` — Nano Banana Pro - `nano_banana_2` — Nano Banana 2 - `nano_banana_2_lite` — Nano Banana 2 Lite Observed live preflight examples on 2026-09-08: - Nano Banana: ~1 credit/image - Nano Banana Pro 2K: ~2 credits/image - Nano Banana Pro 4K: ~4 credits/image - Nano Banana 2 2K: ~2 credits/image - Nano Banana 2 Lite 1K: ~1 credit/image Pricing must be stored as a dated snapshot because provider pricing can change. Other relevant image models include: - GPT Image 2 / 2.5 - Seedream 4.5 / 5.0 - FLUX.2 / Flux Kontext - Recraft V4.1 - Cinema Studio Image - Soul / Soul Cinema - Kling O1 Image - Grok Image ### Image benchmark archetypes - character sheet - turnaround - expression pack - transparent cutout - background plate - storyboard frame - educational diagram - start/end video frame ### Image KPIs Quality: - visual quality - exact character identity consistency - anatomy consistency - silhouette consistency - palette/material consistency - expression readability - prompt adherence - text/diagram accuracy when required - clean-background/transparency quality Efficiency: - credits per requested image - first-pass acceptance - credits per accepted image - generation latency p50/p95 - QA/rework time - rejected-image rate Character assets have a hard identity floor >=95 before economics may influence selection. ## 2. Video model track Relevant live models include Seedance, Kling, Cinema Studio, MiniMax, Wan, Gemini Omni, Veo, FLUX Video and other current catalog entries. ### Video archetypes - character close-up - character dialogue - multi-character - cinematic establishing - short B-roll - Clarity Island - motion transfer ### Video KPIs - visual quality - character identity - anatomy - prompt adherence - motion/physics - temporal continuity - audio/lipsync where applicable - first-pass acceptance - provider success - retries/rework - credits/output second - credits/accepted second - accepted seconds/hour - end-to-end latency **Primary economics metric:** `credits_per_accepted_second`, not raw credits per generation. ## 3. Voice model track The canonical character voice is distinct from the TTS engine. Benchmark engines only against an already-approved mascot voice identity. Current live TTS engine families include: - Seed Audio 1.0 - Qwen Audio 3.0 TTS Flash - Text to Speech V2 variants: ElevenLabs, MiniMax, Seed Speech, Vibe Voice, Cozy Voice - other compatible engines from the live catalog ### Voice KPIs - canonical voice identity similarity - pronunciation accuracy - ES-419 fit - age/personality fit - emotional readability - naturalness - pacing consistency - pitch consistency - audio artifact rate - lipsync compatibility - generation latency - cost per accepted dialogue minute - first-pass acceptance Hard gates for canonical mascot voices: - identity similarity >=95 - canonical lexicon pronunciation >=98 - emotional readability >=90 - lipsync QA pass where required ## 4. Canonical voices Every mascot has one active versioned voice identity: - Lumi Voice v1 - Tiko Voice v1 - Nova Voice v1 - Marina Voice v1 - Rin Voice v1 - Arqueo Voice v1 - Pepe Voice v1 Exact provider `voice_type + voice_id` values remain unbound until audition and explicit approval. Never invent them. Shorts and flagship episodes must reuse the same active mascot voice binding. A provider/engine migration requires matched-sample regression testing and explicit version approval. ## 5. Model selector logic For each generation request: 1. classify modality; 2. classify archetype; 3. determine exact character versions; 4. apply hard quality floors; 5. filter models by live capability snapshot; 6. use matched benchmark cohort only; 7. require enough sample confidence; 8. estimate expected total accepted cost, not raw request cost; 9. account for latency and reliability; 10. select best eligible configuration; 11. retain alternatives for fallback; 12. store decision + reason + benchmark snapshot version. ## 6. Confidence policy - `<5` matched samples: `insufficient_sample` - `5–19`: `provisional` - `>=20`: `decision_ready` No low-sample model is declared a permanent winner. ## 7. Expected-total-cost formula A useful production estimate is: `expected_total_cost = request_cost / first_pass_acceptance_rate + expected_rework_cost + QA_cost` For video, normalize further to accepted seconds/minutes. For voice, normalize to accepted dialogue minutes. For image, normalize to accepted production assets. ## 8. Nano Banana evaluation plan Benchmark all four active Nano Banana variants against the same controlled asset briefs: 1. Lumi six-view turnaround 2. Tiko six-view turnaround 3. Lumi expression pack 4. Tiko expression pack 5. clean 16:9 background 6. Curiosity Graph educational diagram 7. 9:16 Short hook frame 8. start/end keyframe pair Evaluate exact identity, anatomy, consistency across panels, text/diagram quality, cost and latency. Do not crown Nano Banana based on price alone. ## 9. Production recommendation policy The dashboard should output recommendations like: - `best_for_character_turnaround` - `best_for_expression_pack` - `best_for_clean_background` - `best_for_diagram` - `best_for_character_dialogue_video` - `best_for_multi_character_video` - `best_for_fast_short_broll` - `best_for_lumi_voice` - `best_for_tiko_voice` Avoid a single `best_model` field. ## 10. North-star production economics Track at system level: - cost per accepted second - cost per accepted dialogue minute - cost per approved image asset - finished minutes per production hour - first-pass acceptance - late rework rate - total production cost per finished Short - total production cost per finished flagship minute - quality-adjusted production cost The goal is **high-quality accepted output per dollar and production hour**, not lowest nominal API price.