CartoonOS compares matched production tasks by quality, identity consistency, first-pass acceptance, latency, rework and accepted-output cost. There is no single global best model.
Character sheets, turnarounds, expressions, clean backgrounds, diagrams, storyboards and keyframes.
Dialogue, cinematic shots, multi-character scenes, B-roll, Clarity Islands and motion transfer.
Canonical mascot voice identity is measured separately from the underlying TTS engine.
| Modality | Provider | Model | Route | Unit | Samples | Decision state |
|---|---|---|---|---|---|---|
| {model.modality} | {model.providerName} | {model.modelName} {model.modelId} |
{model.route} | {model.benchmarkUnit.replaceAll("_", " ")} | {samples.length} | {decisionState(samples.length).replaceAll("_", " ")} |
Nano Banana, Pro, 2 and 2 Lite are evaluated against the same character/reference briefs. Price alone never overrides identity or anatomy gates.
Lumi, Tiko, Nova, Marina, Rin, Arqueo and Pepe each keep one active versioned voice identity across Shorts/Reels and flagships.
<5 matched samples = insufficient; 5–19 = provisional; ≥{benchmarkDecisionPolicy.decisionReadyMin} = decision-ready.
Measures how often a model clears the exact quality gate without regeneration or repair.
Character assets and dialogue must preserve the approved identity before economics may influence routing.
Queue, generation, QA and total wall-clock time are tracked separately to expose scale bottlenecks.
Rank by matched archetype, exact character/version, quality floor, expected accepted cost, reliability and latency.
A cheap background model can be a poor Lumi close-up model; a strong image model may not be the best storyboard or diagram model.
Production-model optimization ultimately serves durable audience value, not API-price minimization.