# CartoonOS KPI and Measurement Contract Owner: Analytics & Learning capabilities. Reviewed: 2026-09-08. This document defines measurement semantics; it does not claim live data is connected. Runtime status: [`../architecture/CURRENT_RUNTIME_STATUS_2026-09-08.md`](../architecture/CURRENT_RUNTIME_STATUS_2026-09-08.md). ## Executive decision metrics Use three top operating measures: 1. **90-day quality-gated watch hours per economic production dollar** 2. **production cycle time** 3. **first-pass shot acceptance** Quality-gated means the released asset passed required production/release gates. It is not a measured learning score. Safety/factual/IP defects, project-isolation violations and budget breaches remain hard guardrails and are never compensated by watch time. | Metric / unit | Formula and grain | Source / cadence | Decision / limitation | |---|---|---|---| | D90 quality-gated watch hours / USD | eligible first-90-day minutes / 60 divided by matching full economic production cost | platform analytics + reviewed cost ledger; mature cohorts | primary sustainable production value; watch time does not prove education | | Production cycle time / hours | final approval timestamp − brief-ready timestamp per asset; report median/p95/n | workflow evidence; weekly | queue/review bottlenecks; separate active work/waiting | | First-pass shot acceptance / % | shots whose first completed generation attempt passes QA / shots with a reviewed first attempt | generation attempt evidence; weekly | avoidable rework/model routing | | Attempts / accepted shot | all attempts for shots with usable accepted output / accepted shots | production artifact + Observation ledger | lower is generally better subject to quality floor | | Accepted / generated seconds | unique accepted usable seconds / all generated seconds | generation evidence; weekly | yield/efficiency; do not double-count multiple accepted takes | | Cost / accepted second | all attributable attempt costs / unique accepted usable seconds | shot economics; weekly by route/shot family | null when accepted seconds = 0; never present fabricated 0/infinity | | Generation QA acceptance / % | accepted generation attempts / reviewed attempts | generation evidence | not editorial selection rate | | Rough-cut selection / % | selected attempt IDs / usable accepted candidate attempts | RoughCut provenance | editorial efficiency; separate from QA acceptance | | Failure-class rate / % | failed attempts in one canonical failure class / reviewed attempts | generation evidence | drives repair/routing policy | | p50/p95 generation latency | provider execution latency distribution | provider attempt evidence | provider/route capacity and user-facing cycle time | | Asset reuse ratio / % | reused/edit-extended required assets / eligible required assets | AssetReusePlan + actual asset lineage | continuity/cost leverage | | Avoided generation cost / USD | estimated comparable new-generation cost − actual selected reuse/edit cost | AssetReusePlan + actual cost | planning estimate until reconciled with actuals | | Average view percentage / % | provider-reported compatible measure | platform checkpoints | pacing diagnosis; may exceed 100 for loops; preserve provider semantics | | 30-second retention / proportion | supported retention point nearest 30s, preserving coordinates/video length | compatible report/export | not applicable for many Shorts; record uncertainty | | Impressions CTR / % | eligible click views / eligible counted impressions | supported report/export | packaging only; does not describe every traffic source | | Engaged views / count | preserve provider definition/report version | supported platform analytics | do not conflate with starts/all-platform views | | D30 returning-viewer rate / % | returning viewers / eligible viewers at compatible cohort grain | supported platform analytics | franchise retention; cohort/sample caveats required | | Cost / retained viewer-minute | matching full production cost / retained viewer-minutes | economics + platform checkpoints | downstream production efficiency | | Contribution margin | attributed revenue − variable production/distribution costs | monetization/economics ledger | never infer revenue from views | | Cash burn / month | actual cash outflows in period | invoice/payment ledger | runway; separate from allocated content cost | | Queue age / hours | now − entered-stage time for each unfinished item | workflow evidence | bottleneck/WIP control | | Retrieval support / % | correctly supported factual assertions / reviewed assertions | evaluator holdout | research reliability; record abstentions | | Capability task yield / % | accepted/reviewed outputs / completed capability runs | telemetry + reviewer result | replaces “agent task yield”; never reward tool-call/message volume | | Capability ROI | attributable economic/value-created outcome / capability cost | capability telemetry + economics | compare workflows/capabilities, not personas | | D7 memory transfer score | explicit transfer-measure result normalized to contract | reviewed measurement method | engagement is not a substitute | No universal CTR/retention/model benchmark target is asserted without a matched baseline and sufficient sample. ## Production evidence contract Every generation attempt preserves: - project/episode/asset/shot identity; - compiled request ID/checksum + ShotSpec/compiler version; - provider/model/capability + provider job ID; - attempt number; - cost/runtime/output duration; - generation QA accepted/rejected; - failure class/detail + repair action; - output asset ID/checksum/URI for usable accepted output. Generation acceptance and rough-cut selected take are different facts. Canonical failure classes: `identity · continuity · camera · motion · audio · safety · quality · provider · capability · other` ## Truth states Every displayed metric is one of: `live · snapshot · planning · awaiting_data · not_applicable · stale · error` Null always carries a reason. Zero is only a measured zero. Fixture/planning values are never relabeled as live. Every durable metric includes project/organization/channel/asset identifiers as applicable, platform asset ID, metric ID/version, value/unit, checkpoint, measurement range, `data_through_at`, collection time, source/raw-response reference, schema version and revision. Corrections append/supersede revisions rather than silently overwriting evidence. ## Checkpoint timing Schedule: `D1 · D3 · D7 · D14 · D30 · D90` A due time triggers collection; it does not prove exact elapsed-window data is available. Preserve the platform timezone/report interval and provisional/incomplete state. Backfill when provider data matures. Do not sum overlapping cumulative checkpoint values. ## Attribution Character/cast/format/platform/model attribution is descriptive unless the experiment design supports causal inference. Record at minimum: - protagonist/cast mode/specialists; - exact mascot versions; - topic/content pillar; - hook/participation/rewatch/next-curiosity dimensions; - format/runtime/language/platform; - publication age/promotion; - shot family/provider/model/capability; - experiment ID where registered. Compare matched cohorts and display sample size/uncertainty. Do not declare a mascot/model winner from a tiny unmatched sample. ## Revenue Ads/Premium/sponsorship/licensing/merch/subscription have distinct source records. Never infer revenue from views. Provider credits are not automatically USD. Reconcile invoices/consumed credits/overhead/labor without double counting. ## Experiment record Required: - stable experiment ID; - observation/source; - predeclared hypothesis; - one intended change where possible; - primary metric + guardrails; - eligible matched cohort; - expected direction; - review checkpoint/date; - owner capability; - evidence/result/confounders; - learning + next decision. Stop immediately on hard safety/factual/IP failure. Otherwise follow the predeclared decision rule. ## Data integrity law 1. Preserve raw/evidence provenance. 2. Do not discard failed generation attempts. 3. Do not overwrite historic model/capability results. 4. Separate accepted generation from editorial selection. 5. Separate provider latency from CartoonOS queue/review latency. 6. Separate provider cost from human/tool/compute cost. 7. Null ≠ zero. 8. Planning ≠ live. 9. Engagement ≠ learning. 10. Require sufficient matched evidence before autonomous routing/budget changes.