# AI Duel Long-Form Protocol

Version: 0.1 draft

Date: 2026-08-18

Status: Public draft. External review is requested. This version is not frozen.

License: CC BY 4.0.

## Purpose

This protocol measures complete AI video production systems.
It measures finished videos, production behavior, cost, and reliability.

It does not rank isolated clip models.

AI Duel can use this protocol for three result types:

- An open record with one qualified system.
- A head-to-head result with two qualified systems.
- A ranked leaderboard with three or more qualified systems.

An open record does not imply first place.

## Benchmark classes

### Best Complete System

Each system uses its preferred public stack.

This class measures the customer result.
The system can select its normal image, video, speech, music, and language models.

The run must use a public customer configuration.
Private settings require a separate exhibition label.

### Common Visual Engine

Each system uses the same video engine model and API version.

This class measures the production layer above the visual renderer.
It includes writing, planning, prompting, continuity, retries, audio, and assembly.

Auxiliary providers can differ.
The run manifest must identify every auxiliary provider.

The round must give each system the same engine credit limit.

## Interaction classes

The first public protocol uses the `autonomous` interaction class.

An autonomous run receives one brief.
No person can change the plan or output after the run starts.

A future `interactive` class can measure conversational directing.
Do not mix autonomous and interactive results.

## Duration classes

Use these target durations:

| Identifier | Target seconds | Public name |
|---|---:|---|
| `minute_1` | 60 | Complete Narrative Minute |
| `minute_2` | 120 | Narrative Short |
| `minute_5` | 300 | Short Film |
| `minute_15` | 900 | Episode |
| `minute_30` | 1,800 | Extended Episode |
| `minute_45` | 2,700 | Television Special |
| `minute_60` | 3,600 | Feature |
| `minute_120` | 7,200 | Long Feature |

Use the smaller of five percent or 60 seconds as the duration tolerance.

Do not count excessive credits, black frames, frozen frames, or repeated filler.
Freeze the detection thresholds before the round.

Preserve a usable output outside the tolerance.
Mark it as `duration_nonconforming`.

Publish it in the reliability result.
Do not include it in the main preference ranking.

Never combine duration classes in one rating pool.

## Prompt tracks

Use two prompt tracks.
Do not combine their results.

### Brief-to-Film

Give the system one outcome brief.
This track measures the complete consumer workflow.

### Script-to-Film

Give the system a locked script, beat sheet, and character bible.
This track measures production from a fixed creative plan.

Create an atomic requirement checklist for each prompt.
Hash it before generation.

Use it for prompt-adherence diagnostics.

## Prompt commitment

The complete input package must define these items:

- The story outcome.
- Required characters and canon facts.
- The required duration.
- The aspect ratio.
- The required language.
- The audio requirement.
- The prohibited asset classes.

A Brief-to-Film input must not prescribe a shot list.
The production system must create its own plan in that track.

Publish a SHA-256 hash before the first run.
Publish the exact input package when the round opens.

For duration qualification, use one core brief.
Change only the declared duration field.

Hash every duration variant before the first run.

## Run rules

AI Duel must operate every ranked system.
A provider cannot select its strongest output.

Use this sequence:

1. Record the system version and customer tier.
2. Record every selected setting.
3. Reserve the permitted budget.
4. Start the wall clock before submission.
5. Submit the exact committed brief.
6. Accept the first technically valid result.
7. Stop the wall clock after the final asset arrives.
8. Store the original output and checksum.
9. Validate the output without modifying it.
10. Create the neutral review proxy.

Do not retry a weak creative result.

Permit a retry only after one of these events:

- The provider returns an explicit error.
- The output file is empty.
- The output file cannot play.
- The output lacks a required media stream.
- The provider confirms an infrastructure failure.

Record the failed attempt and its cost.
Apply the round retry limit.

An automatic system retry is part of the production system.
It still counts toward cost and reliability.

## Common-engine controls

Freeze these values before a Common Visual Engine round:

- The engine provider.
- The engine model identifier.
- The API version.
- The provider region.
- The account tier.
- The permitted parameter surface.
- The resolution.
- The permitted engine inputs.
- The total credit limit for each system.
- The total production budget for each system.
- The deadline.
- The technical retry limit.
- The human-action limit.
- The rate limit and concurrency.

Route every request through an AI Duel gateway.
Randomly interleave requests from each entrant.

This reduces time and server-load bias.

Do not require one random seed.
Different production plans can contain different scene counts.

Permit each system to allocate its credit across scenes.
This allocation is part of the orchestration result.

Record these engine measures:

- Submitted engine prompts.
- Generated seconds.
- Discarded seconds.
- Provider request identifiers.
- Provider errors.
- Safety blocks.
- Engine cost.

Archive every raw engine output.
Archive the assembled film separately.

Use one public engine endpoint for every entrant.
Do not use private tuning for one entrant.

Freeze the permitted image, language, speech, music, and postproduction tools.
Declare whether native engine audio can be replaced.

Ban these items in the founding class:

- Another video engine.
- Private tuning.
- Private LoRAs.
- Video-to-video replacement.
- Generative upscaling.

## Source assets

The founding autonomous track permits no uploaded media assets.

Every system must create its own visuals and audio.
Stock media use must remain disabled.

Record any future source asset policy in the round manifest.
Do not mix asset policies in one result pool.

## Output requirements

A conforming output must contain:

- One playable video stream.
- One playable audio stream.
- The required aspect ratio.
- A duration inside the declared tolerance.
- A complete beginning, development, and ending.

Do not improve an output during normalization.

The review proxy must preserve:

- The complete duration.
- The frame order.
- The original frame rate.
- The original audio timing.
- The complete image area.

Do not crop, interpolate, upscale, or remix the output.

Fit each result inside the same review canvas.
Use neutral padding when necessary.

Apply the same `AI DUEL | DEMO ONLY` watermark to every proxy.

Do not remove a required provider watermark.
Get clean output or written permission before a blind ranked round.

## Cost reporting

Report these cost values separately:

- The actual billed cost.
- The equivalent public list cost.
- The sponsor credit used.
- The internal compute estimate.
- The storage and delivery cost.

For local compute, record these items:

- The GPU model.
- The GPU count.
- The measured GPU time.
- The peak memory use.
- The electricity estimate method.

Do not report free internal compute as zero customer cost.

## Human actions

Record every human action after submission.

Classify each action as one of these values:

- Observation only.
- Infrastructure recovery.
- Creative direction.
- Manual edit.
- Manual assembly.

Infrastructure recovery does not automatically invalidate a result.
The manifest must describe it.

Creative direction, manual editing, or manual assembly invalidates an autonomous ranking result.
Preserve the result as an exhibition when useful.

## One-minute and two-minute voting

Use blind pairwise voting.
Randomize the playback order for each voter.

Require at least 95 percent watch coverage for both outputs.
Require the ending of each output to play.
Require audio playback for the main result.

Lock the vote before revealing each system.

Use Bradley-Terry maximum likelihood estimation.
Publish confidence intervals and prompt sensitivity checks.

## Five-minute and longer evaluation

Publish two independent results.

### Full-film result

The full-film result measures the complete story.

Randomize the viewing order for each juror.
Counterbalance order with a Latin-square schedule.
Give every system pair equal complete-view exposure.

Require at least 98 percent playback and the ending.
Verify attention with simple story questions.

Do not include attention answers in the quality score.

Use separate sessions for 120-minute entries.
Set fixed breaks and a maximum daily screening time.

Use these minimum targets:

- Five minutes: 25 complete-view voters.
- 15 minutes: 12 complete-view jurors.
- 30 minutes and longer: Seven complete-view jurors.

Label small jury results as descriptive.
Use a power analysis before a formal winner claim.

Exclude entrant employees, investors, contractors, and affiliates.
Record expertise, target-audience fit, compensation, conflicts, and attrition.

Keep audience and craft panels separate.
Do not combine their scores.

### Public segment result

The segment result measures local production quality.

Freeze a deterministic selection algorithm before viewing outputs.
Do not let a person select flattering scenes.

Derive its random seed from the prompt hash and protocol version.
Use fixed relative-time strata across all systems.

Use this schedule:

| Duration | Segment sample |
|---|---|
| 5 minutes | Four non-overlapping segments of 30 to 45 seconds |
| 15 minutes | Six segments of 60 seconds |
| 30 minutes | Eight segments of 60 seconds |
| 45 minutes | Ten segments of 60 seconds |
| 60 minutes | Twelve segments of 60 seconds |
| 120 minutes | Sixteen segments of 60 seconds |

Call this result `within-segment quality`.

Add declared transition-centered excerpts.
Add nonadjacent character and world comparisons.

Publish these continuity diagnostics separately.

Publish segment results separately.
Do not claim that a segment result measures full story quality.

Do not combine full-film and segment scores in version 0.1.

Do not treat segment votes as independent film samples.
Cluster analysis by prompt, run, segment, and voter.

Bootstrap prompts and production runs.
Do not bootstrap individual votes alone.

## Quality dimensions

Ask full-film reviewers to measure these dimensions:

- Story structure and ending.
- Character and world continuity.
- Prompt adherence.
- Pacing and scene purpose.
- Speech quality.
- Music and audio mix.
- Visual quality.
- Overall publish readiness.

Require one holistic paired choice.
Publish each dimension separately.

Define scoring anchors before screening.
Do not create a weighted composite after viewing results.

Publish inter-rater agreement.
Publish the attention-failure exclusion rule.

Use automated critics for diagnostic reports only.
An automated critic cannot decide the public rank.

## Reliability result

Every started run remains in the reliability dataset.

Classify its final state with one value:

- `completed_conforming`
- `completed_nonconforming`
- `provider_failed`
- `system_failed`
- `budget_exhausted`
- `deadline_exceeded`
- `policy_blocked`
- `operator_invalidated`

Do not replace a failure with a provider showcase output.

## Run manifest

Every run must publish a machine-readable manifest.

The manifest must include these fields:

- Protocol version.
- Result type.
- Benchmark class.
- Interaction class.
- Prompt track.
- Duration class.
- System name and version.
- Engine names and versions.
- Test date and time zone.
- Prompt hash.
- Requirement checklist hash.
- Output checksum.
- Target, actual, and effective duration.
- Aspect ratio and resolution.
- Customer tier.
- Credit limit.
- Actual and equivalent cost.
- Wall time.
- Generated and discarded seconds.
- Retry count and reasons.
- Failure events.
- Human actions.
- Raw engine output hashes.
- Gateway log hash.
- Public proxy address.
- Transparency log root.
- Source asset policy.
- Audio provider details.
- Safety and moderation events.
- Sponsor and affiliate disclosures.

Private provider request identifiers can use salted hashes.

## Publication levels

Use these labels:

- `Open Record`: One qualified system.
- `Head-to-Head`: Two qualified systems.
- `Pilot`: Fewer than 100 valid votes.
- `Provisional`: 100 votes, five prompts, and two opponents.
- `Ranked`: 300 votes, six prompts, three opponents, and 75 voters.
- `Leader`: Stable results with non-overlapping intervals.

The standard vote thresholds apply only to one-minute and two-minute tracks.

A long-form leaderboard needs at least three prompts.
It also needs equal run counts for every system.

Publish a one-prompt comparison as a challenge result.
Do not present it as a general system rank.

Freeze these open-record controls before generation:

- The cost ceiling.
- The wall-time deadline.
- The retry limit.
- The human-action cap.

Label every open record as one prompt, one run, and no quality rank.

An endurance jury does not use the standard vote thresholds.
Publish its juror count and confidence limits directly.

Use a power analysis before a formal endurance winner claim.

## Transparency log

Use an append-only public transparency log.

Commit these records before identity reveal:

- The protocol version.
- The prompt hash.
- The requirement checklist hash.
- The budget and deadline.
- The run start.
- Every attempt.
- Every output hash.
- The final manifest.

Timestamp and sign each record.

A correction must add a superseding record.
It must not overwrite the original record.

Publish deidentified ballot records with these fields:

- The battle assignment.
- The side randomization.
- The vote result.
- The exclusion reason.
- The cohort flags.

Publish segment windows, jury assignments, order schedules, and calculation inputs.

Keep each result inspectable through a public review proxy.
Store the private original with an independent auditor when practical.

A checksum does not prove the contents of an inaccessible private file.

Close a challenge season when its required engine version becomes unavailable.

## Challenger path

Accept a challenger only through one of these paths:

- AI Duel operates and reproduces the run.
- An independent auditor witnesses the complete run.

Keep an unverifiable provider submission in the exhibition class.

## Governance

AI Duel must disclose its relationship with UNCEN.

Use this disclosure:

> AI Duel is operated by XP.COM, LLC, which also operates UNCEN. UNCEN is one measured participant.

Publish every sponsor, affiliate relationship, and credit grant.

A sponsor cannot control these items:

- The prompt set.
- The entrant list.
- The selected output.
- The ranking calculation.
- The publication decision.

Publish an unfavorable UNCEN result.

## Version changes

Freeze the protocol version before generation.

Do not modify a completed result after a protocol change.
Run the system again under the new version.

Publish a change log for every protocol version.

Version 0.1 needs external review before its first formal leader claim.
