# Godfortune Seedance Quality Loop

Seedance 2.0 is normalized to 100 as a reference-only video-quality target.
Studio-Core does not call the Seedance API. The reference defines the required
motion stability, physical plausibility, instruction following, micro-expression
quality, multi-subject consistency, camera control, and temporal visual finish.

## Gate Order

1. Approve a production keyframe for each qualification shot.
2. Generate four I2V candidates from the same keyframe and prompt.
3. Reject hard failures before weighted scoring.
4. Shortlist two candidates with technical measurements.
5. Record human Seedance-relative scores for all eight categories.
6. Promote one winner per shot only when the total is at least 90 and no category
   is below 80.
7. Use the four approved winners in a ten-second integration test.

## First Local Candidate

The first measured engine is Wan 2.2 TI2V 5B FP16 through ComfyUI native
offloading. The official ComfyUI guide states that the 5B model should fit well
on 8GB VRAM, while the upstream model supports 720p I2V at 24fps and uses the
Apache 2.0 license.

The first quality recipe is 704x1280, 121 frames, 24fps, 20 steps, CFG 5,
`uni_pc`, `simple`, and model shift 8. No four-step speed LoRA or reduced-step
distillation is allowed in the baseline.

## First Measured Calibration

The first local shot completed in 1002.982 seconds. It is H.264 at 704x1280,
24fps, 121 frames, and 5.041667 seconds. Peak GPU allocation was 11911MiB and
minimum available system memory was 0.855GiB. Music and video generation must
therefore use one exclusive GPU queue; ACE-Step was stopped during Wan inference
and restored after the ComfyUI models were unloaded.

The internal visual calibration is 72.7/100 over the observable single-character
categories. This is not a qualification score. The input was a padded research
state rather than an approved full-scene keyframe, singing synchronization was
not present, and multi-subject interaction could not be observed. The result
proves that the engine is runnable and preserves the simple identity, but it
does not clear the Seedance-relative production gate.

## Approved-Keyframe Run

The user approved the 93.0-point HiDream neutral-front state on 2026-07-24.
Studio composited that state with the approved 94.05-point HiDream stage at
704x1280, recorded the layer hashes and ground-contact placement, and reserved
the lower caption-safe area.

The first production-keyframe motion test uses one bounded action: one gentle
sway, one presenting-hand raise, a smile, and a final hold with a locked camera.
This isolates temporal identity, body mechanics, instruction completion, hand
integrity, and background stability before song timing or camera motion is
introduced. The previous 72.7 calibration remains failure-history evidence and
is not reused as the production input.

The first approved-keyframe candidate scored a provisional 81.3/100 over the
observable categories. The character, stage, and locked composition remained
recognizable, but the left arm rose despite a right-hand-only instruction and
finger topology changed during the gesture. It is evidence of an 8.6-point
improvement over the research calibration, not an approved motion asset.

The second candidate keeps the same approved keyframe, model, resolution,
sampler, steps, and camera. Its defect-control prompt locks the left arm, requests
only a right-forearm presenting gesture, keeps both eyes open, and explicitly
forbids the two-arm pose and finger failures observed in the first candidate.
This run determines whether prompt control is sufficient or whether the pipeline
must introduce an approved end-pose or layered animation path.

The second candidate scored a provisional 84.2. It kept the left arm down and
reduced the hand failure, but ignored the open-eye instruction and did not hold
the final gesture. Prompt-only full-scene I2V therefore remains below the
production target.

## Separated Character Motion

The third candidate allocates 900 vertical pixels to the approved character on a
flat blue field, generates only the character motion, then composites it over
the exact approved background. FFmpeg combines a blue-screen alpha mask with a
spatially limited source-shadow mask, scales the generated canvas by 0.8, and
places the feet on the approved 825-pixel anchor. The source background never
enters the final shot.

The generated arm rise completed too quickly, so the first 0.75 seconds were
retimed to 1.25 seconds and 0.5 seconds of redundant hold were removed. The final
video remains 121 frames, 24fps, and 5.041667 seconds. No generative repaint,
frame interpolation, or character identity edit occurs during this timing pass.

This separated composite initially produced a 94.1 stability diagnostic over the
observable single-character categories. Human review correctly identified that
the score was inflated by low task complexity: the shot has one neutral
expression and one simple arm gesture. It does not contain a speaking or singing
face, a comedy reaction, or multiple independent motion beats.

The 94.1 value is therefore retained only as a separated-composite stability
measurement. Its qualification score is null and its complexity gate fails.
Future candidates must satisfy the shot's expression-count, motion-beat, vocal
face, and start/end eye-topology requirements before weighted Seedance-relative
scoring begins.

## Complexity And Scale Correction

Human review found that the first approved-keyframe candidate contained more
varied and natural motion than the separated stability candidate, but its
720-pixel character height was too large for the approved stage and its opening
and closing eye images did not match. These are independent quality dimensions:
motion variety cannot excuse composition scale or identity drift, and a simple
stable gesture cannot outrank a harder expressive shot without a complexity
gate.

Studio now renders a 540, 600, 660, and 720-pixel scale comparison at the same
825-pixel feet anchor. The 720-pixel placement is rejected and 600 pixels is the
working target for the next full-stage candidate.

The next complexity-matched shot must visibly complete all of these beats before
weighted scoring:

1. neutral start with both round eyes open;
2. a right-palm presentation while the left arm remains down;
3. a surprised mouth-and-brow expression with a small torso recoil;
4. a return to the original neutral eye topology and face.

Missing expression states, missing motion beats, or different opening and closing
eyes cause a hard failure and leave the qualification score null.

The first 600-pixel complexity-matched run completed the expression and motion
count, and the reduced scale fit the stage materially better. It still failed
before scoring: dusty-pink hair became orange, the right hand became an orange
object-like form, the opening round eyes became larger outlined eyes, and the
opening smile returned as a different worried mouth. The eye-region endpoint
MAE was 0.201739, retained only as a diagnostic; human topology review owns the
failure decision.

This establishes the next architecture boundary. Wan 2.2 5B may supply torso
timing and non-critical transitional deformation, but it cannot own exact
identity restoration in an expressive shot. The face, eye, mouth, and hand
states must come from approved layers or approved endpoint states. The approved
background remains static and the 600-pixel full-stage scale remains the working
target.

The first layer-locked candidate reused the stable separated right-arm motion,
played the raise forward and backward, inserted a fixed O-mouth and eyebrow
layer, applied a whole-character recoil, and ended on the opening frame. It
passed the visual expression-count and endpoint identity checks, but human
review rejected it. The eyebrow layer introduced anatomy that the source
hairstyle hides, and the retimed, reversed, and held motion was not natural.

Temporal diagnostics confirmed the review. The layer-locked candidate had a
0.683333 frozen frame-pair ratio and concentrated its discontinuities around
the expression overlay. The research calibration had no frozen pairs and the
first approved-keyframe candidate had a 0.016667 frozen-pair ratio. Low pixel
motion is no longer treated as a quality advantage.

## Direct Temporal Generation And Local Identity Repair

The next candidate rebuilt the approved stage at a 600-pixel character height
and restored direct Wan generation for all 121 frames. It used the same
20-step, 24fps, `uni_pc`, no-speed-LoRA recipe as the smoother historical
candidates. No generated segment was stretched, reversed, frozen, or assembled.

Direct generation restored continuous motion and retained the round eye
topology, but Wan drew eyebrow arcs during the surprise expression despite an
explicit hidden-eyebrow instruction. Prompting alone therefore does not enforce
this character invariant.

The repair candidate preserves the direct Wan frame order and timestamps and
suppresses only the two localized eyebrow regions after the arcs are detected.
It does not modify the eyes, mouth, hands, or body motion. Its frozen frame-pair
ratio is 0.083333, compared with 0.683333 for the rejected layer-locked
candidate. The frame-delta diagnostic finds no late discontinuities after frame
14.

This establishes the current production direction:

1. generate the complete shot motion directly in Wan;
2. keep the approved 600-pixel scene composition;
3. repair only localized identity violations;
4. never build primary body timing from reverse playback, hard holds, or a
   suddenly enabled expression layer;
5. keep the score null until human review confirms that the repair is invisible
   at playback speed.

The current candidate still holds its surprised mouth and raised hand instead of
returning to neutral, and the requested torso recoil is subtle. Those are
instruction-adherence issues for the next shot design, not reasons to discard
the recovered temporal architecture.

## Boundaries

- A good motion model cannot repair a weak input keyframe.
- The same approved input, prompt, resolution, duration, and output probe are
  used for every engine comparison.
- Runtime is measured but is not part of the quality score.
- Automatic image similarity and frame-difference metrics are diagnostics only.
- Human review owns expression, comedy timing, action completion, and the final
  Seedance-relative score.
- Seedance reference assets cannot be represented as locally generated output.

Machine-readable policy:
`config/godfortune_seedance_benchmark_v1.json`.

Primary sources:

- <https://seed.bytedance.com/en/seedance2_0>
- <https://seed.bytedance.com/en/blog/official-launch-of-seedance-2-0>
- <https://docs.comfy.org/tutorials/video/wan/wan2_2>
- <https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B>
