1. 음악·가창 기준 확정
APPROVED · frozen_voice_render_v1 · d090_s20260821 · separated vocal ASR 90.11%
원본 시드 재현 대조군 seed 20260821입니다. semantic plan이 원본과 일치하고, 분리 보컬 ASR은 90.11%로 기록됐습니다.
가창 자체는 유지하고 발음이 뭉개지는 구간만 대본 구조에서 덜 깨지도록 수정 요청을 작성해 보낸 상태입니다.
현재 확인 자료: 운세 62초 연습 작품 v1
무료 OpenUtau 한국어 보컬 바로 듣기OpenUtau 0.1.565 portable · 결제/로그인 없음 · 84 BPM · WORLDLINE-R · 24초. 여성 KO CVC, 남성 KO CVVC로 음절을 직접 고정했습니다.
남성 14음절과 여성 7음절을 한글 포네마이저로 렌더했고, 상견례는 상견녜로 직접 교정했습니다.
두 트랙 모두 Whisper-small이 기대 가사를 회수하지 못했습니다. 남성 음원은 최종 자연 보컬로 부적합합니다.
ASR은 가창에서 오판할 수 있으므로 두 샘플의 실제 발음과 기계적인 음색을 직접 확인합니다.
남성 14음절 + 여성 7음절, MIDI 2트랙, 84 BPM, 24초. 한글 가사 이벤트와 한국어 음소 교정표를 포함합니다.
사용자가 좋다고 평가한 XL Turbo 음악을 기준으로 보컬과 반주를 분리했습니다.
남성 14음, 여성 7음, 한글 가사 21개, 전체 타임라인 24초가 MIDI에 들어 있습니다.
Windows/macOS ACE Studio Verse25에서 MIDI를 열고 Korean으로 합성한 드라이 보컬 WAV를 0초 기준 24초로 내보냅니다.
전부 한글 원문 + ko, 24초, 84 BPM, LM planning off. A와 B는 같은 시드이며 SFT는 공식 50스텝·CFG 7·shift 1 설정입니다.
공식 18.8GB bf16 가중치를 설치했고 12GB GPU 오프로딩으로 3회 생성에 성공했습니다.
세 SFT 시드 모두 요청 가사와 무관한 동일 문장으로 전사됐고 정확 구절은 0개입니다.
XL-SFT의 실제 청취 품질도 매우 나빠 ACE-Step 보컬은 최종 제작 경로에서 제외합니다.
ACE-Step은 반주·음악에만 쓰고, 한국어 음절과 음소를 음표별로 고정할 수 있는 별도 가창 합성기를 검증합니다.
ACE-Step 1.5 XL Turbo, seed 4106564636, 24초, LM planning off. 입력 표기와 vocal_language만 바꿨습니다.
세 후보는 같은 모델 상태, 시드, 길이, 템포, 프롬프트와 추론 설정을 사용했습니다.
영문 친화 표기는 한글 원문보다 개선되지 않았고 세 조건 모두 정확 구절이 0개입니다.
음악과 사람 음색은 좋지만 발음이 가사와 맞지 않고, 세 조건 모두 체감 개선이 없었습니다.
이 변환 규칙은 확장하지 않고 XL-SFT에 한글 원문을 넣어 우선 비교합니다.
No Qwen voice is inserted. Six local XL Turbo seeds were generated with LM planning disabled; A and B are the two highest-ASR candidates.
Dialogue-like recitation, singing and the continuing score are all produced inside one 20-second ACE render; no separately generated TTS is present.
The old 24-second A recovered only its first sentence. New A recovers the full spoken-sung-spoken structure after removing LM planning and reducing lyric load.
A transcribes as “오늘 끝나고 저랑 / 저녁 영화 생활 생견례 / 입 닫아 효도 되길”. It is not release-approved.
The user confirmed that both candidates have poor pronunciation and neither sounds like dialogue.
Test every line as music and align the video as a musical libretto, but require exact Korean diction before expanding beyond a short probe.
E1 was generated separately with Qwen3-TTS and inserted into the ACE soundtrack. It is retained as a male-timbre reference and fallback patch, not as the primary release route.
Faster-Whisper recovered exactly 하나, 둘, 셋 from the raw E1 output.
Male-anchor similarity is 0.610. The 2.377-second phrase fits the 2.770-second window with no stretching.
PCM hashes before and after the patch match B0 exactly; final duration is 90.000000 seconds.
The user confirmed that E1 sounds more natural than the rejected VC candidates.
The user confirmed that the inserted voice reads as male.
This is a separately generated Qwen voice inserted into ACE music. The primary route remains a single ACE generation containing dialogue, singing and transitions.
The user rejected both candidates as female-like, electronic and unstable. Automatic QC now rejects them before review; they remain here only as failure evidence.
The references labeled male had median F0 values of 291/310 Hz and speaker similarity of -0.025/-0.036 to the approved dry male anchor. They were extracted from a mixed soundtrack and are now forbidden as cast references.
The original 하나 scored 0.0 stable voiced-frame ratio and 0.019 voice confidence. A timbre converter can preserve content, but it cannot reconstruct missing natural articulation and prosody.
C1/C2 voice confidence was 0.060/0.010 and raw duration was only 86% of the source. The unseeded upstream runner also changed C1 from 164 Hz to 50 Hz on an identical rerun.
Clean male self-conversion retained 0.83 voiced stability. The seeded runner produced identical SHA-256 outputs, but direct short-word splicing remains blocked until duration handling is redesigned.
The candidate is retained only as failure evidence. It is not awaiting approval and must not be used as the next production base.
B0 target RMS is -22.91 dBFS; the raw ACE target is -57.05 dBFS with no voiced pitch. This was silence/noise replacement, not voice repair.
The 1.7-second mask was below the documented 3-second Repaint range and the full timeline lyrics were replaced by one local word. Either correction alone restored audible output in controlled tests.
The 4.5-second reference is the full B0 mix from source seconds 6.0-10.5, not isolated male speech. ACE Repaint uses it as audio context and does not guarantee speaker identity.
Local ACE-Step 1.5 XL Turbo + 1.7B LM. API cost: 0. Listen in A, B, C order.
A took 76.56s including XL model load; B took 19.48s with the model loaded.
Whisper recovered only “오늘 끝나고 저랑” from A and “마카트 만나고 저런” from B. This is not a release pass.
B changed audio outside the requested window; the post-window PCM control did not match. Compare B directly with C.
A omitted most words and B failed repaint locality. Retain these files as failure evidence only.
The user rejected this assembly route. It remains below only as comparison history.
Separate TTS and music editing did not solve cast and song-to-dialogue continuity.
Use the ACE unified dialogue and music comparison above.
Qwen3-TTS VoiceDesign candidates generated locally on studio. The full preview keeps the approved vocal and BGM at 1.0x; dialogue is natural-speed timing audio.
108.741167 seconds. Dialogue, vocal, BGM, SFX, and master durations match. The locked musical source is not stretched.
Judge natural Korean pronunciation, speaking pace, character fit, and whether A/B/C should anchor the final cast.
The full preview still uses timing-only dialogue and a silent SFX stem. Final dialogue generation and human mix approval remain required.
현재 출시 준비도 40/100 · 필수 출시 게이트 0/5. 이 영상은 조립 구조 증거이며 완성본이 아닙니다.

20개 샷, 62초 캔버스, 자막, 8개 SCAIL 클립 조립은 확인했습니다.
90초 가창 구절 45.68초를 26.67초로 압축했습니다. 평균 1.85배, 최대 4.06배입니다.
대사 자연 길이 측정 → 대본 타이밍 수정 → 26.67초 전용 노래·분리 스템 승인 순서입니다.
7개 기본 팩과 28개 분할 참조가 생성됐습니다. 연습 파이프라인에는 사용 가능하며 제작 승인은 아닙니다.
전 캐스트를 fortune_v1_practice에 등록했습니다.
합창단 동작 크기, 소연 얼굴 근접 눈 비율, 점장 상반신 동작 과장을 기록했습니다.
character_06은 소연 팩에서만 파생하며 별도 정체성을 만들지 않습니다.
두 시트의 역할 구분과 한 작품 안의 조화를 먼저 확인합니다. 승인 후 인물별 4참조 팩으로 확장합니다.
캐스팅 검수 대기입니다. 아직 최종 모션 참조나 제작 승인 캐릭터가 아닙니다.
전신 중립, 얼굴 근접, 상반신 손, 안전 단일동작의 4참조 팩을 만듭니다.
character_06은 소연 원본에서만 파생하며 새 정체성을 생성하지 않습니다.
Human approved on 2026-08-05. The 768x1024 references use aspect-preserving center crop for the 512x896 canvas. Production restriction: one separated moderate-speed action per clip; crossed or overlapping hands remain blocked. 33 frames, 16 fps, 8 steps, DPO, seed 42, runtime 8m 36s.




Human approved on 2026-08-05. Production restriction: one separated moderate-speed hand gesture per clip; crossed hands, overlapping hands, and fast palm rotation remain blocked. 512x896, 33 frames, 16 fps, 8 steps, DPO, seed 42, runtime 8m 47s.



Human review: the source C2 hands have five normal fingers, but fingers deform when the generated hands cross, overlap, and rotate. 512x896, 33 frames, 16 fps, 8 steps, DPO, seed 42, runtime 12m 43s.


Human-approved Yaco references, unchanged official hard arm-motion driver, 512x896, 33 frames, 16 fps, 8 steps, DPO, seed 42, runtime 8m 48s.


Four official references, hardest arm-lift segment, 512x896, 33 frames, 16 fps, 8 steps, DPO, seed 42, runtime 8m 39s.


Official reference and driver, 384x672, 33 frames, 16 fps, 30 steps, seed 42, runtime 4m 54s.

H3 is ignored for this build. Character performance stays in LivePortrait frontal dialogue/singing and SING-LOCK-2 subtle sway; background, props, effects, captions, camera crop, and final timing are editor-owned layers.
The current finished file is 30.0s, 1080x1920, 30fps, with script characters, generated background, effects, Korean captions, and locked song excerpt.
Do not open another quality stage. Fix only visible local issues in this completion file if they are cheap and bounded.
No character prop pressing, same-frame multi-character acting, walking, or independent limb gestures in V0. Those beats are cutaways or editor overlays.
H3 authorization is pending. This board freezes what the existing local route can safely do now: stable frontal singing is safe, subtle sway is conditional, old RIG-5 mouth clips remain blocked.

Stable frontal singing/dialogue close-ups using the approved LivePortrait full phrase path. Machine gate: 125/125 mouth attachment pass.
SING-LOCK-2 subtle whole-character sway can proceed only if human review accepts that it reads as body motion rather than jitter.
Old RIG-5 mouth-layer output, K2 reaction, walking, and independent hand/body gestures are not approved by this route.
This verifies that the approved 5.2s LivePortrait singing mouth remains attached when the whole character is moved slightly. It is deterministic post-motion, not generative body acting.

PASS_SING_LOCK2_SUBTLE_SWAY_REVIEW_READY. 125/125 frames passed mouth attachment against the moving expected socket.
Judge whether the subtle movement improves the locked singing shot without looking like camera jitter.
This does not approve independent hand gestures or model-generated body motion. Those remain separate one-action clip gates.
Expanded from the short syllable probe to the full 5.2s locked vocal phrase. Uses +120ms viseme delay and stable K0/O/E-I/U/A Yaco heads through native LivePortrait mouth motion.

PASS_FULL_PHRASE_TIMING_REVIEW_READY. 125/125 frames kept detected lips inside the approved mouth socket.
Judge whether the 5.2s phrase now feels acceptable for singing timing, especially U and A mouth shapes in the later half.
This approves only stable-head singing if accepted. Body motion, hand gesture, and reaction remain separate one-action clip gates.
This uses the locked vocal d090_s20260821, the approved +120ms viseme delay, and stable K0/O/E-I Yaco heads. Review whether the mouth remains attached while following the sung syllable timing.

PASS_SYLLABLE_TIMING_REVIEW_READY. 36/36 frames kept detected lips inside the approved mouth socket.
Approve only if O and E-I read as attached to the face and timing feels acceptable against the locked vocal.
This still covers stable head singing only. Wider pose, gesture, reaction, and body motion must be tested as separate one-action clips.
Stable K0, A, O, A, K0 heads were run through ComfyUI-LivePortraitKJ directly, without restarting the current ComfyUI service. This is a short recognition and integration probe, not final production approval.
PASS_NATIVE_MOUTH_MOTION_REVIEW_READY. lip_on scored 5/5 mouth-alignment pass; detected lips stayed inside the approved mouth socket.
Please judge whether the open-mouth frames read as attached to the face rather than a floating layer. If accepted, expand to a short syllable-timed clip.
K2 remains rejected for native mouth probing. This route is allowed only for stable K0/A/O-style shots until wider pose states pass recognition.
Detection-only probe using ComfyUI-LivePortraitKJ FaceAlignment + BlazeFace back camera. Green is the approved mouth socket, yellow is detected lips, red is nose, blue is eyes. This checks alignment before any video generation.

ALLOW_STABLE_K0_AO_PROBE_REJECT_K2. Native mouth probing is allowed only for stable K0/A/O-style close shots, not K2 reaction shots.
K0, A, O, and mouthless base place detected lips inside the approved socket. K2 detects an incorrect lower region as the mouth.
This explains why the last attached-mouth clip could work: attachment can pass on stable head states while detector recognition still fails on some pose/expression states.
This clip remains useful as workflow evidence, but it is not a visual approval candidate. The mouth patch appears independent/slipping, so mouth recognition and integration must be gated first.
Do not advance to cleaner final state art until the mouth recognition gate is passed or deterministic head-state integration is explicitly selected.
One prompt creates one clip with one primary action; the editor assembles the full sequence on locked audio.
The current mouth treatment can read as a separate layer. Treat this as a failed mouth-integration sample.
RIG-5 and video-generation-native shots both follow the same rule: a prompt unit may contain exactly one primary action. Multi-beat behavior is assembled by the editor timeline.
One prompt, one clip, one primary action. Mouth, blink, small bounce, background, camera, and effects are secondary layers.
The editor assembles idle singing, hand emphasis, fake walk, shock reaction, and verdict punch into the final beat sequence.
Prompts that combine walking, singing, hand raise, prop action, and reaction in one generation unit are rejected.
RIG-3 proved head and mouth control. RIG-4/RP-3 failed on free arm/body rigging. The next product route is a small approved pose-state pack, short deterministic clips, and editor assembly.
Front singing, narration, reaction cuts, fake walking, verdict beats, camera punch, background pan, and deterministic mouth.
Hand gestures only if the whole arm/prop is authored as one attached state. No loose limb rotation.
Continuous arm choreography, finger-accurate props, real walk cycles, turnarounds, camera orbit, and arbitrary new limb anatomy.
A is the current stabilized assembly. B applies FFmpeg minterpolate at 48fps. C is an editor-only anchor/keyframe baseline. This checks whether interpolation actually fixes shake or only smooths generated deformation.
Interpolation should not be treated as the main shake fix. It changes temporal smoothness, not generated body/hand deformation.
Keep interpolation as final polish only after the source clip is already stable.
Move upstream: shorter generated motions, cleaner first frame, and anchor/keyframe clip planning.
Same timeline as V1, but S04/S03 receive face-box scale and center stabilization before final encode. FFM02 keeps the intentional head lift. This is a review candidate, not a locked pass.
If V2 reduces the scale jitter without locking the body unnaturally, stabilization becomes part of the assembly engine.
If shake remains visible, the fix must move upstream: first-frame framing, shorter motion, or different source clip selection.
Wan remains the clip generator. Timeline, crop, camera feel, and stabilization belong to the separate editor layer.
This validates the product structure: Wan generates short approved clips; the assembly engine handles timeline order, hard cuts, and post-crop camera feel. Camera push-in is applied after generation, not inside the Wan prompt.
Wan can stay as a short-clip generator while a separate editor assembles clips and applies camera feel.
Add audio beat alignment and optional clip-level stabilization/crop rules before final production template lock.
Current version uses hard cuts only. Crossfades are avoided until face morph risk is tested.
Checklist tested: center of gravity, gaze direction, one light direction, hand/prop contact, and motion room. FFM01 combined head lift and camera push-in and failed. FFM02 used head lift only and is the current pass candidate.


Use controlled first frames and one primary motion per generated clip.
Promote FFM02 to 5-second head-lift-only validation, or add camera push later with post-crop/stabilization instead of in the generation prompt.
Do not combine generated subject motion and generated camera push-in at this stage.
Face-box diagnostics confirm the user observation: S04 has materially higher height, area, and center jitter than Realistic S01. Do not promote S04 to 512x896 until scale-lock or post-crop stabilization is tested.

S04 shake is consistent with subject size/framing drift. This must be handled before template lock.
Run one scale-lock prompt test and one post-stabilized crop test at 384x672 before any 512x896 promotion.
This is not a Wan replacement yet. Review it as a separate long-video/avatar candidate while C2 safe remains the active product path. The first decision is whether RTX 3060 12GB single-GPU preflight is realistic.
LongCat targets longer story videos, video continuation, and avatar-style generation. That matches our next concern after short 5s C2-safe clips: continuity across script-driven shots.
Do not install or promote before confirming single RTX 3060 12GB feasibility, INT8/low-step availability, license terms, and whether audio/image/avatar mode can run locally.
Wan C2 safe S04 remains the current 5-second default candidate. LongCat is only a research candidate until it passes preflight.
Automation rule: do not force video-generation tools into loose-part rigging. Use full-frame storyboard/keyframe control for Wan/LTX/Seedance-like routes, and use validated rigpacks only for deterministic character systems.
Video-generation-native: full-frame anchor image, full-frame storyboard/keyframes, model-native I2V/keyframe/control, then visual and runtime gates.
Deterministic rig: validated rigpack, semantic anchors, motion primitives, then optional model polish.
Loose generated body parts cut from finished images and rotated without socket-authored joint metadata.
Music/vocal remains locked. Video generation is now locked to realistic / photoreal character anchors because Wan mouth behavior is materially cleaner on real-photo input and Hunyuan full-quality runtime is rejected.

실사/실사풍 캐릭터 영상을 기본 제품 경로로 고정. Direct 2D Yaco I2V is rejected.
Generate realistic identity anchors, then test speaking, gesture, body/background, and audio-sync composition against the locked vocal.
Fast Probe is now a reject/promote tier only. Final production settings must be validated separately at the selected working or final resolution.
Low resolution pass promotes to mid/final; it never locks final quality by itself.
Current Wan 5s at 512x896 already peaks near 11.25GB, so high-resolution final qualification needs a selected feasible route.
Four 5.04s C2-safe script shots were generated at 384x672. S04 is the current default candidate. S03 is an expression support lane. S01/S02 are kept as rejection evidence because 5-second hold introduced background or graphic artifacts.

S04 is the current C2-safe default candidate for automated 5s speaking shots.
If S04 passes human review, promote S04 to 512x896 5s mid validation before locking the script-server template.
S01 and S02 show that 5-second holds can still create non-product artifacts depending on prompt/seed.
C2 original keeps the right method shape but has a plant-like head artifact. C2 safe removes the large gesture and is the current default candidate. C3 fast is rejected for duplicate arm; C3 midres fixes it and remains a supporting body-motion lane.

C2 safe is the current default method candidate for automated script-server speaking shots.
C3 midres is usable for occasional body-motion shots, but not as the default because it costs 93s for 2s and needs higher resolution.
C2 original has a head artifact. C3 fast duplicates the right arm. C2 clean repair is discarded because it introduced colored graphic artifacts.
Three realistic anchor candidates were generated and tested with identical Wan2.2 Fast Probe settings. This is a promote/reject gate before 5-second Mid Validation.

Pick one or two candidates to promote to 5-second Mid Validation. Do not lock identity from this fast tier alone.
Candidate scope is now HunyuanVideo 1.5 I2V versus Wan2.2 TI2V 5B. This card shows what needs visual review; planning stays in the thread.

Wan S01 is ready for human review against the script-driven mouth/expression gate.
Only low-frame/low-step fast probes are allowed now; full-quality Hunyuan script scenarios are runtime-rejected for the product target.
LTX is out of the current comparison set.
This isolates whether the bad mouth behavior is a general Wan limitation or a 2D character-face interpretation problem.
Real-photo mouth behavior is materially cleaner than the 2D Yaco S01 mouth output.
Wan may need realistic/intermediate face conditioning or a separate 2D mouth layer for Yaco instead of direct 2D talking-face I2V.
Same process as music/vocal: compare model families first, then freeze the best route. Current candidates: HunyuanVideo 1.5 I2V, Wan2.2 TI2V 5B, LTX 2.3 MSR, and Wan2.2 Animate 14B.

No more locking Wan/VG from one promising sample. Tool-family bakeoff comes first.
Wan Animate 14B via current WanGP route is preserved as rejected for Yaco production.
Review Hunyuan against Wan2.2 TI2V 5B and LTX. If Hunyuan passes, run torso translation and speaking close-up cases before locking.
This tests the video-generation-native route: one full-frame character image goes into Wan2.2 TI2V, not loose arm parts. Runtime: 25 frames / 1.04s output / 150.079s elapsed / peak 11.9GB VRAM on RTX 3060.
Unlike RP-3, the arm does not read as a separate rotating object. This supports continuing video-generation-native tests.
Face/expression drift appears and this run only used start-image I2V, not model-native storyboard/keyframe control.
Run full-frame K0/B2/B3/K1 storyboard/keyframe control and compare adherence, identity drift, and runtime.
External/generated right-arm parts were extracted, mapped into the same rigpack contract, validated with file checks, and rendered through deterministic right_hand_raise_v1. Product fit is rejected: the arm behaves like loose pieces rotating separately from the character, so this route must not be used for automated shorts production.

Chroma source ingestion, part extraction, rigpack file contract, validator --require-files, and deterministic render path are working.
The limb reads as separate moving pieces. This is not acceptable even though the rigpack schema and render path pass.
Stop loose generated-part injection. Move to socket-authored rig assets, Live2D/Spine-style rigging, or a full-body keyframe route that preserves integrated limbs.
A real rigpack manifest with required files and shoulder/elbow/wrist anchors. Validator passes with file checks; arm art is procedural placeholder.
Rigpack manifest, required asset files, required anchors, and validator with --require-files are working.
The procedural arm art is not production quality. It proves the product structure, not final character finish.
Replace procedural placeholder arm parts with generated or drawn production arm segments while preserving the same anchor contract.
Fixed body shoulder coordinate plus per-part local shoulder anchors. This checks whether attachment drift is solved before any video model polish.
Placement is now based on a fixed shoulder anchor instead of each part image origin.
The available arm drawings are not true same-joint rig parts, so attachment still looks rough even when the anchor is fixed.
Create or redraw arm parts with explicit shoulder, elbow, wrist anchors before testing model polish.
2D transparent-part composition, no generative interpolation. Rejected because pose parts were placed by image origin rather than shoulder anchors.
Ghosting is structurally removed because old arm pixels are not blended by a video model.
Arm parts do not stay attached to the fixed body joint because no shoulder anchor map was used.
Compare against the anchor-locked v2 test above.
WanGP LTX-2 2.0 Distilled GGUF Q4_K_M 19B. Portrait k0/k1/k2/k3 injected at frames 1/17/33/49, generated in 2m21s.
Portrait output is now locked at 576x768.
Final hand still shows bright smear/afterimage, so this is not approved without human review.
If rejected, move away from free video interpolation for large 2D arm lifts and test rig/render-first composition.
WanGP LTX-2 2.0 Distilled GGUF Q4_K_M 19B. Same C2 k0 to k3, 49 frames, generated in 1m58s.
Right arm rises, but the previous arm shape remains as a visible ghost/afterimage.
The 16:9 source keyframes overrode the requested portrait resolution and produced 896x512 output.
Compare against the portrait Inject Frames test and decide whether LTX should be restricted to shorter/subtler arm motion.
WanGP LTX-2 2.0 Distilled GGUF Q4_K_M 19B. C2 k0 to k3, 25 frames, generated in 1m54s.
Promising for short arm interpolation: identity and face are mostly stable, and the arm moves toward the target pose.
Hand detail is simplified and output resolution changed to 896x512. Not production-approved yet.
If accepted visually, run a 2s C2/C2-mouth-stable test and compare with Wan body-sway route.
Review these inputs before spending GPU time on Wan/LTX generation.



Check identity, whole-body movement, mouth attachment, and arm keyframe readability. Pass only the inputs that are visually useful.
This is not a model output. It only decides whether Wan/LTX short runs are worth launching.
Wan Animate: short body-sway run from the prepared 480x832 input. LTX: first/last keyframe 1-2s run.
Yaco/D-01 route abandoned. Music/vocal remains locked; video research now starts from this benchmark character pack.


Character identity, outfit, open-palm arms, viseme references, and storyboard inputs are now judgeable enough for Wan/LTX/Seedance-like route testing.
Run detector-friendly short preflight: static identity, arm gesture, body sway, mouth/face attachment, then reject or keep by visual review.
Thin chroma-key residue remains around some small eye/mouth parts; final rigging needs edge cleanup before production.
16:9 · 상반신 · 양손/얼굴 가독성 · 전체 프레임 재생성 · 후보 3개 · 총 7분 43초
WA01이 가장 Wan 입력 구조에 가깝습니다. foreground 폭 70%로 얼굴·양손·상체가 크게 잡힙니다.
WA01이 여전히 야코로 보이는지 사람이 먼저 판단해야 합니다. 통과하면 mask와 자체 driver preflight로 넘어갑니다.
팔 cutout, arm paste, 외부 스타일 복제는 사용하지 않았고 앞으로도 금지합니다.
Status: SIMPLE_TEXT_MOTION_ONLY_STORYBEAT_REJECTED_NOT_DEFAULT_PRODUCTION. MSR cannot use direct VG Control Video. Simple text motion is possible; storyboard timing prompt is rejected.

MSR remains useful only for simple selected-shot identity motion. Storyboard-level timing control is not reliable in this route, direct video-driver conditioning is unavailable, and runtime is too high.
Status: NO_OBJECTIVE_UPGRADE_DETECTED. Raw VG is stable, but does not prove quality improvement over the deterministic guide.
SSIM stays around 0.962 and mean absolute difference is only about 3.6, so outputs are close copies. Temporal and edge changes are lower than the guide, indicating reduced motion/edge activity rather than a clear line-quality upgrade.
Same neutral guide and seed 2026072802. Tests whether Raw VG is useful beyond stable guide-copying.
Raw VG remains the most stable LTX2 Q4 route, but strength changes are visually small. It behaves like stable controlled copy/cleanup, not a clear generative quality upgrade.
Neutral deterministic guide · compare guide vs Canny EVG vs Raw VG. Raw is best so far, not final approval.
Raw control is the strongest local LTX2 route so far. It reduces face drift compared with First/Last and Injected Frames. Next: Raw control strength sweep, then LTX-2.3 MSR/reference sheet if needed.
Same seed 2026072802 · video_prompt_type=KFI · compare K1 injected at L vs 13+L.
Injected frames are valid in WanGP CLI and speed passes. They do not materially improve Yaco face identity control. Next validation should use control video or pose/depth/canny, then MSR/reference sheet if identity drift remains.
Same seed 2026072802 · First/Last Frame · compare input_video_strength 1.00 / 0.85 / 0.65.
Speed passes and arm motion is present. Strength alone does not materially fix face/eye/nose drift. Next LTX validation should move to injected frames, then control video or MSR/reference sheet.
WanGP · LTX2 Q4 · image_prompt_type=SE · 25 frames · 24fps · 1m45s. This is the first fair usage-pattern check, not a final route decision.
Movement control is materially better than the smoke test. Arm motion appears, but face/eye/nose style drift remains visible. Next validation should compare lower/higher input strength and injected-frame positioning before judging LTX.
RTX 3060 · 25프레임 · 24fps · cold run 1분 38초 · 얼굴/입 제어는 미승인
속도 후보로는 보류 통과지만, 얼굴·입 제어와 몸짓 지시 이행은 미승인입니다. 다음 테스트는 start/end 키프레임 보간으로 제한합니다.
원본 BF16 · 768x1024 · 97프레임 · 30 steps · 2시간 24분


손을 드는 속도와 멈춤이 자연스러운지, 얼굴이 승인 Yaco에서 부드럽게 변형되지 않는지, 손가락이 중간 프레임에서도 유지되는지 확인합니다.
공식 참조 이미지+드라이버 · 25프레임 · 1 step · 캐시 후 2분 26초 · 모델 정상 입력 대조군

공식 입력에서는 정체성·상체·양손 레이아웃이 유지됩니다. Wan2.2 Animate 자체를 야코 이전 결과만으로 리젝하지 않습니다.
야코도 같은 구조의 상체 참조, 깨끗한 마스크, 비율이 맞는 드라이버로 다시 25프레임 캘리브레이션합니다.
이 외부 이미지는 제작·학습·스타일 복제에 사용하지 않습니다. 구조 분석만 허용합니다.
검출 성공 대조군 · Aligned Open Pose · 105프레임 · 30 steps · 1시간 57분

Yaco K0 · 121프레임 · 30 steps · 2시간 21분 · 모델 자체 최종 리젝 아님

몸 고정 · 목 관절 기준 머리 이동·회전 · +120ms
기존 비교 영상과 같은 조건인 704×1280, 121프레임, 24fps, 5.04초, Wan2.2-TI2V-5B, seed 2026072428로 재현합니다. 캐릭터 설계에 맞춰 눈썹 문구만 “낮은 회의적 눈썹 유지”로 바꿨습니다.
판정 순서: 얼굴·괄호형 옆머리·손·표정·시간축이 모두 버티면 01로 다음 관문을 진행합니다. 하나라도 핵심 정체성이 무너지면 01을 확정하지 않고 다른 후보로 넘어갑니다.
기술 판정은 리젝: 얼굴과 괄호형 머리는 유지됐지만 오른손 단독 동작이 양팔 벌림으로 바뀌었고, O자 입과 눈 확대가 약하며 정지 구간이 길었습니다. 캐릭터 붕괴보다 동작·표정 지시 실패가 주원인입니다.
대체 후보도 리젝: D-04는 뿔 컬과 화난 얼굴로, D-05는 눈 감은 장신형과 흐린 선화로 이동했습니다. 후보 번호 추가를 멈추고 D-01 정체성을 참조로 고정한 포즈 제어형 키프레임 빌더로 전환합니다.
01 영상 분석은 유지 · 04와 05는 이미지 생성 단계에서 방향 이탈로 리젝

APPROVED · frozen_voice_render_v1 · d090_s20260821 · separated vocal ASR 90.11%
원본 시드 재현 대조군 seed 20260821입니다. semantic plan이 원본과 일치하고, 분리 보컬 ASR은 90.11%로 기록됐습니다.
가창 자체는 유지하고 발음이 뭉개지는 구간만 대본 구조에서 덜 깨지도록 수정 요청을 작성해 보낸 상태입니다.
REJECTED_FOR_YACO_CHARACTER_BASELINE · CG-G0 · 서버 RTX 3060 11.63GiB
지금 영상 후보와 HiDream 캐릭터 마스터는 야코 상대 영상용 기준에서 리젝입니다. 영상 점수 계산이나 제작 승격으로 넘기지 않습니다.
FLUX는 표정은 읽히지만 캐릭터가 변했고, Illustrious는 표정 분리가 안 됐으며, 마스크 편집은 입·눈썹 패치가 드러났습니다. 고정 레이어만 VBT-3 대조군으로 유지합니다.
FLUX.2 Klein, HiDream O1, SDXL 제어형 스택을 같은 캐릭터 사양으로 비교합니다. 신규 모델은 공식 실행성을 확인한 뒤에만 추가합니다.
무표정, 가창, 극단적 당황에서 머리·의상·색·비율을 유지하면서 눈·눈썹·입·상체 포즈가 함께 바뀌어야 통과합니다.
놀람 표정, 오른손, 입 모양, 대본 수정 이후 가창/대사 싱크가 장면 의도와 맞는지 다시 봅니다.
공식 복합 동작 샘플을 100점으로 고정했습니다. Seedance API는 호출하지 않고 로컬 결과의 비교 기준으로만 사용합니다.
FP16 모델, UMT5 FP8, Wan VAE를 native ComfyUI offloading으로 구성했습니다. 4-step 속도 LoRA는 사용하지 않습니다.
직접 Wan 7차 눈썹 국소 복원 완료 · 시간축 유지 · 사람 검수 전 점수 보류
승인 키프레임 4종에서 장면당 네 후보를 만들고 총점 90 이상, 항목별 80 이상, 하드 실패 0건을 사람 검수로 통과해야 합니다.
HiDream-O1 Full의 4MP 캐릭터 1차를 상태팩 기준 100으로 고정했습니다. FLUX.2 Klein Base 4B 평면 마감은 사용하지 않습니다.
HiDream 수정본 내부 94.05점으로 안전 구도와 클린 플레이트 편집성을 포함한 모든 수치 항목을 통과했고 사람이 승인했습니다.
동일 원화·seed·짧은 문법의 3차에 작은 눈 하이라이트 레이어만 더한 93.0점 후보를 사람이 승인했습니다. 큰 반사형 눈은 별도 코믹 효과 상태로 분리합니다.
포즈 자체는 참고할 수 있지만 평면 스타일 결과물은 상태팩 자산으로 재사용하지 않습니다. 재현 시험 통과 후 HiDream으로 다시 생성합니다.
발음 · 발음만 보면됨
발음 · 가창은 좋은데 발음만 뭉개짐
semantic plan 해시가 5개 후보에서 동일합니다. 원본 WAV와 정확히 같은 재현 시드는 20260821입니다.
가사 정확도가 19.78%에서 42.86%까지 벌어졌습니다. 음색과 반주 마스킹을 별도 품질축으로 관리해야 합니다.
분리 보컬은 62.64%에서 90.11%입니다. Demucs 재분리 오차가 포함되므로 보조 지표로만 사용합니다.
음악성은 같은 상태이므로 한국어 자음 선명도, 비음·전기음, 보컬 전면감이 가장 나은 후보를 직접 골라야 합니다. ASR만으로 제작 승격하지 않습니다.
각 구간을 원본에서 독립적으로 생성했습니다. 채택하지 않은 부분은 최종 합성에서 전혀 바뀌지 않습니다.
지정 범위 밖 PCM까지 바꾼 후보는 청취 목록에서 제외했습니다.
세 구간에서 하나씩 채택하면 선택된 파형만 원본에 결합할 수 있습니다.
60·75·90초 비교에서 음악성이 가장 좋았습니다. 해당 판정은 자동 ASR보다 우선합니다.
분리 보컬은 90초 83.52%, 105초 71.43%, 120초 62.64%입니다. 길어질수록 반복과 변형이 늘었습니다.
105초와 120초가 더 음악적으로 좋다면 길이를 채택하고 가사 정렬을 별도 개선합니다. 개선이 작으면 90초에서 멈춥니다.
길이 최종 선택과 가사 전달력 보강 전까지 콘텐츠 서버 타이밍 변경이나 본편 렌더를 시작하지 않습니다.
0.6B의 높은 ASR보다 사람이 확인한 음악적 완성도를 우선합니다. 앞으로의 시드 검색도 1.7B 안에서 진행합니다.
90초 seed 20260821의 분리 보컬 기준입니다. 60초와 75초 후보도 함께 들어보고 음악성이 가장 좋은 곡을 선택해야 합니다.
LM 자유 길이는 자연스러운 완곡을 만들지만 현재 콘텐츠에는 과도합니다. 60·75·90초 상한 안에서 한 번에 생성했습니다.
모든 곡은 승인 단어와 순서를 유지한 단일 생성입니다. 구간 이어붙이기와 생성 후 시간 압축 없이 만들었으며 아직 게시 가능 상태가 아닙니다.
전체 믹스 93.41%, 분리 보컬 90.11%입니다. 음악성·자연스러움은 직접 청취 승인 전입니다.
90% 정확도뿐 아니라 순서 단어 100%·누락 0개를 요구하므로 아직 제작 승격하지 않습니다.
현재 설정 28.33초와 결말 7초를 보존한 계산입니다. 62초를 유지하면 설정부가 5초뿐이라 승인된 19개 설정 세그먼트를 보존할 수 없습니다.
곡을 먼저 고른 다음 콘텐츠 서버에는 대사·가사·순서를 바꾸지 않고 세그먼트 타이밍과 목표 길이만 다시 요청합니다.
최대 CUDA 9.73GiB로 26.60초 48kHz 스테레오 WAV를 생성했습니다. 유휴 시 모델은 CPU로 내려갑니다.
분리 보컬 기준 순서 회수 12.50%, 누락 35개입니다. 90%·전 단어 회수 기준을 통과하지 못했습니다.
길이만 72초로 늘리면 50.55% / 순서 47.50%로 상승하지만, 3섹션 단순화는 27.47%에 그쳤습니다.
스타일만 바꾸자 믹스 25.27%, 분리 보컬 35.16%로 개선됐습니다. a cappella는 32.97%로 더 낮아 악기 제거만으로는 부족합니다.
ACE XL은 연속 반주 후보로 유지합니다. 승인 가사는 dry 보컬을 화자별 짧은 구간으로 생성하고, 정확도 통과 후에만 코러스·리버브·마스터링을 적용합니다.
문장 3회 반복 압축을 제거하고, 승인 문장을 1회만 합성했습니다. 보컬 단독 93.41%, 사이드체인 덕킹 적용 믹스 92.31%이며 사람 청취 검수 전 상태입니다.
10초·180 BPM·8개 시드에서 본선 진출만 자동 게이트를 통과했습니다. 전체 가사 게이트 전에는 ACE 훅을 최종 믹스에 넣지 않습니다.
28.03초 MP3를 생성했고 최대 CUDA 할당은 4.86GiB였습니다. 이 서버에서 안정적으로 예약 실행할 수 있습니다.
혼합 음원 ASR은 승인 가사를 회수하지 못했습니다. 분리 보컬 기준 순서 단어 회수 0.00%, 누락 40개로 한국어 제작 후보에서 제외합니다.
공식 진입점은 한글을 other로 거부합니다. 152초 한국어 통제도 실패했지만 공식 중국어 예제는 99.23% 중국어로 판별되어 설치·샘플러 문제를 배제했습니다.
DiffRhythm 2 추가 시드 검색은 중단했습니다. 같은 승인 가사로 XL까지 실측한 결과는 위 최신 묶음에서 확인합니다.
26.72초 WAV 생성 완료. HeartMuLa 7.68GiB, HeartCodec 6.20GiB로 순차 실행됐고 가중치 양자화나 codec BF16 저하는 사용하지 않았습니다.
순서 단어 회수 70.00%, 누락 12개입니다. 자동 기준 90%·전 단어 회수·무누락을 충족하지 못했습니다.
-12.62 LUFS, +0.22 dBTP로 너무 크고 피크가 넘습니다. 후처리로 조정 가능하지만 가사 실패를 마스터링으로 통과시키지는 않습니다.
DiffRhythm 2보다 승인 가사를 훨씬 많이 회수했지만 자체 품질 게이트는 실패했습니다. 검토 자산으로만 유지하며 XL 결과와 함께 비교합니다.
ACE-Step 공식 배포 목록에는 현재 한국어 보컬 전용 LoRA가 없습니다. 자체 학습은 가능하지만 이 서버의 12GB VRAM은 공식 최소 16GB보다 작습니다.
12GB 하드웨어는 통과했지만 가사 정확도 82.42%, 순서 단어 회수 70%, 누락 12개로 제작 품질 게이트는 실패했습니다. 직접 청취 자산은 위에 고정했습니다.
12GB 하드웨어는 4.86GiB로 통과했지만 승인 가사 회수율 0%, 누락 40개로 한국어 제작 품질 게이트를 실패했습니다. 추가 시드 검색은 중단합니다.
12GB에서 4B XL을 구동했지만 승인 길이 가사는 25.27%였습니다. 72초 밀도 통제에서 50.55%로 올라 입력 밀도가 주원인임을 확인했습니다.
비교 기준본으로만 사용하며 Suno API나 외부 생성은 호출하지 않습니다. 로컬 장비로 만든 최종 믹스만 이 기준과 비교합니다.
HeartMuLa·DiffRhythm·YuE는 공식 Apache 2.0, ACE-Step 1.5는 MIT입니다. ACE를 Apache로 적은 기존 보고는 정정하며, YuE는 요청된 출처 표기를 manifest에 남깁니다.
연매출 USD 1M 미만 Community License와 상업 등록 조건을 확인합니다. 공식적으로 명료한 보컬을 만들지 못하므로 한국어 가창 후보가 아니라 반주 A/B 후보입니다.
LeVo는 commercial 또는 production 사용을 금지하고, EXAONE은 모델·파생물·출력의 상업 사용을 금지합니다. 연구 결과도 공개 영상 산출물에 섞지 않습니다.
RVC 보컬 모델·훈련 음성·참조 가수와 BSRoformer 커뮤니티 가중치는 별도 권리입니다. 명시적 상업 허용이나 계약이 없으면 해당 자산을 제작에서 거부합니다.
LM 토큰 생성 뒤 모델을 해제하고 HeartCodec을 로드하는 공식 경로가 있습니다. 단일 GPU인 이 서버에서는 필수이며 codec은 음질 때문에 FP32를 유지합니다.
모델 라이선스가 NC이고, 콘텐츠 서버의 승인 가사를 다시 쓰는 단계 자체가 현재 계약을 위반합니다. Studio 소유 메타데이터만 결정론적으로 편성합니다.
BS-RoFormer도 추정 분리라 누설과 아티팩트가 생길 수 있습니다. RVC는 이미 3.7초 슬라이싱을 하며 보편적인 최적 epoch는 없습니다.
HeartMuLa lazy → DiffRhythm 2 → ACE-Step XL → YuE 1회 → 반주 A/B → 허가된 스템·RVC → 믹스·블라인드 검수 순서입니다.
단일 보컬·1구간·권장 음절 4줄에서 전체 믹스 최고 40%, 분리 보컬 60%로 70% 연구 게이트를 넘지 못했습니다.
최소 캡션은 33.33%, LM을 꺼도 33.33%였습니다. 현재 실패를 긴 프롬프트나 LM 하나로 설명할 수 없습니다.
사람 청취로 실제 발음을 확인한 뒤, 정확한 한국어 가창을 별도 확보하고 ACE는 연속 반주·편곡으로 경쟁시킵니다.
한국어는 감지됐지만 승인 가사를 거의 회수하지 못해 품질 관문에서 탈락했습니다. 사람 청취 전 자동 승격도 금지합니다.
24GB RAM 요구 프로필을 16GB 호스트에서 실행했고 CPU 모델·CUDA 입력 경고도 있었습니다. 같은 YuEGP 조합의 seed 반복은 중단합니다.
장치 배치를 확인한 동일 가사 1회만 허용합니다. 공식 경고대로 양자화가 음악성을 낮출 수 있어 실패 시 YuE 계열을 종료합니다.
a reproducible 75-to-90-second one-pass Korean comic song baseline with human-marked pronunciation defects and optional per-job localized repair; this approval freezes the current research method and does not claim Suno parity
four approved 768x1344 local HiDream or FLUX.2 clean background candidates for each qualification scene, scoring at least 90 without characters or text
approved direct HiDream 4MP full-body states scoring at least 90 relative to the user-selected HiDream reference while preserving its premium face finish, identity, and round two-to-three-head comedy proportions
four human-approved deterministic composites per qualification scene using only approved background and character layers
four independent 4-to-6-second vertical image-to-video shots that score at least 90 against Seedance 2.0 normalized to 100
a final-pipeline ten-second scene using only approved music, voices, character references, motion shots, captions, effects, and mastering
대사·가사·순서·운세 내용은 다시 쓰지 않고 승인 체크섬으로 동결했습니다.
음악 연구 기준선은 승인됐지만 제작 키프레임, Seedance급 모션, 10초 통합 시험이 남아 본편은 아직 생성하지 않습니다.
4종 전체 장면 키프레임을 장면당 네 후보로 만들고 캐릭터 정체성, 표정, 구도, 배경 완성도를 먼저 승인합니다.
Turbo+1.7B가 기존 비교에서는 이겼지만 SFT+1.7B, Base+1.7B, Turbo shift 변형과 조합형 제작 비교가 남았습니다.
보컬 분리 후에도 문자 정확도 36.26%, 누락 29단어여서 반주가 아닌 실제 발음과 가사 생성 문제로 확인됐습니다.
직접 ACE 완곡, 보컬 교체형, 정확한 가이드 보컬+ACE 반주형을 같은 가사·음악·사람 청취 게이트로 비교합니다.
가사 줄 재편곡 최고 27.47%, 원본+단순 프롬프트 32.97%, SFT+1.7B 최고 4.40%로 기존 원본 Turbo 39.56%를 넘지 못했습니다.
19개 구간이 28.333~55초에 연속 배치됐고 47개 발화 문구와 순서는 유지됐습니다.
ACE-Step 8개 후보 대신 19개 beat와 19개 Edge TTS chant를 따로 생성해 이어 붙인 가이드입니다.
결정론적 연구 도형이며 ComfyUI 캐릭터·클린 배경·최종 가창 승자는 아직 없습니다.
미술·애니메이션·음향·스토리보드·모바일 시청자·릴리스 QA의 독립 평가를 항목별 중앙값으로 집계합니다.
목표 90점 기준입니다. 현재 우선 보강 항목은 표정 40.0, 음성·노래·음향 35.5, 포즈·동작 49.0, 캐릭터 디자인 53.5입니다.
최근 변화 +4.12점. 캐릭터는 승인 표정·포즈 상태팩이 필요하고, 노래는 Studio가 ACE-Step으로 생성·합성해야 합니다.
점수는 제작 문법과 기술 품질입니다. 캐릭터 매력, 브랜드 고유성, 자연스러운 가창은 자동 통과로 간주하지 않습니다.