Updated 02:20Z. Expression warmed per your note, and both closer arms are regenerated from it — scripts exact, handles in range, defects I can see listed under them. Questions 1 and 3 below are now settled; question 2 is superseded by the new takes.
Two clips need a human listen before they can pass the source gate. Nothing here is approved yet, and nothing has been rendered against either one.
The RAW gates say speech-to-text can establish which words were said and nothing else — it "normalizes mispronunciations to real words" and cannot certify delivery, pronunciation or lip sync. Six of the eight open gate errors are fields that require ears and eyes. I left them unanswered rather than assert them.
Headphones if you have them. Question 1 is thirty seconds of work; question 2 needs one proper watch.
Both returned their script verbatim and both clear every handle measurement — the first time that has happened. I am not calling them good. You have overruled me on "looks AI" once tonight and you were right, so these go to you unjudged.
Exact. Speech ends 5.94 s · handle 0.56 s · tail 0.54 s.
Exact. Speech ends 5.13 s · handle 1.37 s · tail 0.54 s.
Also resolved: question 3 is moot. Both takes measure
handleAfterMs at 558 ms and 1372 ms, inside the schema's 120–2000 band, so the
ambiguity that blocked the last pair does not arise. It was an artefact of that take's long tail,
not a real schema problem.
You were right: I was generating a starter, but the wrong kind, with the wrong model. Rebuilt against the ad you named. Nothing has been animated from it yet — I wanted you to see the frame before I spend anything on motion.
Script mark crisp, fine print soft and unresolvable, cap on, lips closed, individual stubble hairs, flat window light. Both defects that killed the earlier attempts are gone.
GPT Image 2. Resort pool, golden-hour backlight, open shirt, male-model build. The method doc's exact warned failure — I described an advertisement instead of a capture, and picked the model because this box has an OpenAI key and no Gemini one.
Room, light, wardrobe, ordinary face.
Bottle handling only — up by the cheek, close to the lens. Scene deliberately not used.
save70 arm, iteration 6. The timing on this take is finally correct — it's the only thing standing between it and the gate.
"Barely" and "Raw" are timestamped 0.16 s apart. That is very tight for two separate words, which is the main reason I think this may be an artifact rather than a real insertion — but I can't hear it, so I won't call it.
buy-2-get-1 arm. Words, timing and mouth all measured clean. What's left is judgement.
Muted is its own pass in the gate — does the face and body carry the intent with no sound at all?
Exact match to the script. No added, repeated or truncated words — the only one of six generations to manage it.
competitor-synthesis/field-test-opener.mp4 — the approved human-performance
baseline. The rule is that word accuracy and lip sync cannot make up for delivery that is
flatter, more segmented or less emotionally legible than this.
This is the ending currently running on 3S-V5, GAME and LOOK-TWICE. It says "raw's half off and the roll-on's free" — a rate that matches neither live tier, spoken out loud by a visible actor. This is what we're replacing.
mechanical.handleAfterMs in the source-audit schema accepts 120–2000 ms.
On the b2g1 take it measures 2482 ms, because the 6.5 s edit boundary sits
2.5 s after his last word.
| Item | State |
|---|---|
| UFO-V34 — 4 masters, both arms, both ratios | Delivered & live |
| b2g1 closer — words, timing, mouth, handles | Measured clean |
| b2g1 closer — the six perceptual calls | Question 2 |
| save70 closer — timing, handle, tail | 0.54 s handle, 0.54 s tail |
| save70 closer — the extra word | Question 1 |
handleAfterMs semantics |
Question 3 |
| 3S-V5 / GAME / LOOK-TWICE — 12 masters | Blocked on the above |
| Meta | Untouched — the four wrong-offer ads are still live |
Voice match is already built: the b2g1
take is cloned as a Higgsfield voice element, so once question 1 is settled I can force the
save70 arm onto the identical voice with voice_change, which preserves timing and
lip sync. b2g1 keeps its own native performance.