RAW closer — ear-check adjudication

Updated 02:20Z. Expression warmed per your note, and both closer arms are regenerated from it — scripts exact, handles in range, defects I can see listed under them. Questions 1 and 3 below are now settled; question 2 is superseded by the new takes.

Two clips need a human listen before they can pass the source gate. Nothing here is approved yet, and nothing has been rendered against either one.

Why you're being asked

The RAW gates say speech-to-text can establish which words were said and nothing else — it "normalizes mispronunciations to real words" and cannot certify delivery, pronunciation or lip sync. Six of the eight open gate errors are fields that require ears and eyes. I left them unanswered rather than assert them.

Headphones if you have them. Question 1 is thirty seconds of work; question 2 needs one proper watch.

New — both closer arms, from the approved starter

Both returned their script verbatim and both clear every handle measurement — the first time that has happened. I am not calling them good. You have overruled me on "looks AI" once tonight and you were right, so these go to you unjudged.

70% off
By the way, raw is 70 % off today, that's 3 bottles. Grab it below.

Exact. Speech ends 5.94 s · handle 0.56 s · tail 0.54 s.

Buy 2 get 1
By the way, buy two bottles and the third one's free. Grab it below.

Exact. Speech ends 5.13 s · handle 1.37 s · tail 0.54 s.

What I can see wrong, so you are not looking for it blind:
  • The bottle moves in the back half of the b2g1 take — there is motion blur on it around the five-second mark. I asked for it to stay put and it did not.
  • The framing drifts — he ends up closer and lower than he starts. "Locked propped-phone framing" is not holding.
  • The head movement is bigger than asked for, with a lurch early on, and one or two frames where the eyes squint oddly and the mouth reads gurny.
None of those are things a transcript or a handle measurement can catch. Whether they add up to "looks AI" again is your call, not mine.

Also resolved: question 3 is moot. Both takes measure handleAfterMs at 558 ms and 1372 ms, inside the schema's 120–2000 band, so the ambiguity that blocked the last pair does not arise. It was an artefact of that take's long tail, not a real schema problem.

New — the rebuilt starter

You were right: I was generating a starter, but the wrong kind, with the wrong model. Rebuilt against the ad you named. Nothing has been animated from it yet — I wanted you to see the frame before I spend anything on motion.

The rebuild — Nano Banana Pro, gate-passed
rebuilt starter

Script mark crisp, fine print soft and unresolvable, cap on, lips closed, individual stubble hairs, flat window light. Both defects that killed the earlier attempts are gone.

What I had before — rejected
rejected poolside starter

GPT Image 2. Resort pool, golden-hour backlight, open shirt, male-model build. The method doc's exact warned failure — I described an advertisement instead of a capture, and picked the model because this box has an OpenAI key and no Gemini one.

The two references it was built from
closet reference

Room, light, wardrobe, ordinary face.

car reference

Bottle handling only — up by the cheek, close to the lens. Scene deliberately not used.

1 · Does he say an extra word?

save70 arm, iteration 6. The timing on this take is finally correct — it's the only thing standing between it and the gate.

The question: at about 1.2 seconds, does he say "Barely" before "Raw"? Or is that speech-to-text mis-hearing a breathy "…way, RAW is…"?

If it's a real inserted word, the take is dead — an added word is a fatal source error. If it's an STT artifact, this take passes and I can move straight to rendering.
Just the moment in question — 1.3 seconds, loop it
The whole line, audio only
What speech-to-text returned
By the way, Barely Raw is 70 % off today. That's three bottles. Grab it below.
Word timings
0.00 By0.60 the0.72 way, 1.22 Barely1.38 Raw 1.56 is1.80 702.10 %2.38 off 2.60 today.3.26 That's3.42 three 3.62 bottles.5.14 Grab5.58 it5.72 below.

"Barely" and "Raw" are timestamped 0.16 s apart. That is very tight for two separate words, which is the main reason I think this may be an artifact rather than a real insertion — but I can't hear it, so I won't call it.

With picture, if the mouth helps you decide

2 · Is this performance good enough to ship?

buy-2-get-1 arm. Words, timing and mouth all measured clean. What's left is judgement.

The take

Muted is its own pass in the gate — does the face and body carry the intent with no sound at all?

Audio only
By the way, buy two bottles and the third one's free. Grab it below.

Exact match to the script. No added, repeated or truncated words — the only one of six generations to manage it.

Six calls, yes or no each:
  • Audio alone — does it sound like a human, not a read?
  • Muted — does the performance still convey the intent?
  • Do the audio and the picture want the same thing?
  • Any robotic or teleprompter cadence?
  • Is it materially worse than the golden reference below?
  • Compared directly against that reference — does it hold up?
The golden reference the gate measures against

competitor-synthesis/field-test-opener.mp4 — the approved human-performance baseline. The rule is that word accuracy and lip sync cannot make up for delivery that is flatter, more segmented or less emotionally legible than this.

For contrast — the closer that is on air right now

This is the ending currently running on 3S-V5, GAME and LOOK-TWICE. It says "raw's half off and the roll-on's free" — a rate that matches neither live tier, spoken out loud by a visible actor. This is what we're replacing.

3 · One schema question

mechanical.handleAfterMs in the source-audit schema accepts 120–2000 ms. On the b2g1 take it measures 2482 ms, because the 6.5 s edit boundary sits 2.5 s after his last word.

Does that field mean the untouched tail left in the source after the cut — which is 540 ms and passes — or the clean handle retained after the last word before the cut, which is 2482 ms and fails?

I didn't guess, because guessing here means either fabricating a gate pass or throwing away a good take for no reason.

Where everything stands

ItemState
UFO-V34 — 4 masters, both arms, both ratios Delivered & live
b2g1 closer — words, timing, mouth, handles Measured clean
b2g1 closer — the six perceptual calls Question 2
save70 closer — timing, handle, tail 0.54 s handle, 0.54 s tail
save70 closer — the extra word Question 1
handleAfterMs semantics Question 3
3S-V5 / GAME / LOOK-TWICE — 12 masters Blocked on the above
MetaUntouched — the four wrong-offer ads are still live

Voice match is already built: the b2g1 take is cloned as a Higgsfield voice element, so once question 1 is settled I can force the save70 arm onto the identical voice with voice_change, which preserves timing and lip sync. b2g1 keeps its own native performance.