The four-beat structure (and the timestamps that matter)
Almost every UGC video that holds attention on TikTok follows the same skeleton. Not because it is a gimmick, but because it matches how the feed actually works: the first frame decides the swipe, and the last few seconds decide whether anyone taps the product link.
The four beats are Hook → Problem → Proof → CTA. What separates a video that scrolls past from one that sells is not the script length, it is where each beat lands on the clock. A 30-second video with a hook that takes 6 seconds to arrive has already lost most viewers.
| Beat | Timestamp | Job | Common mistake |
|---|---|---|---|
| Hook | 0–3s | Stop the scroll and name who this is for | Opening with a logo or brand intro |
| Problem | 3–8s | Make the viewer feel the pain or desire | Listing features instead of the felt problem |
| Proof | 8–22s | Show the product working, not just talking | Telling instead of demonstrating on camera |
| CTA | 22–30s | One clear next action, low friction | Three asks at once, or no ask at all |
Use these as ceilings, not targets. A strong hook can land in 1.5 seconds. The point is that by second 8 the viewer should already know who you are talking to, why it matters, and that proof is coming. Think of the timeline as a promise: every second you spend before the proof is a second the viewer is deciding whether you are worth their time. The structure works because it front-loads the payoff signal — the viewer can tell, almost immediately, that something useful is coming.
One nuance most templates skip: the four beats are not equal in weight. The hook carries the majority of your retention risk, the proof carries the majority of your conversion risk, and the problem and CTA are connective tissue. If you only have time to perfect two beats, perfect the hook and the proof.
Copy-paste the template
Below is the full fill-in-the-blank template (in the box further down the page). The bracketed parts are the only things you change. Keep the spoken word count near 2.2 words per second of runtime — a 30-second video is roughly 65–70 spoken words — so you are not rushing or padding.
Two rules make a bigger difference than anything else:
- Say the hook with your mouth. Do not rely on the on-screen caption alone. The feed is a sound-on environment for a large share of viewers, and even the ones watching muted still read your face and your captions — but the spoken hook is what carries energy and intent. A silent caption over a slow pan is the single most common reason a UGC video dies in the first second.
- Get the product physically in frame before second 10. Talking about a product the viewer cannot see yet feels like an ad. Handling it on camera feels like a recommendation. The faster the object appears, the faster the brain reclassifies the video from "someone selling me something" to "someone showing me something."
A practical workflow: write the body and CTA first, because they are the stable parts. Then write five hooks against that fixed body. Counterintuitively, writing the hook last produces sharper hooks — once you know exactly what the proof delivers, you can promise it precisely instead of vaguely. Finally, read the whole thing out loud and time it. If you run over the word budget, cut the problem beat before you cut the proof.
RUNTIME: ~30s | SPOKEN BUDGET: ~65 words | PRODUCT: [product name + price] [0–3s] HOOK (say it out loud) "[Call out WHO this is for] + [the result or pain]." -> e.g. "If you [problem], this is the [product] that [outcome]." -> Pick a formula: Call-out / POV / Confession / Contrarian / Result-first / Question / Curiosity-gap. -> Write 3–5 hook variants. Test these, hold everything below constant. [3–8s] PROBLEM (make them feel it) "[The honest, specific frustration the viewer already has]." -> First person beats lecturing: "I used to ___ and it never ___." -> Felt pain here, NOT a feature list. [8–22s] PROOF (show, don't tell — product on screen) "[Demonstrate the product working] — watch." -> Use observable, honest language: looks / feels / noticed. -> Add a pattern interrupt (cut, zoom, new angle) so it doesn't flatten. -> AVOID: cures, heals, treats, clinically proven (unless the brand can substantiate it). [22–30s] CTA (one ask, low friction) "[One next action]. [One tiny tip to start]." -> e.g. "It's linked below — start with [easiest step]." --- QA before recording: [ ] Hook names the viewer/problem in <3s and is spoken aloud [ ] First frame works as a thumbnail (face / motion / product — not a logo card) [ ] Product physically on screen before 0:10 [ ] Main claim is observable + something the brand can substantiate [ ] Exactly ONE CTA [ ] Spoken script fits the word budget (~2.2 words/sec) [ ] 3–5 hook variants written for this same body
Hook formulas: seven openers you can steal
"Write a strong hook" is not advice, it is a wish. What actually helps is a finite menu of opener shapes you can fill in. Here are seven that hold up across DTC categories, with an illustrative fill-in for each. Treat them as starting points, not scripts — the goal is to have five different angles ready to test, not one you are married to.
| Formula | Shape | Illustrative example |
|---|---|---|
| Call-out | "If you [specific situation], watch this." | "If your skincare shelf is chaos and your skin still looks dull, watch this." |
| POV | "POV: you [relatable scenario]." | "POV: you have a car, a couch, and a dog that sheds like it's a full-time job." |
| Confession | "I was today years old when I realized [thing]." | "I was today years old when I realized I'd been washing my face wrong for a decade." |
| Contrarian | "Stop [common behavior]. Do this instead." | "Stop buying a new serum every month. Here's what actually changed my routine." |
| Result-first | "This is the [product] that [observable outcome]." | "This is the one serum I actually finished a whole bottle of." |
| Question | "Why does no one talk about [problem]?" | "Why does no one talk about how heavy regular vacuums are for your car?" |
| Curiosity gap | "Here's what happened when I [action] for [timeframe]." | "Here's what happened when I used this every morning for two weeks." |
The common thread: every one of these names a specific person or situation. "Check out this amazing product" addresses nobody. "If your skincare shelf is chaos" addresses someone, and that someone stops scrolling because the video just described their bathroom. Specificity is the mechanism — vague hooks fail not because they are boring but because they are aimed at everyone, which means they land on no one.
Two worked examples
Here is the template filled in for two different DTC categories so you can see how the blanks become real lines. Both are illustrative, not pulled from any specific campaign.
Example 1 — a $34 vitamin-C serum (skincare):
Hook: "If your skincare shelf is chaos and your skin still looks dull, watch this."
Problem: "I had six half-used bottles and zero routine. My skin looked tired by 3pm."
Proof: "This is the one serum I actually finished. Two weeks in, my skin looks brighter and a few people have noticed — here's the texture, here's how it sinks in."
CTA: "It's linked below if you want to try it. Start with the morning step."
Notice the proof line says "looks brighter" and "people noticed" — observable, honest, the kind of claim a brand can stand behind. It deliberately avoids "clears acne" or "clinically proven," which are claims a brand would have to substantiate and which can get an ad rejected.
Example 2 — a $59 cordless handheld vacuum (home / DTC gadget):
Hook: "POV: you have a car, a couch, and a dog that sheds like it's a full-time job."
Problem: "The big vacuum is too heavy for the car, and the little dustbuster died in a month."
Proof: "This one charges on USB-C and pulls dog hair out of the seat seam in one pass — watch."
CTA: "Link's in the bio. Grab the one with the crevice tool."
The proof beat in both cases is a demonstration on camera, not an adjective. That single habit — show, do not describe — is the highest-leverage edit you can make to any UGC script. The serum proof shows the texture and the absorption; the vacuum proof shows the one-pass result on a visibly furry seat. In both cases the viewer's eyes do the believing, not the script's promises. Adjectives are free, which is exactly why they don't persuade anyone — a demonstration costs the creator the risk of it not working on camera, and that risk is what makes it credible.
The pre-shoot checklist
Before you (or your creator) hit record, run the script against this list. If any answer is no, fix the script, not the footage. The most expensive mistake in UGC is shooting a bad script beautifully — great lighting cannot rescue a hook that arrives at second six.
- Does the hook name the target viewer or their problem in the first 3 seconds?
- Is the hook spoken out loud, not just shown as a caption?
- Is the product physically on screen before second 10?
- Is the main claim something the brand can honestly back up (an observable result, not a medical or "clinically proven" claim)?
- Is there exactly one CTA, and is it the easiest possible action ("link below," "comment a word")?
- Is the spoken script within the word budget for the runtime (~2.2 words/sec)?
- Did you write at least 3–5 hook variants for the same body, so you can test which one holds?
- Does the first frame work as a thumbnail — is there a face, motion, or visible product, rather than a black frame or a logo card?
- Is there a pattern interrupt somewhere in the 8–15s window (a cut, a zoom, a new angle) so the proof beat doesn't flatten into a monologue?
That seventh point is the one most brands skip. The hook is responsible for the majority of your retention, so it is the variable worth iterating on first. Keep the proof and CTA constant, swap the hook, and let the data tell you which opener earns the watch. Points eight and nine are about the visual layer most scripts ignore — the feed is silent-first for many viewers and motion-driven for all of them, so a script that reads well on paper can still die because nothing moves in the opening frame.
How to test hooks without guessing
One script is not a campaign. The fastest way to find what converts is to treat hooks like the test variable and everything else as the control.
- Hold the body and CTA constant. Change only the first 3 seconds across 3–5 cuts of the same video. If you change the hook and the proof at the same time, a win tells you nothing, because you can't attribute it.
- Watch 3-second view rate and average watch time, not just likes. A hook that wins on retention is the one worth scaling. Likes measure whether people who already watched approved; retention measures whether people kept watching at all, which is the thing the algorithm rewards and the thing that precedes a sale.
- You run the ads and pick the winners. When you stop spending on a tired hook, the value of having a production engine is that you simply order more variations like the one that's working — the engine produces more of the proven shape instead of you starting from a blank page. Sunk-cost loyalty to a clever hook you wrote is the most common reason testing budgets evaporate.
- Double down on the winning angle. Once a hook proves out, produce five more variations of that angle rather than starting from scratch. Winners cluster — if a confession-style opener works, other confessions in the same voice tend to work too.
This is a general pattern, not a guarantee — different products and audiences behave differently. But the discipline of testing the hook in isolation is what turns a decent template into a repeatable system, and it is why volume matters: you cannot find your one winning hook out of a sample size of three videos. A common pattern across DTC teams is that the hit rate on creative is low enough that the only reliable strategy is to produce many disciplined variations and let a small fraction carry the account.
Writing compliant claims for skincare, supplements, and pets
This is where good UGC scripts most often get rejected, demonetized, or — worse — get a brand into regulatory trouble. The fix is a translation habit: take every claim and rewrite it as something observable and first-person rather than medical and absolute.
| Don't say (claim you'd have to prove) | Say instead (observable, honest) |
|---|---|
| "Clears acne" | "My skin looks calmer and a few people noticed" |
| "Clinically proven to brighten" | "Two weeks in, my skin looks brighter to me" |
| "Cures bloating" | "I feel less bloated after dinner" |
| "Boosts immunity" | "It's part of my morning routine" |
| "Heals your dog's joints" | "My older dog seems more eager on walks" |
| "Treats anxiety" | "I reach for it on stressful days" |
The pattern on the right side is consistent: it describes what the speaker personally observed or felt, hedged honestly ("seems," "to me," "a few people"), and it never states a health outcome as a fact. "Clinically proven" is a factual claim — if the brand says it, the brand must be able to substantiate it, and a creator saying it on the brand's behalf carries the same weight.
The honest version isn't just safer, it usually performs better. "Cures acne" triggers skepticism because viewers have heard it a thousand times; "my skin looks calmer and a few people noticed" sounds like a real person, which is the entire point of UGC. Let the brand's own claims policy and legal review gate anything stronger — the script should never be the place where an unsubstantiated claim sneaks in.
What most brands get wrong
After enough scripts, the failure modes rhyme. Here are the patterns that quietly kill UGC performance, and the fix for each.
- The hook is a caption, not a sentence. A pretty text overlay over a slow B-roll pan is not a hook — it's a title card. Fix: say it, with energy, in the first second.
- The problem beat is a feature list. "It has 20% vitamin C and hyaluronic acid" is not a problem. "My skin looked tired by 3pm" is. Features go in the proof, felt pain goes in the problem.
- Proof is told, not shown. "It really works" proves nothing. The proof beat must contain a visible demonstration — texture, before/after framing, a single-pass result.
- Three CTAs at once. "Follow, like, comment, and check the link" splits attention to zero. One ask. The easiest possible one.
- Production value over hook quality. Brands spend on cameras and lighting and skimp on the 30 minutes of hook-writing that actually moves retention. Reverse the priority.
- Testing too few variations. Shipping one or two videos a month and concluding "UGC doesn't work for us." The hit rate on creative is low; the math only works at volume.
- Falling in love with the clever hook. The hook you're proudest of is often not the winner. Let retention data, not your taste, decide.
Notice that almost none of these are production problems — they're scripting and process problems. That's good news, because scripting and process are cheap to fix. You can rewrite a hook for free; you cannot un-shoot a bad one cheaply.
How this scales: doing it yourself vs. running it at volume
The template above is genuinely all you need to write good scripts. The honest question isn't "is this template enough?" — it is, the thinking is the same thinking a studio uses — the question is whether you can produce enough disciplined variations of it to actually find winners.
Here's the math most brands run into. If a winning hook surfaces only once every several attempts, then a handful of videos a month is a coin flip, not a system. Finding reliable winners means producing variations at a cadence that's hard to sustain by hand while also running the rest of the business.
| Human creator, in-house | AI-native production at volume | |
|---|---|---|
| Cost structure | ~$200–$600+ all-in per video (~$150 base before product, shipping, revisions) | A flat monthly engine — 30 to 180+ creatives a month depending on the tier |
| Typical monthly throughput | A handful | 30 creatives on Starter ($3,000/mo), 90 on Growth ($7,500/mo), 180+ on Scale |
| Time to first creatives | Days to weeks (shipping, scheduling, shooting) | On Growth, your first 20 land within 72 hours of brief approval |
| Iteration speed on hooks | Slow — re-shoot required | Fast — new variations from the same brand context |
Neither path is automatically right. If you have a strong in-house creator whose face your audience already trusts, lean into that — authenticity from a known person is hard to beat. The volume path makes sense when you need to test many angles quickly, when re-shooting is the bottleneck, or when the per-video economics of human production make disciplined testing too expensive to sustain. The template is the same either way; what changes is how many at-bats you can afford. And the honest caveat: volume only helps if the variations are disciplined — fifty random videos lose to five tested ones built on a structure like the one above.



