Voiceover, music and animation, with no designer, no filming and no After Effects. This is the exact workflow we use, with the templates we copy from every time. Every rule here exists because something went wrong the first time.
Writes the script, plans the scenes and codes the entire animation.
Turns React code into MP4. One file gives you both 9:16 and 1:1.
A natural voice with a timestamp for every single character.
A track made to the exact length and mood, no license hunting.
Speech to text. Checks that the voice actually said what the script says.
Correct loudness, stills for review and the final polish.
The voice is made first, with timestamps. Then every single animation is locked to the word it belongs to. That is why text, visuals and sound land exactly on time, without anyone sitting in an editing app nudging things around.
The order matters. Most people start with the animation and add the voice at the end. We do the opposite, because the voice sets the pace for everything else.
First decide whether this is an ad for cold traffic or a brand film. Nice typography on an empty background stops nobody in the feed.
Why: Lots of short sentences make the AI voice restart its intonation every time. That is what makes it sound choppy, not the pauses.
Why: We tried speeding up the audio file afterwards once. Everyone heard it right away, even though nobody could say what was wrong.
Why: When the music shifts at the same moment the solution arrives, the whole ad feels thought through. The viewer notices without knowing why.
Safe zone in 9:16. All text sits between y=400 and y=1520 on a 1080×1920 video. Reels, Stories and TikTok cover the top and bottom with buttons and text. Backgrounds and images can go all the way to the edge, and text should be 25 to 30 percent larger than in 1:1.
The voice always sits on top. The music is a bed.
Why even: If you duck the music for every sentence, it pumps up in every little pause. Then the voice seems choppy even when it isn't. Finally, the whole video is normalized to -15 LUFS.
Of everything we have tested, this format has gotten the best response. About 45 seconds, built like this:
Not "poor communication", but the chat that floods until the important message scrolls off the screen. Not "time-consuming", but the clock going from 9:47 PM to 11:12 PM.
Problem 3 gets solution 3. The viewer doesn't need to think, just recognize.
Above each solution, the problem sits with a red line through it: a report full of clicks. You see right away what got fixed.
These are the same templates we start from ourselves. You need an ElevenLabs account and a Mac or PC with Node and Python.
Make a motion graphic ad in Remotion for [company].
Audience: [who]. Offer: [what]. Website: [url].
Take every claim from the website, don't make up numbers.
Format: problem and solution, about 45 seconds, 30 fps, 9:16 and 1:1 from the same component.
- 5 concrete problems as small scenes (not icon lists), then a turning point
where the logo draws itself, then 5 solutions in the same order.
Above each solution: the old problem crossed out in red.
- Script: a few connected sentences. Abbreviations with hyphens. Currency spelled out.
- Voice: ElevenLabs eleven_v4, language_code set to the language of the ad
("en" for English, "no" for Norwegian), the with-timestamps endpoint.
Make 4 takes, check each with speech to text against the script,
pick the one with the correct text and the most even pace.
- Music: ElevenLabs Music in two halves (restless minor, then warm major),
shifting at the turning point.
- Build a scene table in frames at the top of the file, locked to the word timings from the voice.
- All text in 9:16 between y=400 and y=1520.
- Audio mix: voice 0.95, music 0.30 lowered evenly to 0.09 under the whole voice.
- Render both formats and normalize to -15 LUFS.
Show me the script and scene table before you generate the voice.{
"model_id": "eleven_v4",
"language_code": "no",
"voice_settings": {
"stability": 0.5,
"similarity_boost": 0.78,
"style": 0.0,
"use_speaker_boost": true,
"speed": 1.0
}
}import json, base64
r = json.load(open("resp.json"))
open("vo.mp3", "wb").write(base64.b64decode(r["audio_base64"]))
a = r.get("alignment") or r["normalized_alignment"]
chars = a["characters"]
starts = a["character_start_times_seconds"]
ends = a["character_end_times_seconds"]
words, cur, t0, prev = [], "", None, 0
for c, s, e in zip(chars, starts, ends):
if c.strip() == "":
if cur:
words.append({"word": cur, "start": round(t0 * 30), "end": round(prev * 30)})
cur, t0 = "", None
else:
if t0 is None:
t0 = s
cur += c
prev = e
if cur:
words.append({"word": cur, "start": round(t0 * 30), "end": round(prev * 30)})
json.dump(words, open("words.json", "w"), ensure_ascii=False, indent=1){
"prompt": "Instrumental bed for a 47 second Scandinavian software ad in two halves. First 20 seconds: sparse, slightly tense and restless, muted plucked strings and a ticking clock-like pulse, minor key, unresolved. Then at 20 seconds a clear warm turn: soft felt piano and gentle brushed percussion, major key, hopeful, calm, building subtly to a clean resolved ending. 96 BPM, no vocals, no drops, understated.",
"music_length_ms": 47000,
"force_instrumental": true
}import { Easing, interpolate } from "remotion";
const clamp = { extrapolateLeft: "clamp", extrapolateRight: "clamp" } as const;
// Visible from frame a to b: 16 frames in, 14 frames out
export const inOut = (f: number, a: number, b: number, fi = 16, fo = 14) =>
Math.min(
interpolate(f, [a, a + fi], [0, 1], { ...clamp, easing: Easing.out(Easing.exp) }),
interpolate(f, [b - fo, b], [1, 0], clamp)
);import json, subprocess, sys
src, dst = sys.argv[1], sys.argv[2]
o = subprocess.run(
["ffmpeg", "-hide_banner", "-i", src, "-af",
"loudnorm=I=-15:TP=-1.5:LRA=11:print_format=json", "-f", "null", "-"],
capture_output=True, text=True).stderr
d = json.loads(o[o.rindex("{"):o.rindex("}") + 1])
af = (f"loudnorm=I=-15:TP=-1.5:LRA=11:measured_I={d['input_i']}:measured_TP={d['input_tp']}"
f":measured_LRA={d['input_lra']}:measured_thresh={d['input_thresh']}"
f":offset={d['target_offset']}:linear=true")
subprocess.run(["ffmpeg", "-y", "-i", src, "-c:v", "copy", "-af", af,
"-ar", "48000", "-c:a", "aac", "-b:a", "192k", dst], check=True)Already built the whole animation and want to test a different voice? Match the words in the old and new versions, save two lists of frame numbers (OLD and NEW), and let the composition translate every frame to "old time". The whole animation then shifts word by word to follow the new voice. This is how we test several voices on the same video in a few minutes.
import { interpolate, useCurrentFrame } from "remotion";
const extend = { extrapolateLeft: "extend", extrapolateRight: "extend" } as const;
const toOld = (f: number) => interpolate(f, NEW, OLD, extend);
// Use this instead of useCurrentFrame() throughout the composition
export const useWarpedFrame = () => toOld(useCurrentFrame());0 of 7 checked

Hi, Victor here. This is the same recipe we use when we make motion ads for our clients. It gets updated every time we learn something new, so save the link.
Would you rather have us make the ads for you? Send us your website at the bottom of the page, and you'll get a free proposal.

The numbers come from our clients’ own ad accounts.
See also: Google Ads agency · Meta Ads agency · Performance marketing agency · What is UGC? · Google Ads agency Oslo
Send us your website and we’ll put together a concrete proposal for ads and a landing page for your business, showing how it would look with your brand. You’ll get it by email within a couple of days, and it’s yours to keep either way.

Free · no lock-in contract · we never share your info
Rather talk directly? Book 30 minutes in the calendar → · victor@arcads.no