★★★★★ Recommended by our clients · 100% of our clients grew their bottom line · 100% performance-based
Free guide · The full recipe

How we make motion ads with AI.

Voiceover, music and animation, with no designer, no filming and no After Effects. This is the exact workflow we use, with the templates we copy from every time. Every rule here exists because something went wrong the first time.

By Victor Skretteberg, Arcads September 29, 2026 About 12 minutes
8steps from idea to finished video
6templates you can copy
10mistakes we made, so you don't have to
The old way
Yourreportisallclicksandviews.
Views and clicks, week by week
A report full of clicks and views
Onepage.Onlythenumbersthatmatter.
New leads
Cost per lead
Sales from the ads
Ads that pay for themselves.
arcads.no
f 000 / 360
Everything in the video is a function of the frame number, just like in Remotion. Drag the slider to see every frame.
The tools

Six tools. None of them require you to know design.

Claude Code
Director and developer

Writes the script, plans the scenes and codes the entire animation.

Remotion
Video from code

Turns React code into MP4. One file gives you both 9:16 and 1:1.

ElevenLabs speech
Voiceover

A natural voice with a timestamp for every single character.

ElevenLabs Music
Music

A track made to the exact length and mood, no license hunting.

ElevenLabs Scribe
Quality check

Speech to text. Checks that the voice actually said what the script says.

ffmpeg
Audio and checks

Correct loudness, stills for review and the final polish.

The main trick

The voice is made first, with timestamps. Then every single animation is locked to the word it belongs to. That is why text, visuals and sound land exactly on time, without anyone sitting in an editing app nudging things around.

Part 1

The workflow in eight steps

The order matters. Most people start with the animation and add the voice at the end. We do the opposite, because the voice sets the pace for everything else.

01

Brief and angle

First decide whether this is an ad for cold traffic or a brand film. Nice typography on an empty background stops nobody in the feed.

  • An ad for cold traffic needs something concrete in the first 1.5 seconds: a real product, a real screenshot, a situation people recognize, or a number.
  • Take every claim from the company's own website. Don't make up numbers.
  • Use real screenshots and product photos. Drawn mockups look like they were made by a script.
  • Never show a person who isn't the one speaking.
02

Script

  • Write a few long sentences tied together with commas, "and", "because" and "so".
  • Generate the whole script in one call. Never sentence by sentence.
  • Write abbreviations with hyphens, like "C-R-M". Otherwise they are often pronounced as one word.
  • Spell out currency and symbols: "200 dollars", not "$200".
  • About 20 seconds of speech for a regular ad, 40 to 45 seconds for problem and solution.

Why: Lots of short sentences make the AI voice restart its intonation every time. That is what makes it sound choppy, not the pauses.

03

Voice

  • Model eleven_v4 with a language code that matches the ad, like "en" or "no". The older multilingual model drifts into Danish on Norwegian text.
  • Use the with-timestamps endpoint. Then you know exactly when each word starts.
  • Make 3 to 6 takes and pick by measurement: run speech to text on each, check that every word is right, and pick the one with the most even pace.
  • Too slow? Raise speed to 1.05 to 1.15, or shorten the script.
  • Always listen to the voice yourself before you publish.

Why: We tried speeding up the audio file afterwards once. Everyone heard it right away, even though nobody could say what was wrong.

04

Music

  • Generate the music at exactly the same length as the video.
  • If the ad has a turning point, ask for two parts: restless and in a minor key before, warm and in a major key after.
  • Avoid generic "inspiring corporate music". It's the first thing people scroll past.

Why: When the music shifts at the same moment the solution arrives, the whole ad feels thought through. The viewer notices without knowing why.

05

Storyboard and timing

  • Make a scene table in frames (30 per second) where each scene starts on a specific word in the voice.
  • Something new should happen at least every 1.5 to 3 seconds.
  • Start with a lead of about 0.4 seconds before the voice.
06

Animation in Remotion

  • One component per ad. The same code gives you 9:16 and 1:1.
  • Entrances 15 to 20 frames, exits shorter, 10 to 15 frames.
  • Headlines come in word by word, in time with the voice.
  • One big visual event beats ten small icons in the corner.
  • Calm and Nordic: space, thin lines, muted colors. No confetti, spinning stars or big SALE badges.

Safe zone in 9:16. All text sits between y=400 and y=1520 on a 1080×1920 video. Reels, Stories and TikTok cover the top and bottom with buttons and text. Backgrounds and images can go all the way to the edge, and text should be 25 to 30 percent larger than in 1:1.

07

Audio mix

The voice always sits on top. The music is a bed.

Voice
0.95
Music aloneno voice
0.55
Music, starting pointin a video with voice
0.30
Music on the end card
0.16
Music under the voiceeven, the whole way
0.09

Why even: If you duck the music for every sentence, it pumps up in every little pause. Then the voice seems choppy even when it isn't. Finally, the whole video is normalized to -15 LUFS.

08

Render and review

  • Render 9:16 and 1:1 from the same component.
  • Pull stills at key moments and check that the text sits inside the safe zone.
  • Go through the checklist before anything is published.
Part 2

The format that works best: problem and solution, mirrored

Of everything we have tested, this format has gotten the best response. About 45 seconds, built like this:

1

Scenes, not icons

Not "poor communication", but the chat that floods until the important message scrolls off the screen. Not "time-consuming", but the clock going from 9:47 PM to 11:12 PM.

2

Mirrored 1:1

Problem 3 gets solution 3. The viewer doesn't need to think, just recognize.

3

The old way crossed out

Above each solution, the problem sits with a red line through it: a report full of clicks. You see right away what got fixed.

Part 3

The templates. Copy, paste, replace what's in brackets.

These are the same templates we start from ourselves. You need an ElevenLabs account and a Mac or PC with Node and Python.

Prompt for Claude CodeStart here. Claude uses the rest of the templates on its own along the way.
Make a motion graphic ad in Remotion for [company].
Audience: [who]. Offer: [what]. Website: [url].
Take every claim from the website, don't make up numbers.

Format: problem and solution, about 45 seconds, 30 fps, 9:16 and 1:1 from the same component.
- 5 concrete problems as small scenes (not icon lists), then a turning point
  where the logo draws itself, then 5 solutions in the same order.
  Above each solution: the old problem crossed out in red.
- Script: a few connected sentences. Abbreviations with hyphens. Currency spelled out.
- Voice: ElevenLabs eleven_v4, language_code set to the language of the ad
  ("en" for English, "no" for Norwegian), the with-timestamps endpoint.
  Make 4 takes, check each with speech to text against the script,
  pick the one with the correct text and the most even pace.
- Music: ElevenLabs Music in two halves (restless minor, then warm major),
  shifting at the turning point.
- Build a scene table in frames at the top of the file, locked to the word timings from the voice.
- All text in 9:16 between y=400 and y=1520.
- Audio mix: voice 0.95, music 0.30 lowered evenly to 0.09 under the whole voice.
- Render both formats and normalize to -15 LUFS.

Show me the script and scene table before you generate the voice.
Voice settingsPOST /v1/text-to-speech/{voice_id}/with-timestamps
{
  "model_id": "eleven_v4",
  "language_code": "no",
  "voice_settings": {
    "stability": 0.5,
    "similarity_boost": 0.78,
    "style": 0.0,
    "use_speaker_boost": true,
    "speed": 1.0
  }
}
From timestamps to framesPython. Outputs words.json with the start and end of each word.
import json, base64

r = json.load(open("resp.json"))
open("vo.mp3", "wb").write(base64.b64decode(r["audio_base64"]))

a = r.get("alignment") or r["normalized_alignment"]
chars = a["characters"]
starts = a["character_start_times_seconds"]
ends = a["character_end_times_seconds"]

words, cur, t0, prev = [], "", None, 0
for c, s, e in zip(chars, starts, ends):
    if c.strip() == "":
        if cur:
            words.append({"word": cur, "start": round(t0 * 30), "end": round(prev * 30)})
        cur, t0 = "", None
    else:
        if t0 is None:
            t0 = s
        cur += c
    prev = e
if cur:
    words.append({"word": cur, "start": round(t0 * 30), "end": round(prev * 30)})

json.dump(words, open("words.json", "w"), ensure_ascii=False, indent=1)
Music promptElevenLabs Music. Change the length and the timing of the turning point.
{
  "prompt": "Instrumental bed for a 47 second Scandinavian software ad in two halves. First 20 seconds: sparse, slightly tense and restless, muted plucked strings and a ticking clock-like pulse, minor key, unresolved. Then at 20 seconds a clear warm turn: soft felt piano and gentle brushed percussion, major key, hopeful, calm, building subtly to a clean resolved ending. 96 BPM, no vocals, no drops, understated.",
  "music_length_ms": 47000,
  "force_instrumental": true
}
Fade in and outRemotion. Used on almost every element.
import { Easing, interpolate } from "remotion";

const clamp = { extrapolateLeft: "clamp", extrapolateRight: "clamp" } as const;

// Visible from frame a to b: 16 frames in, 14 frames out
export const inOut = (f: number, a: number, b: number, fi = 16, fo = 14) =>
  Math.min(
    interpolate(f, [a, a + fi], [0, 1], { ...clamp, easing: Easing.out(Easing.exp) }),
    interpolate(f, [b - fo, b], [1, 0], clamp)
  );
Loudness to -15 LUFSPython and ffmpeg, two passes. Run: python3 ln.py in.mp4 out.mp4
import json, subprocess, sys

src, dst = sys.argv[1], sys.argv[2]
o = subprocess.run(
    ["ffmpeg", "-hide_banner", "-i", src, "-af",
     "loudnorm=I=-15:TP=-1.5:LRA=11:print_format=json", "-f", "null", "-"],
    capture_output=True, text=True).stderr
d = json.loads(o[o.rindex("{"):o.rindex("}") + 1])
af = (f"loudnorm=I=-15:TP=-1.5:LRA=11:measured_I={d['input_i']}:measured_TP={d['input_tp']}"
      f":measured_LRA={d['input_lra']}:measured_thresh={d['input_thresh']}"
      f":offset={d['target_offset']}:linear=true")
subprocess.run(["ffmpeg", "-y", "-i", src, "-c:v", "copy", "-af", af,
                "-ar", "48000", "-c:a", "aac", "-b:a", "192k", dst], check=True)
+

Bonus: swap the voice without rebuilding the animation

Already built the whole animation and want to test a different voice? Match the words in the old and new versions, save two lists of frame numbers (OLD and NEW), and let the composition translate every frame to "old time". The whole animation then shifts word by word to follow the new voice. This is how we test several voices on the same video in a few minutes.

Time warpRemotion. OLD and NEW are equal-length lists of word start times in frames.
import { interpolate, useCurrentFrame } from "remotion";

const extend = { extrapolateLeft: "extend", extrapolateRight: "extend" } as const;
const toOld = (f: number) => interpolate(f, NEW, OLD, extend);

// Use this instead of useCurrentFrame() throughout the composition
export const useWarpedFrame = () => toOld(useCurrentFrame());
Part 4

Checklist before you publish

0 of 7 checked

Part 5

10 mistakes we made, so you don't have to

01Text and audio were out of syncTimestamps from the voice, never estimated timings
02The voice sounded choppyLong, flowing sentences and even music ducking
03The voice was sped up afterwardsThe speed setting, or a shorter script
04Norwegian was pronounced like Danisheleven_v4 with the right language code
05The abbreviation was pronounced as one wordHyphens between the letters
06The music was too loud0.30 as a starting point, 0.09 under the voice
07Drawn animation got rejectedReal screenshots and product photos as the base
08The text ended up under the Reels buttonsSafe zone y=400 to 1520
09The voice was made sentence by sentenceThe whole script in one call
10A pretty typography intro stopped nobodySomething concrete in the first 1.5 seconds
Victor Skretteberg

Hi, Victor here. This is the same recipe we use when we make motion ads for our clients. It gets updated every time we learn something new, so save the link.

Would you rather have us make the ads for you? Send us your website at the bottom of the page, and you'll get a free proposal.

Some of the companies we work with
Utvendigrenhold Veldigrent.no Rask Flytting Mokki Badstugutta Frifor.app
Before / after

Before and after Arcads.

The numbers come from our clients’ own ad accounts.

See also: Google Ads agency · Meta Ads agency · Performance marketing agency · What is UGC? · Google Ads agency Oslo

Questions and answers.

Do I need to know how to code to use this?+
No, but it helps to be comfortable running commands. Claude Code writes the code and you steer it in plain language. You need Node and Python installed, an ElevenLabs account and access to Claude Code.
How long does it take to make one ad?+
The first time, setup takes a while. Once the templates are in place, most of the time goes into the script, listening to the voice and fine-tuning, not the animation itself.
Is AI voiceover good enough for ads?+
With the right model and the right script, yes. Use eleven_v4 with the correct language code, write few and flowing sentences, generate several takes and always listen yourself before publishing.
Which format works best for motion ads?+
Of what we have tested: problem and solution, mirrored. Five concrete problems as small scenes, a turning point, and five solutions in the same order with the old problem struck through.
What is the safe zone in a 9:16 video?+
On a 1080×1920 video, keep all text between y=400 and y=1520. Reels, Stories and TikTok cover the top and bottom with profile names, buttons and captions.
Free, no strings attached

Would you rather have us make the ads for you?

Send us your website and we’ll put together a concrete proposal for ads and a landing page for your business, showing how it would look with your brand. You’ll get it by email within a couple of days, and it’s yours to keep either way.

Victor and Jona, Arcads
Victor and Jona make the proposal themselves. Ads and a landing page as they would look for you. For established businesses with a monthly ad budget of NOK 30,000+. No pitch, no obligations.

Free · no lock-in contract · we never share your info

Rather talk directly? Book 30 minutes in the calendar → · victor@arcads.no