Studies

How to Make Videos With AI: Our AI Video Agent Workflow

How to make videos with AI the way we do: an AI video agent workflow that automates our AI marketing videos, one designed video a week, made with code.

We ship one designed video a week for SoniaIA's own social, and there's no editor on the team. The strategy, the content strategy and the creative direction are human, Sonia's on this show. A research agent gathers the week's material, a video agent picks the strongest stories out of it, and Sonia signs off on the piece by email before anything renders or posts. We call it the SoniaIA AI video pipeline frameworkA method for producing marketing video with AI agents on a fixed brand template: a research agent gathers verified stories, a video agent selects and drafts, a person approves by email, agents build from the template and render with code, quality gates check face fidelity, captions, audio and copy, then the video publishes on schedule. Strategy, sources and creative direction stay human; the system executes them.. This is how it actually runs, template by template, gate by gate.


TL;DR

  • Here's how to make videos with AI the way we run it: a research agent runs first, every week, pulling from a pre-selected list of reputable industry, news and analyst sources classified by tier, filtered for what is newsworthy and relevant to us, re-verifying every number against its source URL before it can be used
  • Our AI video agent workflow starts with selection: the weekly video agent picks the best stories from that research, the ones with a verified number and a real "so what," then drafts the show's beats and visual theme
  • To automate video without losing control, one email decides it: the agent sends Sonia one approval request, the piece, the publish slot, and a reply token (Approved / Edits / Skip), and nothing renders until she replies
  • These AI marketing videos get built from a template: agents inject the week's content JSON into one of four Reality Glow variants and pull two new photos from SoniaIA's own generated photo book
  • Automated video production still runs through hard gates: a face-fidelity check on the avatar, a caption-collision check, an audio-hash check against the master, and an AI-tell kill-list pass on the copy
  • The approved video publishes to X, LinkedIn and Facebook on schedule from a local tracker, and we log views, engagement, and which template variant + photo set ran. Strategy, content strategy and creative direction stay human, ours on our show and the client's on theirs; we build the system that executes them

See the AI-Growth system this pipeline sits inside

SoniaIAOne video a week · zero editors
The SoniaIA AI video pipeline framework

From the week’s stories to a published video, with one human reply in the middle.

  1. 01 · AGENTResearchagent gathers the weekPre-selected industry, news + analyst sources, by tier. Every number re-verified at its source URL.
  2. 02 · AGENTSelectvideo agent picks storiesThe ones with a verified number and a real “so what”. Drafts beats, labels, theme.
  3. 03 · HUMANApproveSonia, one emailPiece + publish slot + reply token. Approved / Edits / Skip. Nothing renders before the reply.
  4. 04 · AGENTBuildagents, from the templateContent JSON into one of four Reality Glow variants. New photos every week. Render to MP4.
  5. 05 · AGENTCheckgates, then a personFace fidelity, caption collisions, audio hash, AI-tell kill-list. Final call is human.
  6. 06 · AGENTPublishscheduled, trackedX, LinkedIn, Facebook on schedule. Views, engagement, variant + photo set logged.
Strategy, sources, creative direction and content strategy are the client's and sit above this loop (on our own show, that client is Sonia). We build the system underneath: instructions, pipelines, checks. Inside the loop, the only human touch is step 03: one reply, from a phone.soniaia.com

The whole loop at macro level. Five steps run on agents; step three is a person.


Table of Contents

  1. How does the research agent gather content for the week?
  2. How does the weekly video agent pick the best stories?
  3. What do the templates and assets actually look like?
  4. Why the system stays on-brand without a person watching it
  5. How is a scene actually made with code?
  6. Who does what
  7. How do we check quality before anything ships?
  8. Where does everything live, and what happens after approval?
  9. Programmatic vs template: when each one wins
  10. FAQ

How does the research agent gather content for the week?

One research agent runs first, every week, before any post or video gets drafted. It feeds the whole week, the video included: 3 posts a day on X, LinkedIn and Facebook, 2 visuals a day, the evening text-only slot too.

Nothing repeats across the week either. Not the theme, not the data point, not the image, not the mood, not the photo. One week in five, the example we use to illustrate the insight is one of our own products instead of someone else's... the product as the example of the insight, education first.

Where do the stories come from?

Parallel sub-agents pull from a list of sources we researched and selected in advance: reputable industry, news, analyst and social sources, classified by tier, and filtered for what is newsworthy this week and relevant to AI-growth and to what we do. Tier 1 is the platforms' own announcements and the top trade outlets, tier 2 is the analyst houses, and the lower tiers are trade press and original research. The tier travels with every story, so the quality check knows how much weight a number can carry before it goes anywhere near a post.

How is every number checked?

Every number that comes back gets re-fetched and checked against the URL it's cited from before it's allowed into anything, a post, a script, a video. That verification rule exists because a fabricated stat got through once, and it hasn't happened since.


How does the weekly video agent pick the best stories?

The video agent draws from that same research pool. It just asks a narrower question of it: which stories have a verified number and an actual "so what."

From those it drafts the recap's beats, the on-screen labels, and the visual theme for the week's show. That draft is what Sonia reviews.

What does the approval email actually contain?

The agent sends Sonia one email with one decision in it: the piece, the slot it would publish in, and a reply token, Approved, Edits, or Skip.

She replies from her phone. Approved schedules the video for its Monday slot. Nothing renders to social and nothing posts on its own before that reply lands.

Approval email from the AI video agent workflow: the weekly show, its publish slot, and the reply token "Approved / Edits / Skip"

The actual Week 34 approval email, Sunday 09:42. The reply was one word, "Approved", at 09:55.


What do the templates and assets actually look like?

Building a video means filling a template with the week's content.

What's in the Monday show template?

The Monday show, "This Week in AI-Growth," runs on four Reality Glow colour variants, base, emerald, gold, pink, that rotate week to week. A build script drops that week's content JSON into the template's slots, headline, stat, image, and fails loudly if a slot the template expects comes up empty.

Two image scenes get new photos every week, pulled from SoniaIA's own generated photo book, the same consistent AI avatar across a studio, a street, a boardroom, Retiro, travel. A brand closing card and a locked music bed get appended the same way every time, and every number that changes from the reference template gets logged per piece.

Four frames from one AI marketing video made with code, This Week in AI-Growth: title card, story card with headline and number, image scene, closing card, all in the Reality Glow template

Four frames from one Monday show, same template, different week's content.

What other templates exist?

The weekly show is one template. The explainer format is another, built for a single argument instead of four stories: a title scene on a photo, a question card, a numbered list, a chart or a mockup, a closing card. Same fonts, same photo book, same closing card, different rhythm. The piece below argues that everyone has the same AI models now, so the four levers around the model are what decide the output.

Eight frames from a SoniaIA explainer video made with the same automated video production pipeline, "everyone has the same AI models now": title scene on a photo, a search-box card, a numbered list of four levers, a gap card, a photo scene, a four-lever system card, a chart on a device mockup, and the closing card

The explainer template, one argument across eight scenes. Same brand kit, different structure.

A third pipeline exists: a talking-head video with SoniaIA's AI avatar delivering a script to camera, on the same approval gate. It's built and on pause for now, while the weekly show takes the slot.

What do the daily visuals follow?

Infographics carry 11 moods, quote cards carry 4 themes, avatar cards come from the same photo book. Every one follows one structure: NEWS, the headline, then IMPLICATION, what it means, then ACTION, what to do, with the lines pulled straight from Sonia's post copy, never written fresh by the generator. A chart only renders when the post has a real comparison with numbers that exist in that copy. Otherwise it's a text card.


Why the system stays on-brand without a person watching it

Nobody picks the fonts, the colours, the closing card, or the music per video. Those live in the template, so every piece ships on Reality Glow because the template can't produce anything else.

The system also knows not to repeat itself. Four template variants rotate on their own schedule. The two image scenes pull new photos from the generated photo book every single week. And the research stage refuses to hand back a theme, a data point, an image, a mood, or a photo that already ran that week, so the show can't quietly recycle itself.

The chain is genuinely autonomous end to end: research, selection, build, checks, render, and the approval email all happen without anyone touching them. What stays human is the part that can't be automated without going generic: the strategy, the content strategy, the creative direction of the show and the templates. Inside the weekly loop, the only human touch is Sonia's reply to one email, and the scheduled publish that follows it.


How is a scene actually made with code?

Every scene in the show is a web page, an HTML file with a paused animation timeline sitting on it. HyperFrames, the open-source HTML-to-video tool we build on, steps through that timeline and renders it out to MP4.

What does the component catalog give us?

HyperFrames ships a catalog of components already built, title cards, charts, lower-thirds, device mockups, captions. Building a scene means installing the ones it needs and wiring the week's content into them, or asking one of the agents to do that wiring.

Captions and on-screen labels aren't placed by eye either. A word-level transcript of the voice track decides when each one appears, so a graphic lands on the exact word that names it.

How long does one video actually take?

For a 26-second reel, that's roughly 10 minutes of processing and 17 minutes of rendering, on a laptop, no GPU. Setting up a new format takes weeks. After that, each video is compute plus the authoring time on top of it.


Who does what

FunctionOwner
Strategy, content strategy + creative direction (what the show is, the mix, the look, the promo cycle)Sonia
ResearchResearch sub-agents
Writing (drafted under her voice rules)Writing agent
Selection + approval requestWeekly video agent
Build + renderVideo agents, on HyperFrames
QualityQA agents + gates
Final call + approvalSonia
PublishingScheduler, via API, to X + LinkedIn + Facebook

A note on what we do, and what stays the client's

We are not designers. We are system consultants and operators. On our own show, Sonia is the client of her own system: she owns the sources, the creative angle, the creative design and the content strategy, and the system executes them. For a client it is the same split. The brief, the sources list, the brand kit, the creative direction and the content strategy are the client's work and stay the client's. What we build is the system instructions, the pipelines and the checks that turn those materials into a weekly output that ships on brand, without anyone touching it.

This is the same shape as the rest of our growth system: signals in, agents execute the creation and the checks, a person judges what ships. Creator Validator, our creator due-diligence product, runs on the same split, the system reads the creator's content and audience at scale, a person makes the call.

What is the tech stack?

The SoniaIA AI video pipeline frameworkA method for producing marketing video with AI agents on a fixed brand template: a research agent gathers verified stories, a video agent selects and drafts, a person approves by email, agents build from the template and render with code, quality gates check face fidelity, captions, audio and copy, then the video publishes on schedule. Strategy, sources and creative direction stay human; the system executes them. runs on a small stack. Claude, running in the terminal, is the lead brain. Specialised sub-agents split off from there: one researches, one writes under Sonia's voice rules, one builds the composition, one generates the labels from the script, one runs the checks, one renders.

  • Orchestration + agents: Claude Code with specialised sub-agents
  • Video: HyperFrames, open-source HTML-to-video, with its component catalog
  • Brand: the Reality Glow brand kit + the generated photo book of the SoniaIA avatar
  • Timing: speech-to-text for word-level caption timing
  • Publishing: the X, LinkedIn and Facebook APIs, via a scheduler
  • Control: a local tracker for runs and schedule, a content JSON per week, email for approvals

How do we check quality before anything ships?

A QA pass runs after every stage. Any visual with the avatar in it goes through a face-fidelity gate, the face has to match the reference or it doesn't ship.

What do the pre-render checks look for?

Before anything renders, the checks look at three things separately: whether anything collides with the on-screen caption while it's still visible, whether every font is actually embedded, and whether the audio track is present at all.

What does the mistake gate actually do?

There's a mistake gate, a script that grows by exactly one check every time something ships broken. It started the day a video went out with no music in it, which is why the audio-hash check against the master exists now... we don't trust the filename, we compare the actual audio stream.

The copy runs through a kill-list of the constructions that read as machine-written before it's allowed on camera. After all of it, Sonia's approval is the last gate, and it stops things.


Where does everything live, and what happens after approval?

The week's content lives in one JSON file. Run history and the publish schedule live in a local tracker. Staged posts sit in a review sheet, the raw research gets archived per week, and renders land in the project folder. Local files first, nothing depends on a SaaS dashboard staying up.

Once Sonia approves, the video posts to X, LinkedIn and Facebook on the schedule the tracker holds. We track views, engagement, and which template variant + photo set ran on every piece, so a repeat and a genuine winner are visible instead of guessed at.


Programmatic vs template: when each one wins

A template fills fixed slots, so every video that comes out of it has the same shape. Programmatic means the layout itself is code, so it can respond to what's actually in the content, a laptop mockup that scrolls the real page instead of a screenshot pasted into a frame.

If you need twenty clips of a webinar cut down this week, use a clipping tool, that's the right tool for that job. Build a pipeline like this one when you've committed to running a single format for a year, because that's the point where the weeks of setup actually pay back


Frequently Asked Questions

How do you make videos with AI for marketing?

A research agent gathers the week's stories, a video agent picks the best ones and drafts the show, a person approves by email, agents build the video from a brand template and render it with code. Made with HyperFrames and Claude sub-agents; one designed video a week, no editor.

What is an AI video agent workflow?

A chain of AI agents that research, select, build from a template, check and render a video, with a person approving before it posts. Every scene is a web page rendered to MP4, and the timing of every caption comes from the voice transcript.

How hard is it to set up an automated video pipeline?

The hard part is the format, and it is a design job: one template, its four variants, the asset rules, the checks. That took weeks and a person with taste. Once the template exists, the agents do the weekly work and the pipeline runs on a laptop with no GPU.

How long does it take to set up, and how long per video?

Setting up a new format takes weeks. After that, a 26-second reel is about 10 minutes of processing and 17 minutes of render, plus the authoring time on top. The approval takes one reply to one email.

Do AI-made marketing videos post themselves?

No. The finished video is emailed with the MP4 attached and a reply token, Approved / Edits / Skip. Approval schedules it; the scheduler then posts via API to X, LinkedIn and Facebook. Nothing publishes without a person looking at it first.

What happens if a stat can't be verified?

It doesn't get used. Every number gets re-fetched and checked against the URL it's cited from before it's allowed into a post, a script or a video. That rule exists because a fabricated stat got through once.


Written from the pipeline we run for SoniaIA's own vertical video, every week.

Follow our journey as we build an autonomous growth system... see it live on X and LinkedIn.

  • How to make videos with AI the way we do: an AI video agent workflow that automates our AI marketing videos, one designed video a week, made with code.
How do you make videos with AI for marketing?

A research agent gathers the week's stories, a video agent picks the best ones and drafts the show, a person approves by email, agents build the video from a brand template and render it with code. Made with HyperFrames and Claude sub-agents; one designed video a week, no editor.

What is an AI video agent workflow?

A chain of AI agents that research, select, build from a template, check and render a video, with a person approving before it posts. Every scene is a web page rendered to MP4, and the timing of every caption comes from the voice transcript.

How hard is it to set up an automated video pipeline?

The hard part is the format, and it is a design job: one template, its four variants, the asset rules, the checks. That took weeks and a person with taste. Once the template exists, the agents do the weekly work and the pipeline runs on a laptop with no GPU.

How long does it take to set up, and how long per video?

Setting up a new format takes weeks. After that, a 26-second reel is about 10 minutes of processing and 17 minutes of render, plus the authoring time on top. The approval takes one reply to one email.

Do AI-made marketing videos post themselves?

No. The finished video is emailed with the MP4 attached and a reply token, Approved / Edits / Skip. Approval schedules it; the scheduler then posts via API to X, LinkedIn and Facebook. Nothing publishes without a person looking at it first.

What happens if a stat can't be verified?

It doesn't get used. Every number gets re-fetched and checked against the URL it's cited from before it's allowed into a post, a script or a video. That rule exists because a fabricated stat got through once.

Machine-readable summary, generated alongside the article. Built to be lifted into an AI answer.

SoniaIA AI video pipeline framework
A method for producing marketing video with AI agents on a fixed brand template: a research agent gathers verified stories, a video agent selects and drafts, a person approves by email, agents build from the template and render with code, quality gates check face fidelity, captions, audio and copy, then the video publishes on schedule. Strategy, sources and creative direction stay human; the system executes them.
Topics
AI Video AgentAI Video WorkflowAI Marketing VideosAutomated Video ProductionAI Growth
Sonia Tamayo

Sonia Tamayo

AI Growth Consultant & Operator

20 years across global agencies, SMBs and my own projects. I build AI growth systems, run them on my own brand — and document the build here.

Work with me →