Projeto

My-Little-Studio

Transformando semanas de edição, da gravação ao vídeo exportado, em minutos.

Iago Russi FerlaAutor
set 2026Publicado
16 minLeitura
Ler ↓6 partes
Parte 1

De onde veio

A edição que eu fazia

Pra quem não me conhece: fui editor de vídeo por 4 anos e hoje sou engenheiro de software. Trabalhei no Ossos Perdidos, um canal de true crime bem famoso, hoje com mais de um milhão de inscritos, e participava de todas as partes da criação de um vídeo profissional. Se quiser saber mais sobre essa jornada, está tudo em Minha jornada até aqui.

A gente já usava IA na época em que eu editava, mas ela ainda engatinhava. Quando virei engenheiro de software, em algum momento pensei: e se eu pegasse todo aquele processo e automatizasse com as ferramentas que tenho hoje? Foi daí que surgiu o primeiro resquício deste projeto.

Como surgiu o projeto

Hoje eu trabalho para a consultoria do meu amigo e parceiro Dyego Maas, e este projeto caiu como uma luva. Ele vai fazer parte de um ecossistema bem maior, mas a ideia aqui é que os vídeos que a empresa posta em qualquer mídia digital passem por um pipeline automático.

E esse pipeline é exatamente isso: automatizar todo o processo criativo de um vídeo, do arquivo cru ao vídeo exportado, usando IA. O Dyego me ajudou muito com ferramentas e conselhos sobre qual caminho seguir, e assim começamos o projeto, que batizamos de My-Little-Studio.

O vídeo que saiu do My-Little-Studio: da gravação crua ao vídeo pronto.
Parte 2

A engenharia

Pipes and Filters

Já tínhamos a ideia mapeada, mas precisávamos de um padrão de arquitetura decente para ela. Depois de muita pesquisa, o que mais se encaixava no que queríamos era o Pipes and Filters.

Pipes and Filters (canos e filtros) é um padrão com duas peças. Os filtros são etapas independentes: cada uma recebe um dado, transforma e entrega. Os canos ligam uma etapa à próxima, levando o que uma produziu para a outra consumir.

É o mesmo desenho dos pipes do Unix, a família de sistemas operacionais nascida nos anos 70 que deu base ao macOS e inspirou o Linux. No terminal, "cat arquivo | grep erro | sort" liga três programas pequenos pelo caractere "|": cada um faz uma coisa só, e a saída de um vira a entrada do próximo.

Para vídeo, isso serve muito bem, porque precisávamos que cada etapa pudesse ser rodada, testada e trocada sozinha.

vídeo cru
filtrocortar
script
vídeo cruvídeo cortado
filtrolegendar
script
vídeo cortadovídeo legendado
filtroanimar
script
vídeo legendadovídeo pronto
filtro: um script que faz uma coisa sócano: leva o resultado para o próximo
Pipes and Filters: cada script faz uma coisa só, e o que ele entrega segue pela seta até o próximo.

O projeto por dentro

Além da arquitetura, a regra que queríamos seguir era simples: uma parte determinística e uma parte não determinística. A determinística são os scripts, que fazem sempre a mesma coisa. A não determinística é a conversa com o agente: ele segue as instruções que ficam no projeto, mas cada conversa sai de um jeito.

Os scripts, que no projeto chamamos de runners, são as etapas determinísticas: cada um faz uma operação do pipeline de vídeo. Junto deles ficam os recibos de cada execução, para saber exatamente o que aconteceu, e os ajustes, onde ficam os parâmetros e os setups de cor.

~não determinísticoA conversaEu peço em português. O agente lê o manual daquele trabalho e decide o que rodar.
=determinísticoOs scriptsCada um faz uma coisa só. Com a mesma entrada, dão sempre o mesmo resultado.
reciboscada execução anota o que entrou, o que saiu e quanto custou
ajustesparâmetros e setups de cor salvos para os próximos vídeos
As duas metades do projeto: a conversa, que muda a cada vez, e os scripts, que dão sempre o mesmo resultado.

Os filtros: os scripts

Cada script faz uma coisa só. Hoje já são 29, divididos em categorias: corte e limpeza, legenda, cor, corpo e gesto, imagem e animação. Todos têm um parâmetro --help que explica como usar, e dá para rodar qualquer um sozinho pelo terminal. Na próxima parte mostro os principais, categoria por categoria.

E eles são determinísticos: com a mesma entrada, dão sempre o mesmo resultado. As exceções são os scripts que chamam um serviço de IA de fora, como a AssemblyAI na legenda e o gpt-image-2 na geração de imagem, cuja resposta pode variar. Mesmo assim, tudo o que eles pedem e recebem fica registrado no recibo.

Os canos: o acervo e o rastro

Os scripts precisam conversar. Dá para rodar cada um sozinho, mas alguns dependem do resultado de outros: o que coloca a borda do corte no começo da sílaba, por exemplo, precisa antes saber quais trechos ficaram. Para isso existem o acervo e o rastro.

O acervo guarda cada vídeo numa pasta própria: o original, cada arquivo derivado e, dentro dela, o rastro. No rastro, cada etapa grava o que decidiu e anota a impressão digital (um hash sha256) de cada arquivo que leu.

Antes de trabalhar, a etapa confere essas impressões digitais. Se um dado de antes mudou, ela percebe que está trabalhando em cima de algo velho, se recusa a rodar e diz exatamente qual comando falta. Isso aconteceu de verdade neste vídeo: a revisão foi recusada duas vezes, porque uma etapa anterior tinha rodado de novo.

Além disso, toda execução deixa um recibo: o que entrou, o que saiu, os parâmetros, o tempo e o custo. Ela termina de um de três jeitos: deu certo, falhou (e nada pela metade fica no acervo) ou já estava feito (e nada é refeito nem pago).

script de corteRodou em cima do arquivo. A marca do arquivo virou 5f9.arquivomarca 5f9
agenteRevisou em cima do arquivo 5f9 e anotou: “revisei o 5f9”.arquivomarca 5f9anotado: 5f9
script de corte, de novoAlguém rodou o script de corte de novo no mesmo arquivo. A marca mudou.arquivomarca a3c
script que aplicaAnotado 5f9, arquivo a3c. Não bate: a revisão ficou velha. “Rode a revisão de novo.”arquivomarca a3c
O rastro em quatro quadros: a revisão anota a marca do arquivo, e quem aplica recusa quando a marca mudou.

O agente e as skills

O agente, que por enquanto é só o Claude Code, não precisa ler os scripts: isso consumiria muito contexto. Em vez disso, cada tipo de trabalho tem uma skill dentro do projeto.

As skills são manuais de trabalho. Elas não explicam cada script; explicam o procedimento para um tipo de trabalho, como cortar, analisar o enquadramento ou legendar: quando usar, qual script rodar, com quais parâmetros e onde olhar se der problema. Só aí, quando algo dá errado, o agente abre o código. Hoje são 10 skills para os 29 scripts.

Parte 3

Os scripts, categoria por categoria

Corte e limpeza

Para cortar direito, primeiro precisávamos fixar o frame rate. Muita gravação vem com frame rate variável, e aí um corte cai num quadro na imagem e em outro no som. O fix-framerate faz exatamente isso: normaliza o vídeo para 30 quadros por segundo fixos. Depois vêm os scripts básicos de áudio: tirar o áudio alinhado ao vídeo e limpar o ruído, que só é tratado se a medição mostrar que precisa.

O grande diferencial é o transcribe-tape, que monta o que chamamos de fita. Ele usa um modelo de reconhecimento de fala em português (wav2vec2) que roda no próprio computador e diz qual letra soa a cada 20 milissegundos. A transcrição comercial escondia as frases repetidas, juntando as duas tentativas numa só. Na fita, a retomada aparece escrita duas vezes, e é ela que decide o corte.

Com a fita, o resto da cadeia acha os trechos de voz, fica com a última tentativa de cada frase, coloca cada borda do corte no começo da sílaba e prepara um formulário de revisão, que o agente preenche como um editor. Depois vêm o corte em si e o volume nivelado. O roteiro não decide o corte: ele só aparece na revisão, para mostrar se alguma linha ficou sem ser dita.

mídia · prepara imagem e som
fix-framerateextract-audiodenoisetranscribe-tape

O fix-framerate deixa o vídeo com 30 quadros por segundo fixos, o extract-audio tira o áudio alinhado ao primeiro quadro, o denoise mede o ruído e só limpa quando precisa, e o transcribe-tape monta a fita, que diz qual letra soa a cada 20 ms.

ffmpeg · DNSMOS + DeepFilterNet3 · wav2vec2, local
decisão · decide o que fica
detect-speechdedupe-retakespad-edgesreview-draftapply-review

O detect-speech acha os trechos de voz, o dedupe-retakes fica com a última tentativa de cada frase repetida, o pad-edges põe cada borda do corte no começo da sílaba, o review-draft prepara o formulário que o agente preenche como um editor e o apply-review confere e aplica essas decisões.

a fita · Node · Claude, na revisão
saída · entrega o arquivo
render-cutloudnorm

O render-cut corta imagem e som como duas entradas separadas, e o loudnorm nivela o volume em −14 LUFS.

ffmpeg
ferramentas · fora da cadeia
clean-recordingplan-show

O clean-recording roda a cadeia inteira em ordem e para na revisão, e o plan-show mostra onde a cadeia está e o que ficou velho.

orquestrador · só leitura

Legenda

Na legenda, usamos a AssemblyAI como uma segunda camada. A fita já sabe exatamente quando cada palavra é dita, mas escreve do jeito que soa, sem dicionário. A AssemblyAI entra com o texto bem escrito, e a fita cola cada palavra no tempo certo e confere se o que foi escrito bate com o que soou.

Quando uma palavra não bate, ela vira uma dúvida para o agente resolver, em vez de ir errada para a tela. Assim os erros não passam e a legenda sai melhor. Nenhuma das duas sozinha dava conta: a AssemblyAI erra o fim da palavra em até 150 ms, e a fita sozinha cola palavras.

legenda · texto, tempo e conferência
caption

O caption pede o texto à AssemblyAI, usa a fita para dar o tempo de cada palavra, marca como dúvida a palavra que não bate com o som e guarda no glossário a grafia que eu confirmo.

AssemblyAI · wav2vec2, local
texto da AssemblyAI · tempo da fita
As primeiras palavras do vídeo, cada uma no tempo que a fita deu.

Cor

Aqui a ideia é aplicar LUTs e ter sempre um lugar salvo para buscar depois. O color-grade aplica um Setup, que é o LUT mais o ganho e a nitidez calibrados uma vez para uma câmera, uma luz e um cenário, e esse Setup fica guardado para os próximos vídeos.

cor · o visual da gravação
color-grade

O color-grade aplica um Setup, com LUT, ganho e nitidez, calibrado uma vez para uma câmera, uma luz e um cenário.

ffmpeg lut3d

Corpo e gesto

No vídeo lá do começo eu simplifiquei: o OpenCV é responsável só por achar o meu rosto. Na verdade são três ferramentas para captar tudo:

  • o YuNet, um modelo do OpenCV, acha o rosto;
  • o MediaPipe, do Google, recorta a minha silhueta do fundo;
  • o RTMPose segue as duas mãos, o rosto e os ombros, quadro a quadro.

Com esse rastreio, o gesture-events tira cada gesto no quadro exato em que aconteceu, e o framing-analysis monta um mapa do corpo, com cabeça, tronco, braços e mãos, que diz onde um elemento cabe sem encostar em mim. Uma lição que aprendemos medindo: a posição nunca vem de um modelo de linguagem. O Gemini errava 12% da largura da tela.

rastreio · onde está cada parte
track-personhand-detaildetect-cuts

O track-person segue as duas mãos, o rosto e os ombros em cada quadro, o hand-detail lê se a mão está aberta ou fechada e o detect-cuts acha as emendas de um vídeo que não passou pelo corte do studio.

RTMPose · MediaPipe · PySceneDetect
comandos · o que a mão fez
gesture-events

O gesture-events tira do rastreio cada gesto no quadro exato: o pico de movimento, a mão abrindo, a batida do dedo, o arrasto e para onde o indicador aponta.

Node, sobre o rastro
mapa · onde cabe um elemento
framing-analysisframing-query

O framing-analysis nomeia as partes do quadro e desenha o corpo a cada 0,2 s, e o framing-query diz onde um elemento cabe sem encostar no corpo e o que o fez encolher.

YuNet (OpenCV) · MediaPipe · regra C1
regra C1 · onde cabe um elemento faixa da legenda cartão margem por partecabeça 80 pxtronco 40 pxbraço 15 pxmão 0 px encolheu 16%pelo braçoaos 12,4 s
A regra C1: o elemento fica onde não encosta em nenhuma parte do corpo.

Imagem

A imagem é mais tranquila. O gpt-image-2 fica só para as coisas premium: neste vídeo, foram só algumas janelas. Os resultados só com React já são impressionantes: quase todos os cartões e janelas do vídeo saíram dele.

imagem · gerar e ajustar
generate-assetsupscalecompress

O generate-assets cria imagem com o gpt-image-2, o upscale aumenta a resolução pelo Replicate e o compress deixa o arquivo mais leve.

gpt-image-2 · Replicate · ImageMagick

Animação

A animação é onde o HyperFrames faz a mágica acontecer (explico ele melhor no fim). Escolhemos ele em vez do Remotion porque tem uma documentação mais completa e foi pensado para ser usado por agentes.

Hoje dois scripts já preparam a animação: a página onde eu desenho (o quadro de rascunho, que mostro mais abaixo) e o cruzamento desse desenho com a minha mão. O resto ainda foi feito pelo agente em código, e cada parte vira um script na próxima fase.

script hoje
sketchsketch-compare

O sketch gera a página onde eu desenho em cima do vídeo, quadro a quadro, e o sketch-compare cruza esse desenho com a palma da mão e devolve em números onde o elemento nasce e onde assenta.

HTML + Node
fase 2
planoassetscomposiçãosonsrender

O plano é a animação escrita por IA antes de virar código, o catálogo de assets guarda os componentes React prontos para reusar, a composição é uma página HTML animada com GSAP no HyperFrames, os sons acompanham cada movimento e o render junta voz, sons e trilha. Hoje o agente faz isso em código, e cada parte vira um script na fase 2.

IA · React · HyperFrames · ffmpeg
Parte 4

Os quadros de edição visual

Prompts visuais

Por mais que o agente conseguisse fazer as animações, eu tinha muita dificuldade de fazer animações personalizadas. E acho que esse é o maior problema hoje na geração de qualquer conteúdo com IA, seja vídeo ou asset: conseguir pedir algo específico.

Já trabalhei em projetos em que muita gente esbarrava nisso, principalmente na geração de vídeo. O resultado vem bom, mas não vem o que você tinha na cabeça, porque é difícil fazer a ferramenta entender em texto uma coisa visual.

Aqui aconteceu o mesmo. Os scripts já faziam um vídeo tranquilo, mas eu queria algo mais específico. Então criamos um quadro que funciona quase como um Premiere: ele mostra o vídeo quadro a quadro, eu desenho em cima e escrevo instruções para o agente.

Como é HTML, o quadro exporta exatamente em qual quadro eu estava, qual era a instrução e o desenho. E, por incrível que pareça, a IA lê muito bem imagem: vendo o meu desenho, ela entendeu o que eu queria melhor do que eu tinha conseguido explicar antes só com texto. Chamei isso de prompts visuais.

O quadro de rascunho: meus desenhos em cima do vídeo, quadro a quadro (2x).

E tem um bônus: depois de ver o desenho, a IA explica em texto o que entendeu, e melhor do que eu. Essas explicações viram prompts melhores para os scripts automatizados do futuro. Só no primeiro rascunho deste vídeo foram 123 traços e 15 notas.

O quadro de animação

Com a animação foi a mesma coisa: ela errava, e a animação é um pouco diferente, porque trabalha mais com o tempo. Não é tanto desenhar, é dizer quando. Então criei também o quadro de animação, que mostra o vídeo inteiro e cada animação separada, quadro a quadro.

Ali eu consigo definir exatamente quando uma animação entra e quando sai, avançar e voltar quadro a quadro e apontar onde ela estava travada. Do mesmo jeito, isso gera resultados que eu posso reaproveitar para fazer scripts de animação automatizados melhores.

O quadro de animação: o vídeo inteiro com as animações por cima (2x).

O quadro de corte

A mesma ideia foi aplicada ao corte. Os scripts já cortavam bem, mas, estressando o projeto, apareceram problemas específicos: com um respiro, ou com uma frase em que ele não sabia se cortava ou não, ele cortava errado.

No quadro de corte, eu marco na agulha o que está errado, o agente muda e explica o que mudou. E cada correção vira aprendizado para os próprios scripts: o projeto se melhora a cada rodada. É o conceito que mais me interessa aqui, e provavelmente vou levar esse tipo de quadro para outras etapas também.

O quadro de corte: a agulha, a linha do tempo e o que o agente mudou a partir das minhas notas (2x).
Parte 5

Resultado e o que vem

Os números

Hoje, com os scripts prontos, o corte leva minutos, e isso já deixa a etapa de limpeza pronta. A animação é o que mais demora: uns 30 minutos para eu desenhar e ajustar. Mas é tudo com linguagem natural e desenho na mão, é bem tranquilo e sai algo muito funcional e bem feito.

A ideia é que, quando esses scripts ficarem prontos, o vídeo inteiro leve minutos, e os testes que venho fazendo já apontam para isso. O vídeo lá do começo tem pouco mais de 2 minutos; é claro que, quanto maior o vídeo, maior o tempo, mas acredito que ainda vai ser muito bom. Quando eu editava, um vídeo nesse estilo levava mais de uma semana.

E ferramenta de geração pronta não resolve. O resultado vem bom, mas não específico, e quando você precisa de algo que a ferramenta não faz, acaba tendo que buscar outras tecnologias, que geralmente custam dinheiro.

Cada sessão deixa o projeto melhor

Cada sessão do Claude é valiosa. Cada sessão em que trabalhamos deixa o projeto melhor, porque ela tem dados muito bons: o que o agente errou, o que ele precisou, como trabalhou e o que executou.

Um exercício muito interessante é perguntar para o próprio agente que dificuldades ele teve para rodar um script ou fazer um trabalho, onde houve divergência e onde deu problema. Trabalhando em cima disso, o projeto fica cada vez mais robusto. A gente vem usando muito isso, e vale para qualquer projeto.

O que vem

O projeto acabou de começar e ainda está longe de acabar, mas já está entregando. A ideia é ter logo uma versão funcional de ponta a ponta. Hoje estou focado nos scripts de animação, que são os que mais têm problema: o mapa do corpo já entrou, e agora vêm o plano de animação, o catálogo de assets, a composição, os sons e o render.

Depois disso, a maioria dos scripts já está sólida e só precisa de uma leve afiada: corte, legenda e o resto. Com tudo pronto, a ideia é estressar o projeto, ver os pontos fracos e começar a trabalhar nas ideias do backlog.

E teve uma parte que eu nem mostrei aqui: o storyboard. Antes de gravar, uma etapa inteira gera as ideias, escreve o roteiro e monta uma folha de storyboard com um esboço para cada quadro, e o resultado ficou impressionante. Mas ela tem mais a ver com a criação da ideia do que com a edição, então vale outra postagem, quando esse processo estiver mais consolidado.

E quero deixar um adendo sobre o Jev, da TypeSafe AI, lançado agora em setembro. Ele é o primeiro de uma nova classe de modelos que eles chamam de System One, nome tirado da ideia do Kahneman: o Sistema 1 é o pensamento rápido e intuitivo, e o Sistema 2, o lento e deliberado. Em vez de escrever texto palavra por palavra, como um LLM, o Jev recebe a situação e as respostas possíveis, e devolve uma delas já no formato certo, com a probabilidade de cada uma.

Segundo eles, ele responde entre 70 e 500 milissegundos, custa uma fração de um LLM e, como só escolhe entre as opções que você deu, não tem como inventar uma resposta fora delas. Pra este projeto isso é muito interessante: várias decisões pequenas ainda passam pela conversa com o agente, como se uma tentativa fica ou sai na revisão, ou se uma palavra duvidosa da legenda está certa. São perguntas de resposta fechada, e um modelo assim poderia responder rápido, barato e sempre no formato certo, deixando a parte não determinística bem menor. Ele ainda está em acesso antecipado e eu não testei, mas é uma das coisas que mais quero trazer para cá.

LinkIntroducing System One models and Jev ↗
Parte 6

As tecnologias

HyperFrames

O HyperFrames é o coração da animação: é ele que faz tudo funcionar. Pra mim, é provavelmente a tecnologia mais interessante do projeto, e a que mais me surpreendeu.

A ideia é simples: cada animação é uma página HTML, com o movimento feito em GSAP. Para virar vídeo, ele abre essa página num Chrome invisível, avança quadro a quadro, fotografa cada um e junta tudo com o FFmpeg. Como é tudo HTML, qualquer coisa que um navegador consegue desenhar vira vídeo, e a mesma página dá sempre o mesmo resultado.

Além disso, a documentação é incrível. Acredito que não usei nem 30% do que essa biblioteca consegue fazer, e dá para ter muito mais criatividade do que eu tive até aqui. Estou começando agora, e acho que ela é um ponto de partida muito bom para criar qualquer coisa em vídeo.

LinkDocumentação do HyperFrames ↗

A fita: o modelo que lê som a som

A fita, com certeza, foi uma das que mais me surpreenderam. Testei várias coisas para conseguir uma transcrição exata do que eu estava falando, e essa foi a melhor de todas.

É a fita que expliquei no corte: ela sabe os sons do que você falou, letra por letra, inclusive quando você repete uma frase.

Nos testes de estresse, nada chegou perto. A AssemblyAI, em todas as configurações que testei, juntava as duas tentativas de uma frase numa só, e, das outras que pesquisei, a única que marcava as repetições nem instalava direito no Windows. E o melhor de tudo: a fita roda local, no meu computador, e bem tranquila.

LinkO modelo no Hugging Face ↗

O som: DeepFilterNet3 e DNSMOS

Essas duas cuidam da limpeza do áudio, e trabalham em dupla. O DNSMOS, da Microsoft, é um modelo que escuta a gravação e dá três notas, como uma pessoa daria: para a voz, para o fundo e para o conjunto. É por essas notas que o projeto decide se tem ruído para tirar e, depois, se a limpeza preservou a voz.

O DeepFilterNet3 é quem tira o ruído de fato. É um modelo de código aberto que separa a voz do barulho de fundo, e roda no computador sem precisar de placa de vídeo. Antes dele, usávamos um filtro fixo, e ele acabava piorando as gravações que já estavam limpas.

LinkDeepFilterNet no GitHub ↗

O corpo: RTMPose, MediaPipe e YuNet

Essas foram as ferramentas que deixaram as coisas mais legais. As animações já funcionavam, mas eu queria algo que rastreasse o meu movimento. Sabia que não ia ser tão difícil, porque provavelmente já existia algo pronto.

Do OpenCV, usei uma coisa bem específica, o YuNet, que acha o rosto. Sei que o OpenCV tem muito mais, e talvez eu não tenha visto tudo.

Aqui também fiz testes de estresse: coloquei oito bibliotecas lado a lado no mesmo vídeo, entre elas o SAM 2 e o CoTracker, da Meta, o YOLO e até o Gemini, e essas três foram as que deram melhor. E, como eu já tinha falado, o mais surpreendente é que tudo roda no PC tranquilamente, sem nenhum problema.

Linkrtmlib no GitHub ↗

Os serviços de fora

Sobre custos: tudo aqui roda essencialmente local, e para o projeto inteiro eu só precisei de um plano do Claude: tudo foi rodado pelo Claude Code. Mas usei duas coisas de fora, a AssemblyAI e o gpt-image-2. Dá para fazer o vídeo sem elas: sem a AssemblyAI, ele só sai sem legenda.

O gpt-image-2 foi usado só para os assets premium. Os assets que o React fazia já eram suficientes para um vídeo de muita qualidade, mas como eu queria algo um pouco mais personalizado e bonito, gerei alguns com ele. Foram poucos, só algumas janelas.

A AssemblyAI é usada numa etapa só, a da legenda, para trazer o texto bem escrito. Ela é só por enquanto: já conheço alternativas melhores, só ainda não tive tempo de testar, e acredito que dá para fazer o mesmo com um modelo que não custa nada. A gente só está usando porque queria responder rápido uma pergunta e fazer funcionar de um jeito eficiente.

← Home
Project

My-Little-Studio

Turning weeks of editing, from the recording to the exported video, into minutes.

Iago Russi FerlaAuthor
Sep 2026Published
16 minRead
Read ↓6 parts
Part 1

Where it came from

The editing I used to do

For those who don't know me: I was a video editor for 4 years and today I'm a software engineer. I worked on Ossos Perdidos, a very well-known true crime channel, now with over a million subscribers, and I took part in every step of making a professional video. If you want to know more about that journey, it's all in My journey so far.

We already used AI back when I was editing, but it was still crawling. When I became a software engineer, at some point I thought: what if I took that whole process and automated it with the tools I have today? That's where the first trace of this project came from.

How the project started

Today I work for the consultancy of my friend and partner Dyego Maas, and this project fit like a glove. It's going to be part of a much bigger ecosystem, but the idea here is that the videos the company posts on any digital channel go through an automatic pipeline.

And that pipeline is exactly this: automating the whole creative process of a video, from the raw file to the exported video, using AI. Dyego helped me a lot with tools and advice on which path to take, and that's how we started the project, which we named My-Little-Studio.

The video that came out of My-Little-Studio: from the raw recording to the finished video.
Part 2

The engineering

Pipes and Filters

We already had the idea mapped out, but we needed a decent architecture pattern for it. After a lot of research, the one that best fit what we wanted was Pipes and Filters.

Pipes and Filters is a pattern with two pieces. Filters are independent steps: each one receives some data, transforms it and hands it over. Pipes connect one step to the next, carrying what one produced for the other to consume.

It's the same design as Unix pipes, the family of operating systems born in the '70s that became the base of macOS and inspired Linux. In the terminal, "cat file | grep error | sort" connects three small programs with the "|" character: each one does a single thing, and the output of one becomes the input of the next.

For video, that works really well, because we needed every step to be run, tested and swapped on its own.

raw video
filtercut
script
raw videocut video
filtercaption
script
cut videocaptioned video
filteranimate
script
captioned videofinished video
filter: a script that does one thingpipe: carries the result to the next one
Pipes and Filters: each script does one thing, and what it hands over follows the arrow to the next one.

The project from the inside

Beyond the architecture, the rule we wanted to follow was simple: a deterministic part and a non-deterministic part. The deterministic part is the scripts, which always do the same thing. The non-deterministic part is the conversation with the agent: it follows the instructions that live in the project, but every conversation turns out a little different.

The scripts, which we call runners in the project, are the deterministic steps: each one does one operation of the video pipeline. Next to them are the receipts of every run, so we know exactly what happened, and the settings, where the parameters and the color setups live.

~non-deterministicThe conversationI ask in plain language. The agent reads the manual for that job and decides what to run.
=deterministicThe scriptsEach one does one thing. Given the same input, they always give the same result.
receiptsevery run records what went in, what came out and what it cost
settingsparameters and color setups saved for the next videos
The two halves of the project: the conversation, which changes every time, and the scripts, which always give the same result.

The filters: the scripts

Each script does one thing. There are already 29 of them, split into categories: cutting and cleanup, captions, color, body and gesture, image and animation. They all have a --help flag that explains how to use them, and any of them can run on its own from the terminal. In the next part I show the main ones, category by category.

And they're deterministic: given the same input, they always give the same result. The exceptions are the scripts that call an outside AI service, like AssemblyAI for captions and gpt-image-2 for image generation, whose answers can vary. Even so, everything they send and receive is recorded in the receipt.

The pipes: the library and the trail

The scripts need to talk to each other. Each one can run on its own, but some depend on the result of others: the one that places each cut edge at the start of a syllable, for example, first needs to know which stretches were kept. That's what the library and the trail are for.

The library keeps each video in its own folder: the original, every derived file and, inside it, the trail. In the trail, each step records what it decided and notes the fingerprint (a sha256 hash) of every file it read.

Before working, a step checks those fingerprints. If some earlier data changed, it notices it would be working on something stale, refuses to run and says exactly which command is missing. This really happened on this video: the review was refused twice, because an earlier step had run again.

On top of that, every run leaves a receipt: what went in, what came out, the parameters, the time and the cost. It ends in one of three ways: it worked, it failed (and nothing half-done stays in the library), or it was already done (and nothing is redone or paid for again).

cut scriptRan on the file. The file’s mark became 5f9.filemark 5f9
agentReviewed on top of file 5f9 and noted: “reviewed 5f9”.filemark 5f9noted: 5f9
cut script, againSomeone ran the cut script again on the same file. The mark changed.filemark a3c
applying scriptNoted 5f9, file is a3c. No match: the review is stale. “Run the review again.”filemark a3c
The trail in four panels: the review notes the file's mark, and the step that applies it refuses when the mark has changed.

The agent and the skills

The agent, which for now is only Claude Code, doesn't need to read the scripts: that would eat up a lot of context. Instead, every kind of job has a skill inside the project.

Skills are work manuals. They don't explain each script; they explain the procedure for a kind of job, like cutting, analyzing the framing or captioning: when to use it, which script to run, with which parameters and where to look if something goes wrong. Only then, when something breaks, does the agent open the code. Today there are 10 skills for the 29 scripts.

Part 3

The scripts, category by category

Cutting and cleanup

To cut properly, we first needed to lock the frame rate. A lot of recordings come with a variable frame rate, and then a cut lands on one frame in the picture and on another in the sound. fix-framerate does exactly that: it normalizes the video to a constant 30 frames per second. Then come the basic audio scripts: extracting the audio aligned with the video and removing noise, which is only treated if the measurement shows it needs to be.

The big differentiator is transcribe-tape, which builds what we call the tape. It uses a Portuguese speech recognition model (wav2vec2) that runs on my own computer and says which letter sounds every 20 milliseconds. The commercial transcription hid repeated sentences, merging the two attempts into one. On the tape, the retake shows up written twice, and that's what decides the cut.

With the tape, the rest of the chain finds the stretches of speech, keeps the last attempt of each sentence, places each cut edge at the start of a syllable and prepares a review form, which the agent fills in like an editor would. Then come the cut itself and the leveled volume. The script doesn't decide the cut: it only shows up in the review, to point out any line that was never said.

media · prepares picture and sound
fix-framerateextract-audiodenoisetranscribe-tape

fix-framerate locks the video at a constant 30 frames per second, extract-audio pulls out the audio aligned with the first frame, denoise measures the noise and only cleans when it needs to, and transcribe-tape builds the tape, which says which letter sounds every 20 ms.

ffmpeg · DNSMOS + DeepFilterNet3 · wav2vec2, local
decision · decides what stays
detect-speechdedupe-retakespad-edgesreview-draftapply-review

detect-speech finds the stretches of speech, dedupe-retakes keeps the last attempt of each repeated sentence, pad-edges places each cut edge at the start of a syllable, review-draft prepares the form the agent fills in like an editor, and apply-review checks and applies those decisions.

the tape · Node · Claude, in the review
output · delivers the file
render-cutloudnorm

render-cut cuts picture and sound as two separate inputs, and loudnorm levels the volume at −14 LUFS.

ffmpeg
tools · outside the chain
clean-recordingplan-show

clean-recording runs the whole chain in order and stops at the review, and plan-show shows where the chain is and what went stale.

orchestrator · read-only

Captions

For captions, we use AssemblyAI as a second layer. The tape already knows exactly when each word is said, but it writes the way things sound, with no dictionary. AssemblyAI brings properly spelled text, and the tape pins each word to the right time and checks whether what was written matches what was heard.

When a word doesn't match, it becomes a question for the agent to resolve, instead of going wrong onto the screen. That way errors don't slip through and the captions come out better. Neither one could do it alone: AssemblyAI misses the end of a word by up to 150 ms, and the tape alone glues words together.

captions · text, timing and checking
caption

caption asks AssemblyAI for the text, uses the tape to time each word, flags as a question any word that doesn't match the sound, and saves in the glossary the spelling I confirm.

AssemblyAI · wav2vec2, local
text from AssemblyAI · timing from the tape
The first words of the video, each one at the time the tape gave it.

Color

Here the idea is to apply LUTs and always have a saved place to come back to. color-grade applies a Setup, which is the LUT plus gain and sharpness calibrated once for a camera, a light and a set, and that Setup is kept for the next videos.

color · the look of the recording
color-grade

color-grade applies a Setup, with LUT, gain and sharpness, calibrated once for a camera, a light and a set.

ffmpeg lut3d

Body and gesture

In the video at the top I simplified it: OpenCV is only in charge of finding my face. In reality there are three tools capturing everything:

  • YuNet, an OpenCV model, finds the face;
  • Google's MediaPipe cuts my silhouette out of the background;
  • RTMPose follows both hands, the face and the shoulders, frame by frame.

With that tracking, gesture-events pulls out each gesture at the exact frame it happened, and framing-analysis builds a map of the body, with head, torso, arms and hands, that says where an element fits without touching me. A lesson we learned by measuring: position never comes from a language model. Gemini was off by 12% of the screen width.

tracking · where each part is
track-personhand-detaildetect-cuts

track-person follows both hands, the face and the shoulders in every frame, hand-detail reads whether the hand is open or closed, and detect-cuts finds the splices in a video that didn't go through the studio's cut.

RTMPose · MediaPipe · PySceneDetect
commands · what the hand did
gesture-events

gesture-events pulls each gesture out of the tracking at the exact frame: the peak of movement, the hand opening, the finger tap, the drag and where the index finger points.

Node, over the trail
map · where an element fits
framing-analysisframing-query

framing-analysis names the parts of the frame and draws the body every 0.2 s, and framing-query says where an element fits without touching the body and what made it shrink.

YuNet (OpenCV) · MediaPipe · rule C1
rule C1 · where an element fits caption band card margin per parthead 80 pxtorso 40 pxarm 15 pxhand 0 px shrank 16%for the armat 12.4 s
Rule C1: the element goes where it doesn't touch any part of the body.

Image

Images are the easy part. gpt-image-2 is only for the premium stuff: in this video, just a few windows. The results with React alone are already impressive: almost every card and window in the video came out of it.

image · generate and adjust
generate-assetsupscalecompress

generate-assets creates images with gpt-image-2, upscale increases the resolution through Replicate, and compress makes the file lighter.

gpt-image-2 · Replicate · ImageMagick

Animation

Animation is where HyperFrames works its magic (I explain it better at the end). We chose it over Remotion because its documentation is more complete and it was designed to be used by agents.

Today two scripts already prepare the animation: the page where I draw (the draft board, which I show further down) and the step that matches that drawing against my hand. The rest was still done by the agent in code, and each part becomes a script in the next phase.

script today
sketchsketch-compare

sketch generates the page where I draw over the video, frame by frame, and sketch-compare matches that drawing against the palm of my hand and returns, in numbers, where the element appears and where it settles.

HTML + Node
phase 2
planassetscompositionsoundsrender

The plan is the animation written by AI before it becomes code, the asset catalog keeps the React components ready to reuse, the composition is an HTML page animated with GSAP in HyperFrames, the sounds follow each movement, and the render puts voice, sounds and music together. Today the agent does this in code, and each part becomes a script in phase 2.

AI · React · HyperFrames · ffmpeg
Part 4

The visual editing boards

Visual prompts

Even though the agent could make the animations, I had a really hard time getting custom ones. And I think that's the biggest problem today in generating any content with AI, whether it's video or an asset: being able to ask for something specific.

I've worked on projects where a lot of people ran into this, especially with video generation. The result comes out good, but it's not what you had in mind, because it's hard to make the tool understand something visual through text.

The same thing happened here. The scripts already made a decent video, but I wanted something more specific. So we built a board that works almost like Premiere: it shows the video frame by frame, I draw on top and write instructions for the agent.

Since it's HTML, the board exports exactly which frame I was on, what the instruction was and the drawing. And, believe it or not, AI reads images really well: seeing my drawing, it understood what I wanted better than I had managed to explain with text alone. I called these visual prompts.

The draft board: my drawings over the video, frame by frame (2x).

And there's a bonus: after seeing the drawing, the AI explains in text what it understood, and better than I did. Those explanations become better prompts for the automated scripts of the future. The first draft of this video alone had 123 strokes and 15 notes.

The animation board

Animation was the same story: it got things wrong, and animation is a bit different, because it's more about time. It's less about drawing and more about saying when. So I also built the animation board, which shows the whole video and each animation separately, frame by frame.

There I can set exactly when an animation comes in and when it goes out, step forward and back frame by frame, and point out where it got stuck. In the same way, this produces results I can reuse to build better automated animation scripts.

The animation board: the whole video with the animations on top (2x).

The cut board

The same idea was applied to cutting. The scripts already cut well, but when we stress-tested the project, specific problems showed up: with a breath, or with a sentence where it didn't know whether to cut or not, it cut in the wrong place.

On the cut board, I mark what's wrong on the playhead, the agent changes it and explains what it changed. And every fix becomes a lesson for the scripts themselves: the project improves with every round. That's the concept that interests me most here, and I'll probably bring this kind of board to other stages too.

The cut board: the playhead, the timeline and what the agent changed based on my notes (2x).
Part 5

Results and what's next

The numbers

Today, with the scripts ready, cutting takes minutes, and that already finishes the cleanup stage. Animation is what takes the longest: about 30 minutes for me to draw and adjust. But it's all natural language and drawing by hand, it's pretty easy, and what comes out is very functional and well made.

The idea is that, once those scripts are ready, the whole video takes minutes, and the tests I've been running already point that way. The video at the top is a little over 2 minutes long; of course, the longer the video, the longer it takes, but I believe it'll still be really good. When I was an editor, a video in this style took more than a week.

And off-the-shelf generation tools don't solve it. The result comes out good, but not specific, and when you need something the tool doesn't do, you end up looking for other technologies, which usually cost money.

Every session makes the project better

Every Claude session is valuable. Every session we work in makes the project better, because it holds really good data: what the agent got wrong, what it needed, how it worked and what it ran.

A really interesting exercise is asking the agent itself what difficulties it had running a script or doing a job, where things diverged and where something broke. Working on top of that, the project gets more and more robust. We've been using this a lot, and it works for any project.

What's next

The project has just started and is still far from done, but it's already delivering. The idea is to have a working end-to-end version soon. Right now I'm focused on the animation scripts, which are the ones with the most problems: the body map is already in, and next come the animation plan, the asset catalog, the composition, the sounds and the render.

After that, most of the scripts are already solid and just need a light sharpening: cutting, captions and the rest. With everything ready, the idea is to stress-test the project, find the weak spots and start working on the ideas in the backlog.

And there’s a part I didn’t even show here: the storyboard. Before recording, a whole stage generates the ideas, writes the script and builds a storyboard sheet with a sketch for each shot, and the result turned out impressive. But it has more to do with creating the idea than with editing, so it deserves its own post, once that process is more solid.

And I want to leave a note about Jev, from TypeSafe AI, released this September. It's the first of a new class of models they call System One, a name taken from Kahneman's idea: System 1 is fast, intuitive thinking, and System 2 is slow and deliberate. Instead of writing text word by word, like an LLM, Jev takes the situation and the possible answers, and returns one of them already in the right format, with the probability of each.

According to them, it answers in 70 to 500 milliseconds, costs a fraction of an LLM and, since it only picks among the options you gave it, it has no way of making up an answer outside of them. For this project that's really interesting: several small decisions still go through the conversation with the agent, like whether an attempt stays or goes in the review, or whether a doubtful caption word is right. Those are closed-answer questions, and a model like this could answer them fast, cheap and always in the right format, making the non-deterministic part much smaller. It's still in early access and I haven't tested it, but it's one of the things I most want to bring in here.

LinkIntroducing System One models and Jev ↗
Part 6

The technologies

HyperFrames

HyperFrames is the heart of the animation: it’s what makes everything work. For me, it’s probably the most interesting technology in the project, and the one that surprised me the most.

The idea is simple: each animation is an HTML page, with the motion done in GSAP. To turn it into video, it opens that page in an invisible Chrome, steps through it frame by frame, captures each one and stitches everything together with FFmpeg. Since it’s all HTML, anything a browser can draw becomes video, and the same page always gives the same result.

On top of that, the documentation is amazing. I don’t think I’ve used even 30% of what this library can do, and there’s room for a lot more creativity than I’ve had so far. I’m just getting started, and I think it’s a great starting point for making anything in video.

LinkHyperFrames documentation ↗

The tape: the model that reads sound by sound

The tape was definitely one of the things that surprised me the most. I tried a lot of options to get an exact transcription of what I was saying, and this was the best of them all.

It's the tape I explained in the cutting part: it knows the sounds of what you said, letter by letter, even when you repeat a sentence.

In the stress tests, nothing came close. AssemblyAI, in every configuration I tried, merged the two attempts at a sentence into one, and, of the others I looked into, the only one that marked repetitions wouldn’t even install properly on Windows. And best of all: the tape runs locally, on my own computer, and runs smoothly.

LinkThe model on Hugging Face ↗

Sound: DeepFilterNet3 and DNSMOS

These two take care of cleaning the audio, and they work as a pair. DNSMOS, from Microsoft, is a model that listens to the recording and gives it three scores, the way a person would: one for the voice, one for the background and one overall. Those scores are how the project decides whether there’s noise to remove and, afterwards, whether the cleanup kept the voice intact.

DeepFilterNet3 is what actually removes the noise. It’s an open-source model that separates the voice from the background noise, and it runs on the computer without needing a graphics card. Before it, we used a fixed filter, and it ended up making recordings that were already clean worse.

LinkDeepFilterNet on GitHub ↗

The body: RTMPose, MediaPipe and YuNet

These were the tools that made things a lot cooler. The animations already worked, but I wanted something that could track my movement. I knew it wouldn’t be that hard, because something ready-made probably already existed.

From OpenCV, I used something very specific, YuNet, which finds the face. I know OpenCV has a lot more, and maybe I haven't seen it all.

Here I ran stress tests too: I put eight libraries side by side on the same video, among them Meta’s SAM 2 and CoTracker, YOLO and even Gemini, and these three came out on top. And, as I said before, the most surprising part is that it all runs on my PC smoothly, with no problems at all.

Linkrtmlib on GitHub ↗

The outside services

About costs: everything here runs essentially locally, and for the whole project I only needed a Claude plan: everything was run by Claude Code. But I used two things from outside, AssemblyAI and gpt-image-2. You can make the video without them: without AssemblyAI, it just comes out without captions.

gpt-image-2 was only used for the premium assets. The assets React made were already enough for a very high-quality video, but since I wanted something a bit more personal and polished, I generated a few with it. Just a few, only some windows.

AssemblyAI is used in a single stage, captions, to bring in properly spelled text. It's only for now: I already know better alternatives, I just haven't had time to test them, and I believe you can do the same with a model that costs nothing. We're only using it because we wanted to answer a question quickly and make it work efficiently.