tried a open-source Claude Code skill called vox-director this weekend.
The test was a 30 second landscape explainer about why Sichuan food has become more popular outside China. I mostly wanted to see if one workflow could handle the outline, visuals, motion, voice and final assembly without me moving files between like five different tools.
The first step was a six-part beat map. That was probably the most useful part of the whole run because it paused before generating the video and let me check the structure first.
I approved the outline, picked a Chinese ink collage direction from four visual options, and then just let the rest of it run. vox-director used the Atlas Cloud API for the generation steps and ffmpeg for putting everything together.
My first run ended up making ten short shots and a video around 30 seconds long. The generation part took maybe 15 minutes on my setup, give or take.
The result was useful, but definitely wasnt ready to publish untouched.
The opening hook used a dramatic restaurant statistic that I couldnt actually verify. One of the later lines also made the connection between spicy food and endorphins sound way more certain than it really is, so both of those would need rewritten before I used the video anywhere.
A couple transitions were awkward too. One of the shots looked fine on its own, but didnt really connect with the narration, and the middle part moved way faster then the rest.
The Chinese ink style worked better than I expected though. The other options looked kinda like generic tech explainers, while that one at least felt connected to the subject.
My main takeaway is that vox-director seems more useful as an orchestration and first draft system than a one click finished video generator. Having the outline, images, motion, narration, music and assembly in one workflow saved me a lot of jumping around, but the facts and pacing still needed a human looking at it.
I’m attaching the first output instead of only showing a cleaned up version, because honestly the mistakes are probably more useful for talking about how the workflow actually works.
For people building multi-model generation pipelines, where do you normally put the fact checking step?
Do you check the script before generating any visuals, or let the whole first draft finish and review everything after?