Skip to main content
  1. Blog/

Veo 3.1 vs. Jimeng: Stress-Testing an Image-to-Video AI Workflow

·1 min
Author
Chengyu
I’m Chengyu — a final-year Computer Science student at the University of Sydney. I write about the things I build and break, plus hiking, travel, gaming, and gadgets.

Today I tried out a new AI video workflow: first generate a set-environment image with an image-generation tool, then use Veo 3.1’s image-to-video bridge feature to generate a transition clip between two photos.

One of the generated transition frames

Honestly, the final result still has room to improve on smoothness. The main reason is that Veo 3.1’s generation quota is pretty limited, so I couldn’t do repeated fine-tuning or re-rolls — I just had to accept a transition that came out a little rough around the edges.

Another frame from the Veo 3.1 result

That said, there’s no appreciation without comparison. Out of curiosity, I ran the exact same prompt through Jimeng (即梦), and the result nearly gave me a heart attack — the character’s head did a full 180-degree spin in the generated video. For a second it felt like something out of a horror movie; watching that late at night genuinely startled me.

The considerably less reassuring Jimeng result

Takeaway: at this point, Veo 3.1 is clearly more reliable than Jimeng when it comes to understanding physics and human anatomy. It burns through GPU time and quota fast, but at least it doesn’t turn your film set into a horror movie.

Related