We report quantitative benchmarks, fine-grained category-level analysis, human evaluation, and prompt
engineering results.
Quantitative Results on Camera Motion Editing
We compare our method with four video-to-video baselines: ReCamMaster,
CamCloneMaster, Wan2.1-VACE-14B (Wan2.1), and
Wan2.2-VACE-Fun-A14B (Wan2.2). All baselines are evaluated in a
zero-shot setting. Results are reported in terms of Camera Accuracy,
Structural Alignment, and Visual Quality. Triplets denote
basic motions, advanced motions, and the combined set, respectively.
Fine-grained Results Across Motion Subcategories
The radar chart below provides a more detailed breakdown over selected motion subcategories,
complementing the aggregated benchmark results above.
Per-category analysis over representative motion subtypes beyond the aggregated benchmark metrics.
Subjective Evaluation
We complement automatic metrics with human evaluation on ReCamMaster,
CamCloneMaster, and ours. Volunteers assess camera motion, subject
fidelity, background quality, and temporal consistency, and also provide overall pairwise
preferences. Our method performs best across all four dimensions and is strongly preferred over both
baselines.
Success Rates of Human Evaluation (%)
Preference Rates of Human Pairwise Comparison (%)
Prompt Engineering Study
We study whether richer language instructions improve camera-motion editing. By rewriting concise
prompts into more informative, cinematography-aware descriptions for both training and inference, we
observe consistent gains in camera accuracy and structural alignment,
indicating better camera control and scene-consistent editing.