HITCHCOCK-300K

Hitchcock-300K: A Master-Level Camera Control Dataset for Video Generation and Editing

Anonymous Authors
TL;DR. We present Hitchcock-300K, a large-scale dataset for camera-controlled video editing built with Unreal Engine 5 and designed around professional cinematography principles. Compared with prior datasets, Hitchcock-300K supports more expressive camera motions (e.g., spiral orbit, whip pan), dynamic optical control (e.g., zoom, rack focus, Hitchcock zoom), and broad visual diversity from 1,200+ scenes and 250+ characters. Built on this dataset, we train a text-conditioned video-to-video model that edits videos using natural language camera instructions, achieving more accurate cinematic control, better visual quality, and promising generalization to real-world videos.
Overview
Hitchcock-300K combines expressive camera motion, dynamic optical control, and diverse visual content, providing a strong foundation for cinematic video generation and editing.
Overview of Hitchcock-300K
Dataset Showcase
Below we show a small subset of the dataset for visualization.
Demo Videos
Below we present generalization results on natural-world videos using a model trained on Hitchcock-300K, divided into basic motion and advanced motion cases. Basic motion covers canonical camera controls such as push, pull, pan, and truck, while advanced motion involves more complex camera trajectories and coordinated optical control for more cinematic visual expression.
Basic Motion
Representative results on standard camera controls widely used in video editing.
Boom
Truck
Push
Tilt
Pan
Advanced Motion
More challenging cases that require coordinated control of trajectory and optical effects.
Source Videos
Synthesized Videos
Source Videos
Synthesized Videos
Experimental Results
We report quantitative benchmarks, fine-grained category-level analysis, human evaluation, and prompt engineering results.
Quantitative Results on Camera Motion Editing
We compare our method with four video-to-video baselines: ReCamMaster, CamCloneMaster, Wan2.1-VACE-14B (Wan2.1), and Wan2.2-VACE-Fun-A14B (Wan2.2). All baselines are evaluated in a zero-shot setting. Results are reported in terms of Camera Accuracy, Structural Alignment, and Visual Quality. Triplets denote basic motions, advanced motions, and the combined set, respectively.
Method Camera Accuracy Structural Alignment Visual Quality
RotErr ↓ TransErr ↓ Mat. Pix. ↑ CLIP-V ↑ FID ↓ FVD ↓ CLIP-F ↑
Wan2.1 0.30 / 0.52 / 0.45 3.22 / 0.64 / 1.50 318.88 / 246.21 / 266.87 62.08 / 59.68 / 60.36 93.23 747.38 96.57 / 96.44 / 96.48
Wan2.2 0.29 / 0.49 / 0.42 3.35 / 1.08 / 1.88 5.75 / 3.78 / 4.30 64.90 / 62.45 / 63.10 53.19 298.56 96.15 / 96.29 / 96.26
ReCam 0.09 / 0.31 / 0.24 1.06 / 0.66 / 0.79 314.09 / 161.85 / 205.22 81.79 / 73.81 / 76.09 32.85 190.81 97.21 / 96.75 / 96.88
CamClone 0.09 / 0.28 / 0.21 0.50 / 0.48 / 0.49 505.24 / 283.33 / 347.37 88.40 / 83.72 / 85.07 17.53 124.50 97.39 / 96.64 / 96.86
Ours 0.08 / 0.22 / 0.17 0.26 / 0.30 / 0.28 485.29 / 345.73 / 385.40 89.07 / 86.30 / 87.09 21.87 112.33 97.04 / 96.04 / 96.33
Fine-grained Results Across Motion Subcategories
The radar chart below provides a more detailed breakdown over selected motion subcategories, complementing the aggregated benchmark results above.
Radar chart of fine-grained results
Per-category analysis over representative motion subtypes beyond the aggregated benchmark metrics.
Subjective Evaluation
We complement automatic metrics with human evaluation on ReCamMaster, CamCloneMaster, and ours. Volunteers assess camera motion, subject fidelity, background quality, and temporal consistency, and also provide overall pairwise preferences. Our method performs best across all four dimensions and is strongly preferred over both baselines.
Success Rates of Human Evaluation (%)
Method Motion Subject Background Temporal Average
ReCam 15.99 59.67 62.27 83.46 55.34
CamClone 35.87 67.10 81.41 88.10 68.12
Ours 82.34 85.69 88.66 97.58 88.57
Preference Rates of Human Pairwise Comparison (%)
Sample A Sample B Good Same Bad
ReCam CamClone 5.39% 48.33% 46.28%
Ours ReCam 87.92% 9.48% 2.60%
Ours CamClone 79.00% 9.29% 11.71%
Prompt Engineering Study
We study whether richer language instructions improve camera-motion editing. By rewriting concise prompts into more informative, cinematography-aware descriptions for both training and inference, we observe consistent gains in camera accuracy and structural alignment, indicating better camera control and scene-consistent editing.
Setting Camera Accuracy Structural Alignment Visual Quality
RotErr ↓ TransErr ↓ Mat. Pix. ↑ CLIP-V ↑ FID ↓ FVD ↓ CLIP-F ↑
w/o PE 0.11 / 0.27 / 0.22 0.50 / 0.45 / 0.47 405.70 / 328.05 / 350.12 85.33 / 86.08 / 85.87 21.03 122.84 96.60 / 95.94 / 96.13
w/ PE 0.10 / 0.24 / 0.19 0.41 / 0.34 / 0.36 482.73 / 366.12 / 399.18 88.98 / 86.45 / 87.17 21.42 122.91 97.09 / 96.14 / 96.41
Visual Comparison with Baselines
Existing baselines are mainly limited to basic camera motions and cannot reliably handle source videos with complex camera movement. In contrast, our model trained on Hitchcock-300K supports coordinated control of complex camera trajectories and optical effects, enabling more expressive cinematic editing.
Synthetic Data Domain
Source Video
ReCamMaster
CamCloneMaster
Ours (Hitchcock-300K)
From top to bottom, the five cases correspond to boom, spiral orbit, zoom, Hitchcock zoom, and rack focus.
Natural Data Domain
Source Video
ReCamMaster
CamCloneMaster
Ours (Hitchcock-300K)
From top to bottom, the four cases correspond to rotation, orbit lift, whip pan, and Hitchcock zoom.