Prompt for AI Video Generation
Generate cinematic AI video clips with a detailed, model-ready prompt for Veo, Runway, Kling, Pika and Luma.
Copy-ready prompt
Create a text-to-video prompt for [tool: Veo/Runway/Kling/Pika/Luma]. Subject and action: [who or what, and exactly what they do]. Setting and time of day: [location, weather, morning/golden hour/night]. Camera framing and movement: [wide/medium/close-up; static, slow dolly-in, orbit, tracking, handheld]. Lighting and color: [soft daylight, neon, warm tungsten; color palette or mood]. Style and rendering: [cinematic live-action, 35mm film, 3D animation, anime, hyperreal]. Aspect ratio: [16:9 / 9:16 / 1:1]. Duration: [seconds]. Write it as one flowing, specific paragraph. Front-load the subject and action, then layer setting, camera, lighting, and style. Avoid vague adjectives; name concrete visual details.
Want a version tailored to you?
Answer a few quick questions and the AI Video Prompt Generator builds a custom prompt from your exact details.
🎬 Open the AI Video Prompt GeneratorWhy specific visual language beats vague prompts
Text-to-video models are only as good as the picture you paint for them. A prompt like "a cool city scene" gives the model almost nothing to anchor on, so it fills the gaps with whatever is statistically average — a nondescript street, flat lighting, an aimless camera. The result feels generic because the instructions were generic. The template above fixes this by making you commit to concrete choices. Instead of "a person walking," you write "a woman in a red raincoat walking briskly across a wet crosswalk." Instead of "nice lighting," you name "warm streetlight reflecting off puddles at dusk." Each specific noun and verb narrows the space of possible outputs toward the shot you actually have in mind. The single most reliable upgrade you can make to any video prompt is to replace abstract adjectives with observable visual details a cinematographer could point a camera at.
Camera and motion cues are what make it feel filmed
The difference between a clip that looks like a moving photograph and one that looks like real footage is almost always the camera. AI video models respond strongly to framing and movement language, so the template asks you to state both. Framing sets the scale: a wide shot establishes a location, a medium shot shows body language, a close-up delivers emotion or product detail. Movement gives the shot life: a slow dolly-in builds intimacy, an orbit reveals dimension, a tracking shot follows action, and handheld adds documentary energy. Naming these explicitly — "slow push-in on the subject's face" rather than leaving it unspecified — tells the model how the frame should evolve over the clip. Keep the motion simple and singular for short durations; asking for a dolly, an orbit, and a crane move in one three-second clip usually produces a confused, warping result.
Lighting, mood, and color carry the emotion
Lighting is the fastest way to set tone, and it is the detail most people leave out. The same subject reads completely differently under soft overcast daylight, hard midday sun, warm tungsten interior light, or cold neon. The template asks you to name the light source and quality, plus the time of day, because "golden hour" and "blue hour" and "high noon" each imply a whole palette and shadow behavior the model already understands. Color direction reinforces the mood: a warm amber palette feels nostalgic, teal-and-orange feels cinematic and modern, desaturated greys feel bleak. State the mood you want in visual terms — "moody and cinematic with deep shadows" — rather than emotional abstractions the model cannot render directly. When lighting, palette, and time of day agree with each other, the clip reads as a deliberate, coherent shot.
Iterate one variable at a time, and know your tool
Video generation is inherently a loop, not a one-shot. When a clip is close but not right, resist the urge to rewrite the whole prompt. Change one variable — swap the camera movement, or the time of day, or the lens feel — and regenerate, so you can see what that single change did. This disciplined approach turns a slot machine into a controllable process and gets you to the intended shot far faster. It also helps to remember that the tools differ. Google Veo leans toward realistic physics and coherent audio; Runway gives fine control and strong stylization; Kling handles complex motion and longer takes well; Pika is fast and flexible for stylized clips; and Luma excels at smooth, dreamlike camera moves. A prompt tuned for one may need small adjustments on another, so treat the first generation as a draft and iterate deliberately toward the result.
Why this prompt works
AI video tools turn text into motion, so vague prompts produce generic, drifting clips. This template forces you to specify subject, camera, lighting, style, aspect ratio, and duration up front — the exact cues the model needs to render a shot that looks intentional rather than random.
How to customize it
- Front-load the subject and action, then layer setting, camera, and light.
- Use one clear camera movement per short clip to avoid warping.
- Iterate by changing a single variable so you can see what each change does.
Example output
Sample onlyGoal: A cinematic establishing shot of a lone hiker at sunrise (Veo, 16:9, 5 seconds).
Prompt:
A lone hiker in a weathered orange jacket stands on a rocky ridge at sunrise, mist filling the valley below. Wide cinematic shot with a slow dolly-in toward the figure. Soft golden-hour light rakes across the rocks, warm amber palette against cool blue shadows. 35mm film look, shallow depth of field, gentle grain. Aspect ratio 16:9, 5 seconds.
What it produces: A steady, filmic clip that opens on the landscape and slowly pushes toward the hiker as warm light catches the ridge. Because the subject, camera move, lighting, palette, and film style are all named, the model renders a coherent establishing shot instead of a drifting, generic mountain scene.
Note: The single slow dolly-in keeps motion clean over five seconds; adding an orbit or crane on top would likely warp the geometry.
Prompt variations to try
Vertical social clip
Create a 9:16 text-to-video prompt for [tool] built for a social feed. Subject and action: [what happens]. Keep it punchy and full-frame with the subject centered for vertical viewing. Setting and time of day: [where, when]. Camera: mostly static or a subtle push-in so the subject stays readable on a phone. Lighting and color: [bright and clean, or moody neon]. Style: [live-action/animation]. Aspect ratio 9:16, [3-6] seconds. Write it as one specific paragraph and front-load the most eye-catching moment for the first second.
Product B-roll shot
Write a text-to-video prompt for [tool] for a clean product B-roll clip. Product: [name and key visual features]. Camera: slow orbit or macro dolly-in that reveals the product's shape and texture. Setting: minimal surface with soft studio lighting and a shallow background. Lighting and color: soft key light with a gentle rim highlight, [palette]. Style: hyperreal commercial photography in motion. Aspect ratio [16:9 or 1:1], [4-6] seconds. Keep the motion single and smooth so the product stays sharp and undistorted.
Animated / anime style
Write a text-to-video prompt for [tool] in a [2D anime / 3D animated] style. Subject and action: [character and what they do]. Setting and time of day: [where, when]. Camera framing and movement: [wide establishing, then a slow pan]. Lighting and color: [describe the palette and mood, e.g. warm sunset tones with soft cel shading]. Style and rendering: [Studio-Ghibli-like hand-painted look / clean modern 3D]. Aspect ratio [16:9], [4-6] seconds. Emphasize expressive character motion and a consistent art style throughout the clip.
Common mistakes to avoid
- Writing vague adjectives instead of visual details. "Beautiful cinematic scene" gives the model nothing to render. Name the subject, action, and light in concrete, observable terms.
- Leaving out the camera. Without framing and movement cues the clip feels like a moving photo. State a shot size and a single motion such as
slow dolly-inororbit. - Stacking too many camera moves in a short clip. A dolly plus an orbit plus a crane in three seconds warps the image. Use one clear movement per short shot.
- Forgetting aspect ratio and duration. A 16:9 prompt looks wrong cropped to a vertical feed. Always set the ratio and length to match where the clip will be used.
- Expecting one prompt to work identically across tools. Veo, Runway, Kling, Pika, and Luma differ. Treat the first output as a draft and iterate one variable at a time.
Frequently asked questions
Which AI video tool should I use this prompt with?
It works with any modern text-to-video model. Google Veo leans realistic with coherent audio, Runway offers fine control and stylization, Kling handles complex motion, Pika is fast and flexible, and Luma gives smooth camera moves. Start with whichever you have access to and adjust small details for that tool.
How long should an AI-generated clip be?
Short clips of three to six seconds are the most reliable, because the model has to keep motion and geometry stable across every frame. For longer sequences, generate several short shots and edit them together rather than asking for one long take, which is more likely to drift or warp.
Why does my video look warped or morphing?
Usually the prompt asked for too much motion at once, or the subject and camera move fought each other. Simplify to one clear camera movement, keep the action singular, and shorten the duration. Regenerating with a single simpler motion almost always stabilizes the result.
Do I need a reference image?
Not necessarily. A detailed text prompt can produce strong results on its own. That said, many tools accept a starting image, and using one gives you tighter control over the subject and framing. If consistency matters, generate or choose an image first, then animate it with a camera and motion prompt.
Tip: replace the parts in [square brackets] with your own details before you send. The more specific you are — audience, tone, goal, constraints — the better the AI output.