Video AICinematic ShotsAdvanced180 minSaves 2+ hours

Cinematic Diner Booth Two-Shot: Intimate Dialogue Scene

For AI filmmakers, generate a detailed cinematic two-shot of two characters in a diner booth, capturing subtle dialogue and emotional depth with specific camera, lighting, and sound cues for a film-look.

This prompt helps AI filmmakers create a nuanced, intimate two-shot scene set in a diner booth. It specifies camera movement, lens choice, lighting, wardrobe, and score cues to achieve a character-driven, atmospheric film look, ensuring emotional depth and visual storytelling.

READY-TO-USE PROMPT

Copy Prompt

prompt.txt
Role: You are a seasoned film director preparing a scene for a character-driven documentary.

Context: The scene is an intimate two-shot in a classic diner booth. Two individuals are mid-conversation, the atmosphere is subdued, and the focus is on subtle emotional exchange. Lighting is practical and warm.

Task: Develop a comprehensive cinematic scene prompt for a video AI model. Detail every aspect necessary to achieve a specific film aesthetic: shot list, camera movement, lens choice, lighting setup, wardrobe, and a specific score cue. The output should evoke a deep, character-driven mood.

Constraints:
*   **Frame Rate:** Specify 24 frames per second (fps).
*   **Film Look:** Emphasize shallow depth of field, naturalistic lighting, and a slightly desaturated color palette.
*   **Focus Pull:** Include a specific rack focus instruction.
*   **Dialogue Driven:** The visual and auditory elements must support an intimate, reflective conversation.
*   **Atmosphere:** Prioritize a sense of quiet intensity and connection.
*   **No CGI Overlays:** All elements should appear organically within the scene.
*   **Length:** The generated scene prompt should be detailed enough to convey the precise vision.

Output:
Provide a cinematic scene prompt structured as follows:

**Scene Title:** "Diner Confession"
**Location:** Classic American Diner, night. Booth seating.
**Characters:**
*   **Character A:** {{character_a_description}}
*   **Character B:** {{character_b_description}}
**Dialogue Snippet:** "{{dialogue_snippet}}"

**1. Shot List & Camera Movement:**
*   **Opening:** Medium close-up (MCU) on Character A, slightly off-center, framed by the booth. Camera very slowly pushes in, almost imperceptibly.
*   **Mid-scene:** As Character A finishes speaking their line, a deliberate, smooth rack focus shifts attention from Character A to Character B, who is listening intently. The camera subtly re-frames to an MCU on Character B.
*   **Closing:** Hold on Character B's reaction, a slight, almost imperceptible nod or change in expression. Camera slowly pulls back slightly, revealing a bit more of the booth environment, but maintaining focus on Character B.

**2. Lens & Depth of Field:**
*   **Lens:** Prime lens, 50mm equivalent (full-frame).
*   **Depth of Field:** Very shallow, isolating the characters from the background and foreground elements. Bokeh should be soft and creamy.

**3. Lighting:**
*   **Key Light:** Warm, practical light emanating from a single hanging pendant lamp directly above the booth, casting soft shadows.
*   **Fill Light:** Minimal, naturalistic fill from ambient diner lighting, just enough to lift shadows without flattening the image.
*   **Backlight:** Soft, subtle backlighting from distant window reflections or neon signs outside, adding separation.
*   **Overall Mood:** Warm, intimate, slightly melancholic. High contrast in texture, low contrast in color.

**4. Wardrobe:**
*   **Character A:** Earth tones, slightly worn but clean, casual jacket or sweater. (e.g., dark olive green, muted brown).
*   **Character B:** Cool tones, simple, understated shirt or blouse. (e.g., slate grey, deep navy).
*   **Texture:** Soft, natural fabrics. Avoid reflective or distracting patterns.

**5. Score Cue:**
*   **Music:** Sparse, melancholic piano chords, low in the mix. Underscore, not overt. A single, sustained cello note enters subtly during the rack focus.
*   **Sound Design:** Gentle hum of the diner, clinking of distant cutlery (very subtle), muffled street sounds outside. Prioritize silence and the characters' breathing.

**6. Technical Specifications:**
*   **Frame Rate:** 24fps
*   **Aspect Ratio:** 2.35:1 (Cinemascope)
*   **Color Grade:** Slightly desaturated, with warm highlights and cool shadows. Film grain emulation.

Estimated results

DifficultyAdvanced
Setup time180 min
Time saved2+ hours
Best modelsVeo, Kling, Runway
Best audienceFilm Production, Documentary Filmmaking

Editor's note

Why this prompt matters

Crafting nuanced, character-driven scenes with AI video models presents a distinct challenge. Directors often struggle to translate specific cinematic visions—like the subtle interplay of light and shadow, precise camera movement, or the emotional weight of a dialogue exchange—into prompts that AI models can interpret accurately. This workflow addresses that gap, providing a structured approach to generate intimate, film-look scenes.

It is designed for AI filmmakers and documentarians who prioritize atmosphere and emotional depth over flashy effects. When your project demands a specific visual and auditory texture for a quiet moment, rather than a generic shot, this framework becomes essential. It helps ensure that the AI understands not just *what* is in the shot, but *how* it should feel and function within the broader narrative. Use this when you have a clear directorial intent for a scene's mood, pacing, and technical execution, and need to communicate that vision with precision to a generative video model.

Anatomy

Prompt engineering breakdown

Role

You are a seasoned film director preparing a scene for a character-driven documentary.

Context

The scene is an intimate two-shot in a classic diner booth. Two individuals are mid-conversation, the atmosphere is subdued, and the focus is on subtle emotional exchange. Lighting is practical and warm.

Goal

Develop a comprehensive cinematic scene prompt for a video AI model. Detail every aspect necessary to achieve a specific film aesthetic: shot list, camera movement, lens choice, lighting setup, wardrobe, and a specific score cue. The output should evoke a deep, character-driven mood.

Constraints

* **Frame Rate:** Specify 24 frames per second (fps). * **Film Look:** Emphasize shallow depth of field, naturalistic lighting, and a slightly desaturated color palette. * **Focus Pull:** Include a specific rack focus instruction. * **Dialogue Driven:** The visual and auditory elements must support an intimate, reflective conversation. * **Atmosphere:** Prioritize a sense of quiet intensity and connection. * **No CGI Overlays:** All elements should appear organically within the scene. * **Length:** The generated scene prompt should be detailed enough to convey the precise vision.

Output format

Provide a cinematic scene prompt structured as follows: Scene Title, Location, Characters (A & B descriptions), Dialogue Snippet, Shot List & Camera Movement, Lens & Depth of Field, Lighting, Wardrobe, Score Cue, Technical Specifications (Frame Rate, Aspect Ratio, Color Grade).

Why this structure works

The prompt effectively uses role priming to align the AI as a film director, ensuring the output reflects a professional cinematic perspective. Explicit constraints, such as 'no CGI overlays' and specific film look instructions, guide the AI away from undesirable elements. The structured output section provides a clear template, ensuring all necessary details for a video AI model are included and organized.

Pick your version

Prompt variations

BeginnerWorks with any model

For users new to video AI or filmmaking, focusing on foundational visual and atmospheric elements without overly technical specifications.

prompt.txt
Imagine you're making a short film. Create a video scene of two people talking in a diner booth. The light is warm and comes from a lamp above them. Show Character A ({{character_a_description}}) speaking, then slowly shift the camera's focus to Character B ({{character_b_description}}) as they listen to the dialogue: "{{dialogue_snippet}}". Use a normal lens, so the background is a bit blurry. The overall look should feel like an old movie, slightly faded colors, and run at 24 frames per second. Dress characters in simple, comfy clothes (like a jacket or shirt). Add quiet, sad piano music in the background, with diner sounds.
ProfessionalBest with veo

When precise, detailed control over every cinematic element is required, for experienced filmmakers and directors aiming for a specific aesthetic.

prompt.txt
As a film director, create a detailed cinematic scene for a video AI model. The scene is an intimate two-shot in a classic American diner booth at night, focusing on Character A ({{character_a_description}}) and Character B ({{character_b_description}}) during a reflective conversation: "{{dialogue_snippet}}". Specify a 24fps frame rate and 2.35:1 aspect ratio. The shot list begins with an imperceptible slow push into an MCU on Character A. A deliberate rack focus then shifts to Character B's reaction, followed by a subtle pull-back. Use a 50mm equivalent prime lens with very shallow depth of field for soft bokeh. Lighting should be practical, warm, from a hanging pendant, with minimal naturalistic fill. Wardrobe: natural fabrics, earth tones for A, cool tones for B. Score: sparse, melancholic piano with a sustained cello during the rack focus, subtle diner ambience.
Short VersionBest with runway

For quick iterations, rapid prototyping, or when working with models that respond best to concise, direct instructions.

prompt.txt
Generate a 24fps video scene: two individuals in a dimly lit diner booth, mid-conversation. Character A ({{character_a_description}}) speaks, then a slow rack focus shifts to Character B ({{character_b_description}}) reacting to "{{dialogue_snippet}}". Use a 50mm equivalent lens with shallow depth of field. Lighting is warm, practical from a pendant lamp, creating a soft, melancholic mood. Wardrobe in natural, muted tones. Include subtle, sparse piano music and quiet diner ambience. Aim for a desaturated film look with film grain.
EnterpriseBest with kling

In large-scale productions or studio environments where adherence to brand guidelines, metadata tagging, and compliance checks are critical alongside creative execution.

prompt.txt
As the lead director for '{{project_title}}', generate a video asset for an intimate diner scene, adhering to established brand and creative guidelines. The scene features Character A ({{character_a_description}}) and Character B ({{character_b_description}}) engaged in a reflective dialogue ({{dialogue_snippet}}). Ensure a 24fps, 2.35:1 aspect ratio, and 'Classic Film Look' color grade. Implement an initial MCU on Character A with a slow push, followed by a precise rack focus to Character B, and a subtle pull-back. Mandate a 50mm equivalent prime lens for shallow depth of field and soft bokeh. Lighting must be warm, practical (pendant lamp), with minimal fill. Wardrobe choices (earth/cool tones, natural fabrics) require compliance approval. Score cue: sparse, melancholic piano with a cello entry during the focus shift. All generated elements must be tagged for intellectual property and asset management systems. Avoid any non-approved branding or visual distractions.

What you'll get

Expected output

Scene Title: "Diner Confession" Location: Classic American Diner, night. Booth seating. Characters:

  • Character A: Elara, mid-40s, a documentary photographer. Wears her experiences on her face, but with a quiet resilience.
  • Character B: Marcus, late 30s, a former student of Elara's, now a journalist. Observant, a touch guarded.

Dialogue Snippet: "Elara: 'Sometimes, you just have to trust the light will find the story, even when you can't see it yet.' Marcus: 'And what if the story's in the dark?'"

1. Shot List & Camera Movement:

  • Opening: Medium close-up (MCU) on Elara, slightly off-center, framed by the booth. Camera very slowly pushes in, almost imperceptibly.
  • Mid-scene: As Elara finishes speaking their line, a deliberate, smooth rack focus shifts attention from Elara to Marcus, who is listening intently. The camera subtly re-frames to an MCU on Marcus.
  • Closing: Hold on Marcus's reaction, a slight, almost imperceptible nod or change in expression. Camera slowly pulls back slightly, revealing a bit more of the booth environment, but maintaining focus on Marcus.

2. Lens & Depth of Field:

  • Lens: Prime lens, 50mm equivalent (full-frame).
  • Depth of Field: Very shallow, isolating the characters from the background and foreground elements. Bokeh should be soft and creamy.

3. Lighting:

  • Key Light: Warm, practical light emanating from a single hanging pendant lamp directly above the booth, casting soft shadows.
  • Fill Light: Minimal, naturalistic fill from ambient diner lighting, just enough to lift shadows without flattening the image.
  • Backlight: Soft, subtle backlighting from distant window reflections or neon signs outside, adding separation.
  • Overall Mood: Warm, intimate, slightly melancholic. High contrast in texture, low contrast in color.

4. Wardrobe:

  • Character A: Earth tones, slightly worn but clean, casual jacket or sweater. (e.g., dark olive green, muted brown).
  • Character B: Cool tones, simple, understated shirt or blouse. (e.g., slate grey, deep navy).
  • Texture: Soft, natural fabrics. Avoid reflective or distracting patterns.

5. Score Cue:

  • Music: Sparse, melancholic piano chords, low in the mix. Underscore, not overt. A single, sustained cello note enters subtly during the rack focus.
  • Sound Design: Gentle hum of the diner, clinking of distant cutlery (very subtle), muffled street sounds outside. Prioritize silence and the characters' breathing.

6. Technical Specifications:

  • Frame Rate: 24fps
  • Aspect Ratio: 2.35:1 (Cinemascope)
  • Color Grade: Slightly desaturated, with warm highlights and cool shadows. Film grain emulation.

Under the hood

Why this prompt works

This prompt generates precise cinematic scene descriptions because it employs several effective prompt engineering techniques. Role priming establishes the AI's persona as a "seasoned film director," which guides its understanding and output toward professional filmmaking terminology and considerations rather than generic descriptions. The use of explicit constraints is critical, detailing requirements such as a specific frame rate (24fps), a "film look," and a precise "rack focus instruction." These constraints prevent the AI from defaulting to common, less specific outputs, ensuring the generated scene adheres to a director's exact vision.

Furthermore, the structured output ensures comprehensive coverage of all essential filmmaking elements: shot list, camera movement, lens, lighting, wardrobe, and score cue. This methodical breakdown forces the AI to consider each component individually, then integrate them into a cohesive scene. This multi-faceted approach, combining a specific role, granular constraints, and a clear output structure, yields a far more detailed and actionable cinematic blueprint than a simple, open-ended request.

Model fit

Best AI models for this prompt

Veo

Veo excels at generating highly detailed and atmospheric scenes, making it suitable for nuanced character work. Its ability to interpret complex lighting cues and deliver a consistent film look is a significant strength. However, achieving precise, subtle rack focus transitions can sometimes require multiple iterations. See the full Veo hub for deeper guidance.

Kling

Kling demonstrates strong capabilities in producing realistic human characters and environments, which is crucial for an intimate diner scene. It handles practical light sources effectively, contributing to a believable filmic aesthetic. Its current limitation often lies in maintaining very specific, long-duration camera movements without slight deviations. See the full Kling hub for deeper guidance.

Runway

Runway is adept at generating diverse visual styles and offers good control over scene composition and mood. It can interpret descriptive prompts well to establish a coherent atmosphere. The challenge with Runway can be consistently rendering very fine details in facial expressions or maintaining subtle, slow camera pushes across longer durations. See the full Runway hub for deeper guidance.

When to use

  • When creating intimate, character-focused documentary scenes that emphasize emotional exchange.
  • For dialogue-heavy sequences where subtle reactions and internal states are central to the narrative.
  • When a specific film aesthetic with shallow depth of field, warm practical lighting, and a melancholic mood is desired.
  • To generate a specific two-shot with a deliberate rack focus that enhances the emotional arc of a conversation.
  • When working with AI video models like Veo, Kling, or Runway that respond well to detailed cinematic instructions.

When not to use

  • For fast-paced action sequences or scenes requiring dynamic, handheld camera work.
  • When the primary goal is to showcase elaborate sets or broad environmental context rather than character intimacy.
  • If a bright, high-key lighting scheme or a vibrant, saturated color palette is the desired aesthetic.
  • For scenes where dialogue is incidental and visual spectacle or complex CGI elements are the focus.
  • When an overt, dramatic musical score is intended to drive the scene's emotional impact.

Get more from it

Pro tips

  • 1

    To prevent generic outputs, specify character backstory and relationship dynamics concisely; this informs subtle gestures and expressions.

  • 2

    For consistent lighting, describe the specific type and color temperature of the practical lamp, ensuring it feels organic to the diner setting.

  • 3

    Avoid flat images by detailing subtle camera drifts and re-frames, not just static holds, to enhance emotional pacing.

  • 4

    To ensure emotional resonance, provide dialogue snippets that hint at underlying tension or revelation, guiding character performance.

  • 5

    Prevent wardrobe clashes by specifying textures and exact color palettes for both characters, aligning with the scene's subdued mood.

  • 6

    To avoid generic soundscapes, describe specific ambient diner sounds and their volume, prioritizing character breathing and natural quiet.

  • 7

    Ensure the rack focus feels natural; specify the exact moment in the dialogue or character action it should occur for narrative impact.

Don't ship this

Common mistakes

  • Omitting specific character details, leading to generic portrayals lacking emotional depth and distinctiveness.

    Fix — Provide 2-3 defining traits for each character and their relationship, informing their demeanor and interaction within the scene.

  • Vague lighting descriptions resulting in flat, uninspired visuals that miss the intended warm, intimate mood.

    Fix — Detail the practical light source's position, intensity, and color temperature for realistic ambiance and mood creation.

  • Generic camera movements lacking emotional intent, failing to guide the viewer's focus or enhance the dialogue.

    Fix — Specify subtle pushes, pulls, and re-frames tied directly to character emotional beats, making the camera an active participant.

  • Choosing overly complex dialogue that distracts from visual subtlety and the intended intimate character study.

    Fix — Keep dialogue concise and understated. Allow the visuals, sound design, and character reactions to carry much of the scene's weight.

  • Neglecting sound design, leaving a sterile or unfitting audio track that undermines the scene's atmospheric goals.

    Fix — Include specific ambient sounds like distant cutlery or street noise at very low levels, reinforcing the diner's intimacy.

  • Inconsistent wardrobe choices that conflict with the scene's subdued, melancholic tone or character descriptions.

    Fix — Ensure wardrobe aligns with character backstory and the scene's mood, favoring soft, natural fabrics and muted colors.

People also ask

Frequently asked questions

Q.Can I change the location from a diner to another intimate setting?

Yes, but maintain the intimate, confined setting. Adapt lighting and sound design to match the new environment while preserving the two-shot focus and emotional tone. Examples include a quiet bar booth or a living room couch.

Q.How specific should the character descriptions be for optimal results?

Provide 2-3 key personality traits and a brief relational context. Avoid overly long backstories; focus on details influencing their on-screen presence and interaction, such as 'hesitant' or 'world-weary'.

Q.Will this prompt work for non-documentary fiction films, such as a drama?

Absolutely. While framed for documentary, its focus on character intimacy, nuanced performance, and specific cinematic craft is highly effective for any fiction dialogue scene, particularly dramas requiring emotional depth.

Q.How long should the dialogue snippet be to effectively guide the AI?

A brief exchange, 1-3 lines per character, is usually sufficient. The goal is to hint at the conversation's depth and emotional subtext, not to script a full scene. Focus on impact, not length.

Q.Can I use different lens choices than the suggested 50mm equivalent?

While 50mm is suggested for intimacy and a natural perspective, a slightly wider 35mm or tighter 85mm prime can work. Ensure your choice still supports shallow depth of field for character isolation.

Q.What if my AI model doesn't support the 2.35:1 aspect ratio?

If 2.35:1 is not directly supported, use the closest available widescreen ratio, such as 16:9. You can also mention 'cinematic framing' in the prompt to guide the model toward a similar aesthetic.

Q.How do I ensure the rack focus is effective and noticeable in the generated video?

Clearly state the precise trigger for the rack focus (e.g., a specific dialogue word, a character's subtle action) and the exact subject it shifts to. Emphasize 'smooth' and 'deliberate' for clarity.

Version 1.0Last reviewed July 13, 2026
Reviewed by PromptInFlow Editorial Team