LocalBanana
© 2026 LocalBananaFollow us on X

January 30, 2026·Tutorials·30 min read

The Six-Slot Method for Writing AI Image Prompts

Learn how to express visual intent before writing prompts: references, mood, composition, output format, prompt structure, and practical iteration.

The Six-Slot Method for Writing AI Image Prompts

On this page

  • The problem this guide solves
  • How long should an AI image prompt be?
  • The six slots every prompt should fill
  • Prose or structured slots?
  • How to write the light slot
  • How to write the camera slot
  • How to route a prompt to the right model
  • What to do when the output is wrong
  • A skeleton you can copy
  • FAQ
  • How long should an AI image prompt be?
  • What should an AI image prompt include?
  • Should I write AI image prompts in JSON?
  • Do negative prompts work in AI image generation?
  • Which model should I use for which kind of prompt?
On this page
  • The problem this guide solves
  • How long should an AI image prompt be?
  • The six slots every prompt should fill
  • Prose or structured slots?
  • How to write the light slot
  • How to write the camera slot
  • How to route a prompt to the right model
  • What to do when the output is wrong
  • A skeleton you can copy
  • FAQ
  • How long should an AI image prompt be?
  • What should an AI image prompt include?
  • Should I write AI image prompts in JSON?
  • Do negative prompts work in AI image generation?
  • Which model should I use for which kind of prompt?

The problem this guide solves

You describe an image, the model returns something adjacent to it, and you have no idea which word to change. That is not a talent gap — it is a missing slot in the prompt.

This guide is built on a count rather than on habit. We measured 9,599 published prompts in the LocalBanana gallery — every item with a prompt longer than 20 characters, each stored with the model that ran it and the image it produced — and turned the patterns into a writing procedure you can follow line by line.

83

median words per prompt in the corpus

mean 128; p25 = 36, p75 = 174, p90 = 296

80

words — where median engagement steps up

median views go 5 → 19 across that boundary

13%

of prompts are written as labelled slots, not prose

1,249 of 9,599, and they behave differently

How to read the numbers in this guide

Word frequencies are exact counts over the corpus. Anything described as engagement uses on-site view counts, which are also affected by how long an item has been published and where the feed surfaced it — so every engagement comparison here is correlation, not proof of causation. Where a comparison had an obvious confound, we controlled for it and say so. The full analysis is in what 9,667 AI prompts actually contain.

How long should an AI image prompt be?

The median prompt is 83 words. The mean is 128, so a long tail of people write far more than the middle.

Prompt lengthShare of corpusMedian views
Under 30 words20% (1,966)4
30–79 words28% (2,670)5
80–149 words22% (2,129)19
150+ words30% (2,834)39

The interesting part is that the rise is not gradual. Nothing much happens between 30 and 79 words. Then median views roughly quadruple between the 30–79 band and the 80–149 band.

The mechanism is not length. It is what fits. Below about 80 words you can name a subject and one or two adjectives. Above it you have room for subject, setting, light, lens and mood in the same breath — which is the point at which the model stops filling gaps with its own defaults.

✕ Bad

A moody portrait of a woman, cinematic, high quality, 8k

✓ Good

Moody cinematic portrait of a young woman, barely-there makeup, loose dark brown hair, fitted black turtleneck. Dark minimal background fading into black. Gentle diffused front light, soft shadows, subtle film grain, organic color grading, shallow depth of field, high-end editorial feel.

The second version is a real gallery prompt of roughly 60 words. It sits below the 80-word threshold and still works, because every word is a specification rather than a compliment. Words like cinematic, high quality and 8k cost you length and buy nothing; dark background fading into black and diffused front light decide the picture.

Orange Sunglasses Selective Color Portrait

Orange Sunglasses Selective Color Portrait

Nano Banana Pro
Prompt▾
A studio-style close-up editorial portrait of a person with strong, well-defined facial features and slightly imperfect, natural skin texture. The subject wears a black tailored turtleneck with sharp, clean lines, layered under a high-collared black jacket in a minimalist contemporary fashion style.

The subject wears semi-transparent orange acetate sunglasses — rectangular frames with softly rounded edges, glossy finish, and amber gradient lenses — serving as the only colored element in the image.

Color concept: selective color photography — monochrome black-and-white image with only the sunglasses in vivid orange.

Mood is calm and confident, serious expression, direct gaze into the camera.

Lighting is soft frontal studio light with gentle shadows, even skin tones, cinematic contrast, and visible natural skin texture. Shot on a professional portrait camera, f/2.0, ISO 100, 1/125s. High resolution, ultra-sharp focus on the face.

Style: editorial luxury fashion portrait, photorealistic, professional studio photography, no illustration, no painterly effects.
Prose, but every clause is a specification — wardrobe, collar height, skin texture, background separation.Open in LocalBanana

The six slots every prompt should fill

Read enough high-performing prompts in the corpus and the same skeleton keeps appearing. Written out, it is six questions:

SlotThe question it answersWhat the corpus shows
SubjectWho or what, in physical detailAlways present — this is the slot nobody forgets
SettingWhere they are, and what is behind themFrequently missing; background is the most common accidental artefact
LightWhere the light comes from and how hard it isstudio lighting leads at 6.5%; most prompts name no light at all
CameraFocal length, aperture, angle, distancedepth of field 14.5%, 85mm 4.9%, f/1.8 only 2.0%
Mood and styleThe emotional register and the reference mediumUsually present, often as the only non-subject slot
ExclusionsWhat must not appearRare — and the cheapest fix in this whole guide

Here is a real gallery prompt that fills all six, reformatted so the slots are visible:

{
  "subject": "young woman, voluminous messy long hair, black sporty crop top
               with white trim, sitting perched on a high dark wooden counter,
               leaning forward, sultry nonchalant expression",
  "setting": "white wall with vintage posters taped up, brown wooden blinds,
              electric guitar neck and liquor bottles in the foreground",
  "composition": "full body, 35mm focal length, slightly low angle,
                  hard flash fall-off, messy but balanced framing",
  "lighting": "direct on-camera flash, hard light, sharp drop shadow on the
               wall behind her, high contrast, no diffusion",
  "color_palette": "vintage film aesthetic, Kodak Gold 200 simulation, warm
                    wood tones against cool white flash, deep blacks, grainy",
  "mood": "indie sleaze, candid, rebellious, snapshot style",
  "negative_prompt": "soft lighting, studio lighting, bokeh, airbrushed skin,
                      3d render, illustration, over-processed, HDR"
}
Indie Sleaze Cool Girl Flash Portrait

Indie Sleaze Cool Girl Flash Portrait

Nano Banana Pro
Prompt▾
{
  "portrait_prompt": {
    "subject": "Based on <User Portrait>, a young woman with voluminous, messy, long hair, wearing a black sporty crop top with white trim and matching black shorts, white calf-high socks. She is sitting perched on a high dark wooden counter or piano top. Her pose is casual and edgy, leaning forward with one hand resting near her mouth, biting her finger slightly, gazing directly at the camera with a sultry, nonchalant expression. A silver bracelet on her wrist. Background includes a white wall with vintage posters taped up and brown wooden blinds. Foreground details include the neck of an electric guitar, a glass jar of cookies, and liquor bottles",
    "composition": "Full body shot, 35mm focal length, slightly low angle to emphasize leg length, sharp focus on subject with hard flash fall-off, messy but balanced framing",
    "camera_angle": "Eye-level relative to the seated subject, slightly low angle from the floor, medium distance",
    "lighting": "Direct on-camera flash photography, hard lighting creating a sharp drop shadow on the wall behind the subject, high contrast, reminiscent of 90s point-and-shoot aesthetics, no diffusion",
    "color_palette": "Vintage film aesthetic, Kodak Gold 200 simulation, warm tones from wooden blinds and furniture contrasted with cool white flash light, deep blacks, slightly grainy texture, lo-fi indie vibe",
    "mood": "Indie sleaze, candid, rebellious, cool girl aesthetic, raw and authentic, Y2K retro fashion editorial, snapshot style",
    "negative_prompt": "soft lighting, studio lighting, bokeh, professional studio portrait, airbrushed skin, 3d render, cartoon, illustration, distorted hands, missing guitar strings, floating objects, anatomical errors, stiff pose, over-processed, HDR"
  }
}
The prompt above, and what it returned. Note the negative list — it is what stops the model reverting to a soft studio look.Open in LocalBanana

The negative_prompt slot is worth singling out. Most models drift toward a house style: soft light, smooth skin, shallow background blur. If your brief is the opposite of the house style, saying so explicitly is usually more effective than piling on more positive adjectives.

Prose or structured slots?

This was the most surprising result in the corpus, so it gets its own section.

13% of prompts (1,249) are written as structured objects with keys like subject, lighting, camera, style. The remaining 8,350 are sentences.

Prompt styleCountMedian views
Prose8,3509
Structured slots1,24998

An 11× gap — but structured prompts are also longer on average, so the obvious objection is that this is just the length effect again. It is not. Within each length band, structured still wins:

Length bandProse median viewsStructured median views
80–149 words17141
150–299 words3267
300+ words3398

Same length, different structure, consistently different outcome. The effect is strongest in the 80–149 word band — exactly the band most people should be writing in.

Why it plausibly works has nothing to do with the model parsing JSON. It is that a labelled slot cannot be left empty without you noticing. Prose lets you write a beautiful sentence that never once says where the light comes from. A list with lighting: on it does not.

You do not need strict JSON syntax. This is enough:

Subject: [who, physical detail, wardrobe, expression]
Setting: [location, what is behind them, foreground objects]
Light: [source, direction, hardness, colour]
Camera: [focal length, aperture, angle, distance]
Mood: [emotional register, reference medium]
Avoid: [what must not appear]
Upward Gaze in Minimalist Space

Upward Gaze in Minimalist Space

Nano Banana Pro
Prompt▾
{
"image_type": "photographic portrait",
"style": "studio portrait, cinematic, minimalist",
"composition": {
"orientation": "portrait",
"framing": "full body from extreme high angle",
"subject_position": "centered",
"camera_angle": "top-down (bird’s-eye / overhead)",
"lens_distortion": "strong wide-angle perspective exaggeration",
"negative_space": "extensive surrounding empty space",
"perspective": "dramatic overhead with subject looking up"
},
"subject": {
"count": 1,
"description": "young person with glasses",
"pose": "standing upright, shoulders slightly forward, arms relaxed",
"expression": "soft, introspective, mildly curious",
"gaze": "looking directly up at the camera",
"accessories": [
"round eyeglasses"
],
"clothing": {
"outerwear": "dark brown jacket",
"innerwear": "light-colored knit or textured shirt",
"style": "casual, understated"
}
},
"facial_details": {
"features": "rounded face, soft jawline",
"emotion": "calm, thoughtful",
"eye_emphasis": "enhanced by glasses and upward gaze"
},
"lighting": {
"type": "studio lighting",
"setup": "top-centered soft light with gradual falloff",
"contrast": "low to moderate",
"shadows": "subtle shadows beneath chin and body",
"vignette": "strong radial vignette darkening toward edges"
},
"color": {
"palette": [
"cool gray",
"charcoal",
"muted brown",
"soft beige"
],
"temperature": "cool-neutral",
"saturation": "low",
"mood": "quiet, contemplative"
},
"background": {
"environment": "studio",
"surface": "smooth seamless floor",
"gradient": "radial gradient from light center to dark edges",
"distractions": "none"
},
"technical_details": {
"camera_type": "digital",
"lens": "ultra-wide or fisheye-style wide-angle",
"depth_of_field": "deep (entire subject in focus)",
"sharpness": "high center sharpness with slight edge softness",
"noise": "minimal",
"post_processing": [
"contrast shaping",
"cool color grading",
"vignette enhancement",
"perspective exaggeration"
]
},
"artistic_elements": {
"concept": "isolation and vulnerability through scale and perspective",
"visual_metaphor": "small subject surrounded by vast empty space",
"aesthetic_influences": [
"editorial portrait photography",
"modern studio minimalism",
"cinematic overhead compositions"
]
},
"typography": {
"presence": false
},
"overall_mood": "intimate, introspective, slightly surreal",
"intended_use": [
"editorial portrait",
"conceptual photography reference",
"AI image generation style guide"
]
}
A structured prompt whose camera slot does the heavy lifting: full body, extreme high angle, centred subject.Open in LocalBanana

How to write the light slot

Light is the highest-leverage slot and the most commonly skipped. Here is what the corpus actually asks for, as a share of all 9,599 prompts:

Lighting termShare of prompts
studio lighting6.5%
cinematic lighting4.8%
rim light3.9%
natural light3.8%
flash3.5%
golden hour2.7%
soft light2.6%
backlight1.8%
warm light1.5%
overcast1.5%
diffused light1.3%
volumetric light1.3%
window light1.1%
neon light0.9%
side lighting0.5%
harsh light0.3%
blue hour0.2%
candlelight0.2%
top light0.05%

Two practical readings. First, rim light at 3.9% has quietly become standard vocabulary rather than an advanced move — you can use it without it reading as a trick. Second, and more useful: the bottom of this table is a list of looks almost nobody asks for. side lighting at 0.5%, harsh light at 0.3%, top light at 0.05%. They work fine. They are simply under-requested, which makes them the cheapest way to make an image not look like everything else.

The upgrade that costs nothing is naming direction and hardness, not just a mood word:

✕ Bad

cinematic lighting, dramatic, moody

✓ Good

single hard source from camera left at 45 degrees, no fill, deep shadow on the far cheek, cool teal spill on the back wall
Ice Blue Interrogation Femme Fatale

Ice Blue Interrogation Femme Fatale

Nano Banana Pro
Prompt▾
{
  "scene_description": {
    "location": "Cold interrogation room with teal-blue brick tiled wall, cinematic still from 'Basic Instinct' (1992)",
    "vibe": "Intense seductive power, classic film-noir tension, commanding presence",
    "lighting": "Teal-blue ambient room light, strong directional key from left casting long shadows, rim light sculpting figure and legs, high contrast, dramatic neo-noir mood",
    "camera_settings": "35mm anamorphic lens, f/2.0, shallow depth of field, cinematic film grain, subtle lens breathing, ultra-photorealistic 8k, 'Basic Instinct' color palette (cool teal shadows, warm skin highlights)"
  },
  "primary_subjects": [
    {
      "character": "attached celebrity lookalike",
      "role": "Interrogation subject / ultimate femme fatale",
      "action": "Seated in black metal chair, legs crossed seductively (right over left), right arm casually draped over chair back, left hand resting on thigh, upper body leaning back slightly, hips angled forward, piercing direct stare at camera, lips parted in confident half-smile",
      "details": "Sleeveless white high-neck mini dress clinging to curves, bare toned legs with glossy sheen, white pointed high heels, long dark glossy hair with loose waves, smoky cat-eye makeup, bold red lips, gold bracelet on right wrist"
    }
  ],
  "environment_details": {
    "crowd": "None – solitary cinematic intensity",
    "architecture": "Teal-blue brick tiled wall, black metal chair, concrete floor",
    "fidelity": "Hyper-realistic 8k, lifelike skin under cold light, fabric detail, hair shine, precise 'Basic Instinct' Catherine Tramell pose, lighting, and atmosphere"
  }
}
The colour of the light is named, not implied — teal-blue on the tiled wall, which is what makes the scene read as an interrogation room.Open in LocalBanana
Black-and-white silhouette behind frosted glass

Black-and-white silhouette behind frosted glass

Nano Banana Pro
Prompt▾
Abstract fine art black and white photography of a human silhouette seen through frosted glass, translucent plastic, or diffused acrylic. She is fully obscured, on her knees, with no facial detail, only a soft shadowed outline. Hands are pressed against the surface, fingers slightly spread, creating subtle depth and emotional tension. Extreme soft focus, heavy diffusion, smooth gradients, and gentle falloff of light. High-key monochrome tones, low contrast, minimal texture. Clean minimalist composition with negative space. Ethereal, surreal, haunting mood. Studio lighting with soft backlight behind the subject, cinematic fine-art photography, conceptual and introspective. Sensual pose. Aspect ratio 4:5
Light through diffusion as the actual subject. The prompt specifies the material between the source and the figure.Open in LocalBanana

How to write the camera slot

The corpus is unusually specific about camera vocabulary, which makes the gaps easy to spot:

Camera or composition termShare of prompts
depth of field14.5%
shallow depth of field11.1%
close-up9.3%
bokeh5.4%
symmetry5.1%
85mm4.9%
35mm3.5%
full body3.3%
negative space2.7%
low angle2.5%
50mm2.4%
wide-angle2.4%
f/1.82.0%
medium shot1.8%
eye level1.4%
f/2.81.3%
centered composition1.3%
24mm0.75%
high angle0.74%

One in seven prompts mentions depth of field. It is the single most common technical instruction in the whole corpus. And here is the gap worth exploiting: 11.1% ask for shallow depth of field by name, but only 2.0% write f/1.8. People name the lens and then hope for the opening. Naming the aperture is one token and it removes the ambiguity entirely.

A camera slot that actually constrains the frame has four parts:

Camera: 85mm, f/1.8, eye level, waist-up — background falls off to soft bokeh
Camera: 35mm, f/2.8, slightly low angle, full body — background legible, not blurred
Camera: 24mm, f/4, high angle looking down, subject centred with headroom

Focal length also decides how the face reads, which matters more than most people expect. 85mm leads the corpus at 4.9% — about 1.4× the rate of 35mm — and that ordering reflects how portrait-heavy the library is. Longer focal lengths compress features and separate the subject; wider ones exaggerate whatever is closest to the lens and keep the room in the story.

Urban Still Shadow Crowd Flow

Urban Still Shadow Crowd Flow

Nano Banana Pro
Prompt▾
Ultra-realistic cinematic street portrait of a young woman standing still in a crowded city street, sharp focus on her face with calm, intense expression. Long-exposure effect with people moving around her creating strong directional motion blur, streaked crowd movement while the subject remains perfectly still. Natural soft daylight, shallow depth of field, creamy bokeh, realistic skin texture, subtle freckles, neutral makeup, dark winter coat and scarf. Emotional, introspective mood, urban storytelling photography, DSLR quality, 85mm lens look, f/1.8, high dynamic range, 8K, professional color grading.
Shutter speed as a compositional instruction: the subject stays sharp while the crowd streaks into directional motion blur.Open in LocalBanana

How to route a prompt to the right model

The same corpus, split by which model ran the prompt, shows people have already worked out an informal division of labour:

Nano BananaGPT ImageMidjourney
Prompts in corpus3,3052,0434,305
Ask for text or typography21.8%51.7%7.9%
Carry camera specifications45.5%28.0%9.6%
Ask for illustration or anime19.4%36.2%17.0%

Read as instructions:

  • Photographic brief with a camera slot → over 45% of Nano Banana prompts carry camera specs, the highest of the three. Write the full lens-and-aperture line. Browse Nano Banana prompts.
  • Anything with legible words in the image → more than half of GPT Image prompts ask for text or typography, versus 7.9% for Midjourney. Put the exact copy in quotes in the prompt. Browse GPT Image prompts.
  • Mood-led work with no hard constraints → Midjourney prompts are the least technical in the corpus, and that is a fit, not a failure. Browse Midjourney prompts.
Premium Poster for Dan Dan Noodles

Nano Banana Pro

Premium Poster for Dan Dan Noodles

Futuristic Earbuds Ad

GPT Image

Futuristic Earbuds Ad

Hong Kong Office Lady Outfit Guide

GPT Image

Hong Kong Office Lady Outfit Guide

Three text-forward briefs from the gallery. The badge on each card shows which engine ran it.

The full split, including how the three engines handle the same brief, is in Nano Banana vs DALL·E and the 2026 generator comparison.

What to do when the output is wrong

The single most expensive habit is rewriting the whole prompt after a bad generation. You lose the information about which change did what.

Change one slot at a time, in this order:

  1. Composition wrong (subject too small, cropped badly, wrong angle) → fix the camera slot only. Distance and angle before anything else.
  2. Mood wrong but framing right → fix the light slot only. Direction and hardness, not adjectives.
  3. Style drifting toward generic → add exclusions. Name the house style you are getting and forbid it.
  4. Subject details wrong (wardrobe, age, hair) → move those details earlier in the prompt. Earlier tokens carry more weight in most models.
  5. Everything roughly right, want variations → hold the prompt and change one adjective per run, not five.

Reference-driven work follows the same rule. When you supply an input image, the prompt's job shifts from describing to constraining what may change:

Premium Industrial Product Diagram

Premium Industrial Product Diagram

Nano Banana Pro
Prompt▾
{
  "reference": {
    "type": "image",
    "usage": "Use the uploaded image as the sole source of truth. Do not assume product category in advance."
  },
  "scene": {
    "description": "Minimal, premium product presentation layout inspired by high-end industrial design and product documentation.",
    "background": "Pure white, clean, distraction-free"
  },
  "layout": {
    "top_left": {
      "content": "Brand name",
      "style": "Modern sans-serif typography, subtle, elegant"
    },
    "left_column": {
      "content": "Multiple auxiliary product views",
      "views": ["Primary front or main view", "Secondary side view", "Alternate angle or rear view", "Detail or top view if applicable"],
      "arrangement": "Vertically stacked, evenly spaced, aligned"
    },
    "right_section": {
      "content": "Main hero product render",
      "style": "Large, dominant, photorealistic",
      "lighting": "Soft studio lighting",
      "materials": "Visually accurate to the uploaded product"
    }
  },
  "visual_style": {
    "aesthetic": "Scandinavian, modern, industrial documentation",
    "mood": "Calm, precise, professional",
    "color_palette": "Neutral and restrained"
  },
  "rendering": {
    "quality": "Ultra-high resolution",
    "shadows": "Soft and realistic",
    "accuracy": "Exact proportions and geometry fidelity"
  }
}
A reference-driven brief that opens by declaring the uploaded image the sole source of truth, then specifies only the layout around it.Open in LocalBanana
One Image Generates 9 Different Shots

One Image Generates 9 Different Shots

Nano Banana Pro
Prompt▾
<instruction>
Analyze the entire composition of the input image. Identify all key subjects present (whether a single person, group/couple, vehicle, or specific object) and their spatial relationships/interactions.
Generate a coherent 3x3 grid "Contact Sheet" showcasing 9 distinct shots of exactly these subjects within the same environment.
You must adapt standard cinematic shot types to fit the content (e.g., if a group, keep the group together; if an object, frame the entire object):
Row 1 (Establishing Context):
1. Extreme Long Shot (ELS): Subjects appear very small within a vast environment.
2. Long Shot (LS): Full subject or group visible from top to bottom (head to toe / wheels to roof).
3. Medium Long Shot (American Shot/Three-Quarter): Framed from above the knees (for people) or a 3/4 view (for objects).
Row 2 (Core Coverage):
4. Medium Shot (MS): Framed from above the waist (or the central core of an object). Focus on interaction/action.
5. Medium Close-Up (MCU): Framed from above the chest. Intimate framing of the primary subject(s).
6. Close-Up (CU): Tight framing on the face or the "front" of an object.
Row 3 (Detail & Angles):
7. Extreme Close-Up (ECU): Macro detail with intense focus on a key feature (eyes, hands, logo, texture).
8. Low-Angle Shot (Worm's-Eye View): Looking up at the subject from ground level (spectacular/heroic feel).
9. High-Angle Shot (Bird's-Eye View): Looking down on the subject from above.
Ensure strict consistency: The same people/objects, same clothing, and same lighting across all 9 panels. Depth of field should vary realistically (background blur in close-ups).
</instruction>
A professional 3x3 cinematic storyboard grid containing 9 panels.
This grid showcases the specific subject(s)/scene from the input image across a comprehensive range of focal lengths.
Top Row: Wide environmental shots, full view, 3/4 crop (above-knee view).
Middle Row: Above-waist view, above-chest view, face/front close-up.
Bottom Row: Macro detail, low angle, high angle.
All frames feature photorealistic textures, consistent cinematic color grading, and correct framing for the specific number of subjects or objects analyzed.
Nine variations from one input in a single run — a faster way to explore than nine separate prompts.Open in LocalBanana

If you want a whole grid of consistent variations rather than single frames, the mechanics are in the multi-panel image guide.

A skeleton you can copy

Prose version, roughly 90 words when filled — right in the band where the corpus effect is strongest:

[Subject: age, build, hair, wardrobe, expression, pose].
[Setting: location, background surface, one or two foreground objects].
Lit by [source] from [direction], [hard or diffused], [colour of the light],
[what the shadows do].
Shot at [focal length], [aperture], [camera height], [framing distance].
[Mood in three words], [reference medium].
No [thing to exclude], no [thing to exclude].

Structured version, same content, labelled slots:

{
  "subject": "",
  "setting": "",
  "lighting": "source, direction, hardness, colour, shadow behaviour",
  "camera": "focal length, aperture, height, framing distance",
  "mood": "",
  "style": "reference medium, grain, colour treatment",
  "negative_prompt": ""
}

Fill every slot. If you genuinely do not care about one, write what you want the model to default to instead of leaving it blank — that is the whole trick.

WormseyeUltra Realistic Deconstruction PortraitCity RibbonPhotorealistic Overhead Macro PhotographyBrowse all prompts

When you have a skeleton you like, run it in the generator and change one slot at a time. To read more well-structured prompts before writing your own, the portrait category is the densest place in the gallery.

FAQ

How long should an AI image prompt be?

The median in a 9,599-prompt corpus is 83 words, and median engagement steps up sharply across the 80-word boundary — from 5 to 19 median views. Aim for 80–150 words: enough to carry subject, setting, light, camera and mood without padding.

What should an AI image prompt include?

Six slots: subject, setting, light, camera, mood and style, and exclusions. Subject and mood are the two almost everyone writes; light and camera are where most images are won or lost; exclusions are the cheapest and rarest fix.

Should I write AI image prompts in JSON?

Structured prompts show markedly higher median engagement than prose at the same length — 141 versus 17 median views in the 80–149 word band. Strict JSON syntax is not the point. Labelled slots are, because an empty slot becomes visible.

Do negative prompts work in AI image generation?

They are useful specifically when your brief runs against the model's default look. If you want hard flash and the model keeps returning soft studio light, listing soft lighting, studio lighting, bokeh, airbrushed skin as exclusions is more reliable than adding more positive adjectives.

Which model should I use for which kind of prompt?

Photographic work with camera specifications goes to Nano Banana, where 45.5% of prompts carry them. Anything with legible text in the image goes to GPT Image, where 51.7% of prompts ask for typography versus 7.9% for Midjourney. Mood-led work with no hard constraints suits Midjourney. Every prompt counted here is published in the LocalBanana gallery with the image it produced and the model that ran it.

LocalBanana Team

We run Nano Banana, GPT Image and other models side by side, and publish every prompt in our gallery. These guides are written from that corpus.

Updated August 5, 2026@LocalBanana_io

Try it yourself

Recreate these looks in LocalBanana

Every prompt in this article is ready to run — and there are plenty more in the gallery.

Start creatingAll prompts

Related articles

Nine AI Image Prompt Patterns That Cover Most Jobs

Nine AI Image Prompt Patterns That Cover Most Jobs

Tutorials·Feb 2, 2026·25 min read

10 Prompt Mistakes Beginners Make (And How to Fix Them)

10 Prompt Mistakes Beginners Make (And How to Fix Them)

Tips & Tricks·Feb 2, 2026·5 min read

AI Image Editing Prompts Tested: Replace, Relight, and Preserve Identity

AI Image Editing Prompts Tested: Replace, Relight, and Preserve Identity

Tutorials·Sep 3, 2026·12 min read