Tips & Tricks5 min read
10 Prompt Mistakes Beginners Make (And How to Fix Them)
Avoid these common pitfalls and instantly improve your AI image generation results with simple fixes.

On this page
- How these ten were chosen
- Mistake 1: The prompt stops at the subject
- Mistake 2: Asking for blur without naming the aperture
- Mistake 3: Leaving the light unspecified
- Mistake 4: Writing one long paragraph instead of labelled slots
- Mistake 5: Stacking quality adjectives instead of specifications
- Mistake 6: Giving a lens but no camera position
- Mistake 7: Describing what you do not want
- Mistake 8: Contradicting yourself
- Mistake 9: Sending the job to the wrong model
- Mistake 10: Copying a prompt without changing the parts that only worked for its author
- A prompt that fixes all ten at once
- FAQ
- Why are my AI images not turning out how I imagined?
- How long should a beginner prompt be?
- Do negative prompts work in AI image generators?
- Why do copied prompts give me different results?
- Which model should a beginner start with?
How these ten were chosen
Most "beginner mistakes" lists are one person's habits generalised into rules. These ten were picked by comparing what a short prompt typically contains against 9,599 published prompts in the LocalBanana gallery — every item with a prompt longer than 20 characters, stored alongside the model that ran it and the image it produced.
So each mistake below carries a number: how often the fixed version actually appears in the corpus, and what the fix looks like as text you can paste.
20%
of prompts are under 30 words
1,966 of 9,599 — the most common shape of a beginner prompt
11.1%
ask for shallow depth of field
but 77.6% of them never name an aperture
4 vs 19
median views, under 30 words vs 80-149 words
the step happens at 80 words, not gradually
What these numbers are and are not
Word frequencies are exact counts over the corpus — they are simply what people wrote. View counts are on-site engagement, which is also affected by how long an item has been published and where the feed surfaced it, so treat every view comparison here as correlation, not proof of causation.
Mistake 1: The prompt stops at the subject
A fifth of the corpus is under 30 words. Almost all of those prompts name a subject and stop — which leaves lighting, lens, framing, wardrobe and mood to the model's defaults, and the defaults are the same for everyone.
Median views by prompt length tell the same story from the other side: 4 for prompts under 30 words, 5 for 30–79, then 19 for 80–149. The step is not gradual. It lands at roughly the length where a prompt can carry subject and light and camera and mood at once.
✕ Bad
a woman in a cafe, beautiful, high quality✓ Good
A woman in her late twenties sitting at a cafe window seat, mid-morning, oversized cream knit sweater, hands around a ceramic cup, looking out of frame to camera left. Soft window light from the left, warm bounce off the table. 85mm, f/1.8, eye level, shallow depth of field. Quiet, unposed, editorial.
Indoor Casual Fashion Portrait
Nano Banana ProPrompt▾
A young woman standing in a simple indoor setting with a cream-colored ceramic tile wall behind her. She wears a tight-fitting chocolate brown long-sleeved ribbed Henley top with the top buttons casually left undone, revealing a natural collarbone line. Her posture is relaxed and upright, arms resting naturally at her sides. Medium shot framed from mid-thigh up, eye-level camera angle. Soft diffused indoor lighting from a ceiling source above, casting a gentle downward shadow on the tiled wall behind her. Her expression is calm and composed with direct eye contact toward the camera, radiating a confident natural beauty. High-fidelity photorealism with accurate skin texture, fabric ribbing detail, and natural hair texture. Casual smartphone portrait aesthetic, as if captured spontaneously in everyday life. Natural warm color palette with chocolate brown, cream, and soft skin tones. Sharp focus, accurate texture rendering, 3:4 portrait aspect ratio.
Mistake 2: Asking for blur without naming the aperture
The sharpest contradiction in the whole corpus: 11.1% of prompts ask for shallow depth of field by name, and 77.6% of those never write an aperture like f/1.8. People describe the effect and withhold the setting that produces it.
Focal lengths are much better covered — 85mm appears in 4.9% of prompts, 35mm in 3.5%, 50mm in 2.4%. So the typical prompt names the lens and leaves the opening blank, which is the half that actually controls the blur.
✕ Bad
portrait with beautiful blurry background, bokeh, professional✓ Good
85mm, f/1.8, focus on the eyes, background falls off into soft bokeh, subject 2 metres from a brick wall 6 metres behind herDistance matters as much as the number. f/1.8 with the subject pressed against a wall produces no separation at all; the same aperture with six metres behind them produces the look people are asking for.
Mistake 3: Leaving the light unspecified
studio lighting is the single most common lighting term in the corpus at 6.5%. Every other term is rarer: cinematic lighting 4.8%, rim light 3.9%, natural light 3.8%. Add them all up and most prompts still say nothing at all about light.
That is the cheapest available upgrade. Naming a direction, a quality and a source puts you ahead of the median with about eight words.
✕ Bad
good lighting, dramatic, cinematic✓ Good
Single hard key from camera right at 45 degrees, no fill, deep shadow down the left side of the face, cool practical light in the background
Ice Blue Interrogation Femme Fatale
Nano Banana ProPrompt▾
{
"scene_description": {
"location": "Cold interrogation room with teal-blue brick tiled wall, cinematic still from 'Basic Instinct' (1992)",
"vibe": "Intense seductive power, classic film-noir tension, commanding presence",
"lighting": "Teal-blue ambient room light, strong directional key from left casting long shadows, rim light sculpting figure and legs, high contrast, dramatic neo-noir mood",
"camera_settings": "35mm anamorphic lens, f/2.0, shallow depth of field, cinematic film grain, subtle lens breathing, ultra-photorealistic 8k, 'Basic Instinct' color palette (cool teal shadows, warm skin highlights)"
},
"primary_subjects": [
{
"character": "attached celebrity lookalike",
"role": "Interrogation subject / ultimate femme fatale",
"action": "Seated in black metal chair, legs crossed seductively (right over left), right arm casually draped over chair back, left hand resting on thigh, upper body leaning back slightly, hips angled forward, piercing direct stare at camera, lips parted in confident half-smile",
"details": "Sleeveless white high-neck mini dress clinging to curves, bare toned legs with glossy sheen, white pointed high heels, long dark glossy hair with loose waves, smoky cat-eye makeup, bold red lips, gold bracelet on right wrist"
}
],
"environment_details": {
"crowd": "None – solitary cinematic intensity",
"architecture": "Teal-blue brick tiled wall, black metal chair, concrete floor",
"fidelity": "Hyper-realistic 8k, lifelike skin under cold light, fabric detail, hair shine, precise 'Basic Instinct' Catherine Tramell pose, lighting, and atmosphere"
}
}The full frequency table for twenty lighting terms, and what each one does to a face, is in the lighting keyword guide.
Mistake 4: Writing one long paragraph instead of labelled slots
13% of the corpus (1,249 prompts) is written as structured JSON rather than prose. Those prompts show markedly higher median engagement — 98 vs 9 overall — and the gap survives when you control for length: in the 80–149 word band, structured prompts sit at 141 median views against 17 for prose of the same length. Again: correlation, and long prompts get more of everything. But the mechanism is easy to believe.
A labelled slot cannot be silently skipped. Prose lets you write a lovely sentence that never mentions where the light comes from.
{
"subject": "man, early 30s, short dark hair, three-day stubble",
"wardrobe": "charcoal wool overcoat, grey crew neck underneath",
"setting": "empty underground car park, concrete pillars, wet floor",
"lighting": "overhead fluorescent strips, hard top light, green colour cast",
"camera": "35mm, f/2.8, low angle, full body",
"mood": "still, waiting, unglamorous",
"style": "documentary photography, fine grain"
}

High-angle Chic Influencer Portrait
Nano Banana ProPrompt▾
{
"subject": {
"description": "A stunning high-angle shot of a chic Asian fashion influencer with a cool, alluring attitude.",
"age": "20s",
"expression": {
"eyes": {
"look": "sharp fox-eyes, piercing gaze directed at camera",
"energy": "confident, slightly cold, seductive",
"details": "defined eyeliner, emphasized aegyosal"
},
"mouth": {
"position": "relaxed lips, subtle smirk",
"energy": "chic"
},
"overall": "stunning, high-visual-impact beauty"
},
"face": {
"preserve_original": false,
"makeup": "high-contrast makeup, pale porcelain skin, reddish gradient lips, sharp jawline, V-shape face",
"style": "cool-toned beauty, K-pop idol visual"
},
"hair": {
"color": "black",
"style": "long sleek straight hair with full straight bangs",
"effect": "glossy, high-fashion finish"
},
"body": {
"frame": "slim, petite, fragile aesthetic",
"pose": {
"position": "leaning forward significantly",
"overall": "dynamic foreshortening, emphasis on head and upper torso"
},
"skin": {
"tone": "cold fair skin",
"lighting_effect": "brightened face, soft beauty lighting, no dark shadows"
}
},
"clothing": {
"top": {
"type": "ultra-fine gauge knit top",
"color": "cool grey",
"details": "mock neck, skin-tight fit, lightweight thin fabric (not thick)",
"effect": "perfectly sculpting body curves, smooth texture"
},
"bottom": {
"type": "dark pencil skirt",
"details": "high waisted with thin luxury belt"
}
}
},
"photography": {
"camera_style": "High-end social media snapshot",
"angle": "High angle POV",
"shot_type": "Medium close-up",
"aspect_ratio": "9:16",
"texture": "clear, sharp, slightly filtered for beauty",
"lighting": "overcast cool daylight, soft diffuse light"
},
"background": {
"setting": "European classic architecture",
"atmosphere": "fashionable street corner",
"blur": "bokeh background to emphasize subject"
},
"negative_prompt": [
"round face",
"plain face",
"no makeup",
"warm yellow skin",
"chunky knit",
"thick sweater",
"loose clothing",
"wrinkled fabric",
"dull eyes",
"friendly boring smile",
"low resolution",
"dark lighting"
]
}You do not need strict JSON. A plain list — Subject: / Setting: / Lighting: / Camera: / Mood: — reproduces most of the effect, because the empty line is visible.
Mistake 5: Stacking quality adjectives instead of specifications
masterpiece, 8k, ultra detailed, award winning, hyperrealistic — these are inherited from an earlier generation of models and they cost you the words you could have spent on facts. Length correlates with better outcomes because long prompts carry more specifications, not more words. Two hundred words of praise is still a prompt with nothing in it.
✕ Bad
masterpiece, best quality, 8k, ultra detailed, award winning photography, stunning, breathtaking, hyperrealistic portrait of a woman✓ Good
Portrait of a woman in her forties, visible skin texture and fine lines left intact, no retouching. Soft north-facing window light, white wall two metres behind. 85mm, f/2.8, chest-up framing.The good version is shorter and contains four more decisions.

Moody Cinematic Portrait
Nano Banana ProPrompt▾
Moody cinematic portrait of a young woman with clear fair skin and barely-there makeup, warm natural lips. Loose dark brown hair, softly textured. Calm, self-assured expression. She's wearing a fitted black turtleneck sweater, clean lines, understated elegance. Dark minimal background fading into black. Gentle diffused front light, soft shadows, subtle film grain, organic color grading, shallow depth of field, high-end editorial feel.
Mistake 6: Giving a lens but no camera position
Camera position is the most under-specified control in the corpus. low angle appears in 2.5% of prompts, eye level in 1.4%, high angle in 0.74%. Meanwhile close-up is at 9.3% — people specify how tight the crop is far more often than where the camera is standing, even though the height is what changes how the subject reads.
✕ Bad
close up portrait, 50mm lens, nice composition✓ Good
50mm, chest-up crop, camera slightly below eye level looking up, subject centred with negative space above the head
Upward Gaze in Minimalist Space
Nano Banana ProPrompt▾
{
"image_type": "photographic portrait",
"style": "studio portrait, cinematic, minimalist",
"composition": {
"orientation": "portrait",
"framing": "full body from extreme high angle",
"subject_position": "centered",
"camera_angle": "top-down (bird’s-eye / overhead)",
"lens_distortion": "strong wide-angle perspective exaggeration",
"negative_space": "extensive surrounding empty space",
"perspective": "dramatic overhead with subject looking up"
},
"subject": {
"count": 1,
"description": "young person with glasses",
"pose": "standing upright, shoulders slightly forward, arms relaxed",
"expression": "soft, introspective, mildly curious",
"gaze": "looking directly up at the camera",
"accessories": [
"round eyeglasses"
],
"clothing": {
"outerwear": "dark brown jacket",
"innerwear": "light-colored knit or textured shirt",
"style": "casual, understated"
}
},
"facial_details": {
"features": "rounded face, soft jawline",
"emotion": "calm, thoughtful",
"eye_emphasis": "enhanced by glasses and upward gaze"
},
"lighting": {
"type": "studio lighting",
"setup": "top-centered soft light with gradual falloff",
"contrast": "low to moderate",
"shadows": "subtle shadows beneath chin and body",
"vignette": "strong radial vignette darkening toward edges"
},
"color": {
"palette": [
"cool gray",
"charcoal",
"muted brown",
"soft beige"
],
"temperature": "cool-neutral",
"saturation": "low",
"mood": "quiet, contemplative"
},
"background": {
"environment": "studio",
"surface": "smooth seamless floor",
"gradient": "radial gradient from light center to dark edges",
"distractions": "none"
},
"technical_details": {
"camera_type": "digital",
"lens": "ultra-wide or fisheye-style wide-angle",
"depth_of_field": "deep (entire subject in focus)",
"sharpness": "high center sharpness with slight edge softness",
"noise": "minimal",
"post_processing": [
"contrast shaping",
"cool color grading",
"vignette enhancement",
"perspective exaggeration"
]
},
"artistic_elements": {
"concept": "isolation and vulnerability through scale and perspective",
"visual_metaphor": "small subject surrounded by vast empty space",
"aesthetic_influences": [
"editorial portrait photography",
"modern studio minimalism",
"cinematic overhead compositions"
]
},
"typography": {
"presence": false
},
"overall_mood": "intimate, introspective, slightly surreal",
"intended_use": [
"editorial portrait",
"conceptual photography reference",
"AI image generation style guide"
]
}Mistake 7: Describing what you do not want
Negative wording forces the model to represent the thing you are trying to exclude. It is also unnecessary: every negative has a positive form that is more specific.
✕ Bad
no people, not blurry, no text, without weird hands, not cartoon✓ Good
Empty street at dawn, sharp focus throughout, clean untextured surfaces, hands resting flat on the table in full view, photographic realismMistake 8: Contradicting yourself
Physically incompatible instructions get resolved by the model picking one, and you do not get to choose which. The common pairs:
| Contradiction | Why it cannot resolve | Pick one |
|---|---|---|
24mm wide-angle + heavy bokeh | Wide lenses have deep focus; that is the optics | Wide and deep, or long and shallow |
close-up + full body | Two different crops | close-up (9.3% of prompts) or full body (3.3%) |
golden hour + overcast | Two different skies | Warm and directional, or flat and soft |
symmetry + rule of thirds | Two different placements | symmetry appears in 5.1% of prompts; centre it or offset it |
candid, unposed + looking directly at camera, posed | Two different moments | Decide whether the subject knows the camera is there |
✕ Bad
35mm wide angle shot, extreme close-up of her face, heavy bokeh, golden hour, overcast day✓ Good
35mm, waist-up, subject two metres away, background compressed but still readable, late afternoon sun from behind camera leftMistake 9: Sending the job to the wrong model
Split the same corpus by which model ran the prompt and the specialisation is obvious:
| Nano Banana | GPT Image | Midjourney | |
|---|---|---|---|
| Prompts in corpus | 3,305 | 2,043 | 4,305 |
| Ask for text or typography | 21.8% | 51.7% | 7.9% |
| Carry camera specifications | 45.5% | 28.0% | 9.6% |
| Ask for illustration or anime | 19.4% | 36.2% | 17.0% |
Over half of GPT Image prompts ask for words in the picture, against 7.9% for Midjourney. If your image needs a legible headline, the model matters more than the prompt.

Clean Modern YouTube Thumbnail
GPT ImagePrompt▾
A clean modern YouTube thumbnail on a bright white background, split visually between bold text on the left and a cropped presenter on the right. On the left, place a massive stacked headline in heavy black sans-serif uppercase reading CLAUDE DESIGN, aligned top-left across two lines. Below it, add a large rounded rectangular orange banner with a subtle drop shadow containing bold white uppercase text with an orange outline that reads IN 6 MINS. Near the headline, include 2 orange starburst spark shapes, one small beside the text and one much larger in the far right background. In the lower-left quadrant, show a floating app interface mockup for a design tool dashboard with a soft off-white panel, thin borders, rounded corners, and a faint shadow. The interface should have a left sidebar with 4 menu items labeled exactly: "Recent", "Your designs", "Examples", and "Design systems". The main panel header should say "Recent" and contain 3 project cards arranged in a grid, labeled exactly: "Thumio Design System", "Thumio Startup", and "Design System". On the right half, show a cropped photorealistic young adult man from chest up wearing a white turtleneck sweater, with messy light brown hair and light facial hair, positioned close to camera. His face is intentionally obscured by a solid tan rectangle placeholder covering most of the face area. He points with one hand toward the interface panel on the left, finger extended. Use a warm orange and white brand palette, high contrast, minimal clutter, polished creator-economy thumbnail styling, subtle shadows, and a composition designed for tech tutorial content.
Full breakdowns: Nano Banana vs GPT Image and Nano Banana vs Midjourney.
Mistake 10: Copying a prompt without changing the parts that only worked for its author
This one is invisible until you look at real prompts side by side. Two things travel with a copied prompt and quietly break it:
Bracketed placeholders. Plenty of published prompts are written as reusable templates with fields like [BRAND NAME], [FOOD], [CITY] left in. Paste one unchanged and the model will render the bracket text, invent a brand, or produce a generic stand-in.

Full High-End Product Promotional Photo
Nano Banana ProPrompt▾
[BRAND NAME] is launching a new functional wellness elixir (e.g., adaptogenic, nootropic, or natural energy drink). As the Creative Director, devise a product name and visualize a complete high-end promotional shot. The aesthetic is "Cosmic Premium"—technological, clean, and sophisticated, like top-tier Apple product photography. THE PRODUCT: Design a sculptural, multi-layered beverage bottle suspended in the center. The form is engineered and futuristic. The materials are hyper-tactile: bead-blasted titanium details, frosted borosilicate glass, and textured haptic polymer grips. **Crucial Color Instruction:** The liquid inside must have a distinct, natural color relevant to its invented function (e.g., vibrant turmeric yellow, deep berry red, earthy matcha green, or pale calming blue). The liquid should look real with subtle natural sediment. **Crucial Graphic Detail:** On the clear glass section of the bottle, apply a layer of subtle, minimalist, technical typography printed in matte white ink. This design should feel utilitarian and futuristic (e.g., small technical specs like 'SPACE GRADE FORMULA', 'BATCH: OZ-9', volume indicators, or coordinate markings), adding a functional aesthetic similar to aerospace labeling, without overwhelming the bottle's clean lines. THE ENVIRONMENT & LIGHTING: The bottle is in a seamless studio. **Crucial Background Instruction:** The background must be a solid, clean, very light pastel tone that is specifically chosen to complement the liquid color (e.g., a soft cool mint background for a warm orange liquid, or a pale blush background for a deep green liquid). No gradients. Ultra-soft, diffused studio lighting creates sleek highlights on metal and deep subsurface scattering in the glass and liquid. PHOTOGRAPHY STYLE: High-resolution 100mm macro lens shot. Shallow depth of field, sharp focus on bottle textures and the printed graphics on the glass, smooth pastel background bokeh. 8k resolve, hyper-realistic textures. GRAPHIC OVERLAYS: Include subtle dark gray UI elements. Bottom Left Corner: Very small, minimalist text (like Manrope Regular font) describing the product's name and function in two sentences. Bottom Right Corner: A small, minimalist dark gray logomark for [BRAND NAME].
Account-bound flags. Midjourney prompts in the corpus routinely end with things like --profile wstgdj4 or --sref 2143660546. A profile code points at that person's personalisation data; an --sref points at a specific style reference. Copy them and you are asking for someone else's taste with none of it attached — or nothing at all, on a platform that does not read those flags.

Cat by the Sea in Orange Light
MidjourneyPrompt▾
Sea, cat, orange warm bright light --ar 1:2 --profile wstgdj4
Before you reuse any prompt: replace every bracket, strip flags that belong to another platform, and check whether a reference image was doing work the text does not describe.
A prompt that fixes all ten at once
Subject: woman in her early thirties, shoulder-length dark hair pushed
behind one ear, no visible makeup, faint smile
Wardrobe: oversized ecru linen shirt, sleeves rolled to the elbow
Setting: kitchen table beside a north-facing window, mid-morning,
half-finished coffee and an open notebook in frame
Lighting: soft directional window light from camera left, white wall
bouncing a little fill into the shadow side, no artificial light
Camera: 50mm, f/2.0, eye level, waist-up, subject offset to the right
third with negative space to the left
Motion: caught mid-turn towards the window, not looking at the lens
Style: editorial documentary photography, fine natural grain,
neutral colour, skin texture preserved
Ninety words. It names a lens and an aperture, a light direction and a quality, a camera height, a crop, a moment, and no adjectives about how good it should be. Run it in the generator and change one line at a time to see which slot is carrying the image.
FAQ
Why are my AI images not turning out how I imagined?
In most cases the prompt named a subject and nothing else. A fifth of the prompts in our 9,599-item corpus are under 30 words, which leaves lighting, lens, framing and mood to defaults. Add a light direction and a camera line first — those two changes alter the image more than any other edit.
How long should a beginner prompt be?
Aim for 80–150 words. That is where median engagement steps up in our data (from 5 to 19 median views between the 30–79 and 80–149 bands), and it is roughly the length needed to carry subject, light, camera and mood without padding. The full length analysis is in what 9,667 prompts contain.
Do negative prompts work in AI image generators?
Writing what you do not want is a weaker instruction than writing what you do want, and it varies by platform. Convert every negative into a positive: instead of "not blurry", write "sharp focus throughout"; instead of "no weird hands", describe where the hands are and what they are doing.
Why do copied prompts give me different results?
Three usual reasons: a bracketed placeholder such as [BRAND NAME] was left unreplaced, the prompt carries platform-specific flags like --sref or --profile that mean nothing outside Midjourney, or the original used a reference image the text never mentions. Every prompt in our gallery shows the model that ran it, so you can check the first two in seconds.
Which model should a beginner start with?
For photographic work, Nano Banana — 45.5% of its prompts in our corpus carry camera specifications, and it is built to be briefed like a photographer. For anything with readable text in the frame, GPT Image, where 51.7% of prompts are text jobs. The comparison guide covers the rest.
LocalBanana Team
We run Nano Banana, GPT Image and other models side by side, and publish every prompt in our gallery. These guides are written from that corpus.
Updated @LocalBanana_io





