Gemini Image Generator Prompts: The Complete Guide (2026)
Google's Gemini image generator (the model many people still call by its old nickname, "Nano Banana") works differently from most AI image tools. It does not just generate a picture from a text description. It can hold a character's face steady across many images, put clean readable text inside a photo, and change one small detail in a picture without touching anything else.
This guide explains what makes it different, gives you 9 original prompt templates built for its actual strengths, shows you what usually goes wrong and how to fix it, and compares it fairly against the other big image models people use today.
What Google's Gemini Image Generator Actually Is
Most AI image generators are diffusion models. They start from random noise and slowly clean it up into a picture, guided by your text prompt. This approach is good at making striking, varied images, but it struggles with two things: keeping a character looking the same twice, and spelling words correctly inside an image.
Google's Gemini image model was built with a different focus. It treats an image more like something to be understood and edited step by step, not just generated from scratch each time. That is why it tends to be much better at:
- Keeping the same character, face, or object consistent across separate images.
- Rendering text that is spelled correctly and placed where you asked.
- Making one small, specific change to an image while leaving the rest exactly as it was.
The nickname "Nano Banana" came from an early version of this model. Google has since released newer versions under the Gemini name, so this guide uses the generic term "Gemini image generator" since that is what will still make sense next year, even after the model gets renamed again.
The one thing to remember before you write any prompt for this model: it rewards detail. A vague prompt gets you a generic result. A long, exact, specific prompt gets you exactly what you described. This is the opposite of some other image tools, where too much detail can confuse the output.
The 3 Core Strengths, and Why They Work
1. Character Consistency
Ask most image tools to draw "the same person" in a new scene, and you usually get someone who looks similar but not identical. Gemini holds identity much better because it does not treat each generation as a fresh, unrelated task. If you give it one complete, exact description of a character and reuse that exact same wording every time, it treats that description as a fixed anchor point rather than a rough suggestion.
2. Text That Is Actually Readable
Getting an AI image tool to write a sign, label, or book cover correctly has always been hard, because most models treat letters as shapes to guess at, not as text to spell out. Gemini treats the actual text you write in quotes as something to reproduce exactly, which is why it is far more reliable at signs, labels, and typography inside images.
3. Precise, Local Edits
Most models regenerate a big chunk of the image even when you ask for one small change, because they do not have a strong sense of what should stay fixed. Gemini can be told, in plain language, to change one detail and leave everything else untouched, and it generally listens to that instruction rather than reinterpreting the whole scene.
The pattern behind all three strengths is the same: describe what should stay fixed in full detail, then describe what should change. Gemini follows that structure closely.
9 Prompt Templates for Gemini Image Generator
These templates are written from scratch. Copy the structure, then fill in your own details.
1. Character Consistency
Use this when you need the same character to appear across many images, like a comic, a mascot, or a story.
Template: "A [age] [gender] character with [hair color and style], [eye color], [skin tone], wearing [exact clothing description]. [One distinguishing feature, like a scar, glasses, or a tattoo]. Now show this exact character [doing an action] in [a setting]. Keep the face, hair, and clothing identical to the description above."
Tip: Save this character description in a text file. Paste the same wording, word for word, every time you generate a new scene with that character.
2. Legible Text in an Image
Use this for signs, book covers, product labels, or anything where the words need to be spelled correctly.
Template: "A [type of object or surface, like a wooden cafe sign or a paperback book cover] with the text '[your exact text]' written in [font style, like bold serif letters or handwritten chalk lettering]. The text is [color] on a [background color] background. [Describe the surrounding scene]."
Tip: Always put your exact wording inside quotation marks. Describe the text as a physical object sitting on a surface, not as something floating in the air.
3. Product Placed Naturally in a Scene
Use this for product photos, mockups, or online store listings.
Template: "A [product name] made of [material], in the shape of [exact shape and proportions], with the label reading '[label text]' in [color]. Place this exact product on/in [describe the setting, like a marble kitchen counter or a grassy picnic table]. The lighting is [describe lighting, like soft morning light from the left]."
Tip: Describe the product completely in the first sentence, then describe the scene in a separate sentence. Keeping these two parts separate helps the model keep the product looking correct while still placing it convincingly.
4. Changing the Art Style
Use this when you want a photo turned into a painting, cartoon, or another art style, without changing the pose or composition.
Template: "A photorealistic image of [subject, pose, and setting in full detail]. Now redraw this exact same subject, pose, and setting in the style of [art style, like a watercolor painting or a 1990s anime]. Keep the composition, pose, and identity of the subject exactly the same. Only the art style should change."
Tip: Always describe the photo-realistic version first, in full detail, even though you will not use that version. This gives the model something exact to hold onto while it changes only the style.
5. One Specific Edit, Nothing Else
Use this when you want to change a single detail in an existing description without touching the rest.
Template: "[Full description of the scene as it currently is, in detail]. Make one change: [the single specific thing to change, like turn the red jacket to blue, or remove the sunglasses]. Everything else in the scene, including the pose, background, and lighting, should stay exactly the same."
Tip: Put the one change you want in its own short sentence, separate from the rest of the description. Do not bury it in a long paragraph of other details.
6. Combining Two Images Into One Scene
Use this when you want to place a person or object from one description into a different setting or combine two separate ideas into one image.
Template: "Take [subject A, described in full detail] and place them into [scene B, described in full detail]. [Subject A] should look natural in this new setting, matching the lighting and scale of [scene B], while keeping their appearance from the original description unchanged."
Tip: Describe the subject and the new setting as two separate, complete blocks of text before asking for them to be combined.
7. Controlling Pose and Camera Angle
Use this when the character or object needs to appear from a specific angle, like a product turnaround or a comic panel.
Template: "[Full description of the subject]. Show this subject from a [camera angle, like a low angle looking up, or a three-quarter view from the left]. The subject's pose is [describe the exact pose, like standing with arms crossed, or sitting cross-legged]. The camera is [distance, like close-up on the face, or a wide full-body shot]."
Tip: Name the camera angle and distance directly, using simple terms like "low angle," "eye level," "close-up," or "wide shot," rather than describing it indirectly.
8. Changing Lighting or Mood
Use this when you want the same scene shown in a different time of day or emotional tone.
Template: "[Full description of the scene, unchanged]. Now show this exact same scene but with [new lighting, like warm golden hour light, cold blue moonlight, or harsh overhead fluorescent light]. Keep the composition, subjects, and objects the same. Only the lighting and overall mood should change."
Tip: Name a specific light source and time of day (morning sun, candlelight, streetlight) rather than a vague word like "moody," since specific light sources render more predictably.
9. A Consistent Set of Product Photos
Use this when you need the same product shown from several angles or in several settings for a catalog.
Template: "A [product] with [exact shape, color, and material details], and a label reading '[label text]'. This is photo [number] of a set. Show it [specific angle or setting for this photo, like from the front on a white background, or held in a hand outdoors]. Keep the product's shape, color, and label identical in every photo in this set."
Tip: Number each photo in the set within the prompt itself, and repeat the exact same product description word for word across every photo in that set.
Common Failure Modes and How to Fix Them
Even with good templates, things go wrong. Here is what usually happens and how to correct it.
The character's face keeps changing slightly between images. This almost always means the character description was paraphrased instead of copied exactly. Save the description once and paste it unchanged every time, instead of retyping it from memory.
Text in the image is misspelled or garbled. This usually happens when the text is described loosely, like "a sign that says something about coffee," instead of the exact words in quotation marks. Always write the exact text you want inside quotes.
An edit changed more than you asked for. This usually means the single change was buried inside a long paragraph. Separate the one change into its own short sentence and explicitly say what should stay the same.
The product's shape or proportions look slightly different than the real product. This usually happens when the product description and the scene description are mixed together in one long sentence. Describe the product completely on its own first, then describe the scene separately.
The art style change also changed the pose or identity of the subject. This happens when the instruction only mentions the new style and not what to preserve. Always add a sentence that says the composition, pose, and identity should stay the same.
A Repeatable Prompt Framework You Can Reuse
Across all nine templates above, the same pattern repeats. Use this checklist for any new prompt you write:
- Describe what must stay fixed, in full detail, first. This could be a character, a product, or a scene.
- Describe what should change, in a separate, short sentence. Keep it isolated from the fixed description.
- Say what should stay the same, out loud. Do not assume the model will guess. Add a sentence like "everything else should stay the same."
- Put any exact text in quotation marks. Never describe words loosely.
- Reuse fixed descriptions word for word. If you are building a series, copy and paste the same description each time instead of rewriting it.
- Name specific details instead of vague moods. "Warm afternoon sunlight" works better than "nice lighting."
Gemini vs GPT Image 2 vs Midjourney vs FLUX
Each of the major image models has a different strength. Here is how to think about which one fits your task.
Gemini image generator is the strongest choice when you need the same character, face, or product to stay consistent across many images, or when the words inside the image need to be spelled correctly.
GPT Image 2 (from OpenAI, the current model behind ChatGPT's image tool) is a strong general-purpose generator with wide flexibility. It is a good default when you want something new and creative rather than something that needs to stay consistent with a previous image.
Midjourney is best known for artistic and stylized output. If the goal is a striking, painterly, or highly stylized image rather than accuracy or consistency, Midjourney tends to produce the most visually impressive results.
FLUX is a strong choice for realistic photography-style images at speed, often used by people who need a large number of natural-looking images quickly.
If your task is "keep this exact thing the same, or change one small part of it," Gemini is built for that. If your task is "make something new and visually striking," look at Midjourney or GPT Image 2 instead.
Real Workflow Examples

Example 1: A four-panel comic strip with the same character.
Write out the full character description once and save it. For panel one, paste the character description and add the first scene. For panel two, paste the exact same character description again and only change the scene and action. Repeat for panels three and four. Because the character description never changes wording, the character's look stays steady across all four panels.
Character description (save this, reuse it exactly in all 4 prompts):
"A young woman in her twenties with short curly red hair, green eyes, light freckled skin, wearing a yellow raincoat and black rubber boots. She has a small star-shaped birthmark on her left cheek."
Panel 1:
"A young woman in her twenties with short curly red hair, green eyes, light freckled skin, wearing a yellow raincoat and black rubber boots. She has a small star-shaped birthmark on her left cheek. Show this exact character standing at a bus stop in the rain, looking at her watch. Comic book art style with bold outlines and flat colors. Keep the face, hair, and clothing identical to the description above."
Panel 2:
"A young woman in her twenties with short curly red hair, green eyes, light freckled skin, wearing a yellow raincoat and black rubber boots. She has a small star-shaped birthmark on her left cheek. Show this exact character running to catch a bus, splashing through a puddle. Comic book art style with bold outlines and flat colors. Keep the face, hair, and clothing identical to the description above."
Panel 3:
"A young woman in her twenties with short curly red hair, green eyes, light freckled skin, wearing a yellow raincoat and black rubber boots. She has a small star-shaped birthmark on her left cheek. Show this exact character sitting on the bus, smiling and looking out the window. Comic book art style with bold outlines and flat colors. Keep the face, hair, and clothing identical to the description above."
Panel 4:
"A young woman in her twenties with short curly red hair, green eyes, light freckled skin, wearing a yellow raincoat and black rubber boots. She has a small star-shaped birthmark on her left cheek. Show this exact character stepping off the bus into bright sunshine, folding her umbrella. Comic book art style with bold outlines and flat colors. Keep the face, hair, and clothing identical to the description above."

Example 2: An online store product shoot with five photos. Write out the full product description once, including material, color, shape, and label text. For each of the five photos, paste that same product description and only change the angle or setting sentence at the end, like "from the front on white" or "held outdoors in natural light." Number each photo in the prompt so the set stays organized.
Product description (save this, reuse it exactly in all 5 prompts):
"A matte black ceramic coffee mug with a curved handle, holding 350ml, with the label reading 'BREW & CO' printed in white in a clean sans-serif font on the front."
Photo 1:
"A matte black ceramic coffee mug with a curved handle, holding 350ml, with the label reading 'BREW & CO' printed in white in a clean sans-serif font on the front. This is photo 1 of a set. Show it straight on from the front, centered on a plain white background, studio lighting. Keep the product's shape, color, and label identical in every photo in this set."
Photo 2:
"A matte black ceramic coffee mug with a curved handle, holding 350ml, with the label reading 'BREW & CO' printed in white in a clean sans-serif font on the front. This is photo 2 of a set. Show it from a 45-degree angle on the same plain white background, studio lighting. Keep the product's shape, color, and label identical in every photo in this set."
Photo 3:
"A matte black ceramic coffee mug with a curved handle, holding 350ml, with the label reading 'BREW & CO' printed in white in a clean sans-serif font on the front. This is photo 3 of a set. Show it held in a hand outdoors, with a wooden table and morning light in the background. Keep the product's shape, color, and label identical in every photo in this set."
Photo 4:
"A matte black ceramic coffee mug with a curved handle, holding 350ml, with the label reading 'BREW & CO' printed in white in a clean sans-serif font on the front. This is photo 4 of a set. Show it from above, looking straight down into the mug, on a marble kitchen counter. Keep the product's shape, color, and label identical in every photo in this set."
Photo 5:
"A matte black ceramic coffee mug with a curved handle, holding 350ml, with the label reading 'BREW & CO' printed in white in a clean sans-serif font on the front. This is photo 5 of a set. Show it filled with coffee, steam rising, next to a small plate of pastries on a cafe table. Keep the product's shape, color, and label identical in every photo in this set."

Example 3: Turning a photo brief into three style options. Write a full, detailed description of a subject and scene as if it were a real photograph. Generate it three separate times, each time adding one extra sentence naming a different art style (watercolor, pencil sketch, oil painting) and repeating the instruction to keep the pose and identity the same. This produces three versions of the same scene in three different styles for a client to choose from.
Base photorealistic description (used in all 3, only the last sentence changes):
Watercolor version:
"A photorealistic image of an elderly man with a white beard, wearing round glasses and a brown wool sweater, sitting on a park bench reading a newspaper, autumn leaves falling around him. Now redraw this exact same subject, pose, and setting in the style of a soft watercolor painting. Keep the composition, pose, and identity of the subject exactly the same. Only the art style should change."
Pencil sketch version:
"A photorealistic image of an elderly man with a white beard, wearing round glasses and a brown wool sweater, sitting on a park bench reading a newspaper, autumn leaves falling around him. Now redraw this exact same subject, pose, and setting in the style of a detailed graphite pencil sketch. Keep the composition, pose, and identity of the subject exactly the same. Only the art style should change."
Oil painting version:
"A photorealistic image of an elderly man with a white beard, wearing round glasses and a brown wool sweater, sitting on a park bench reading a newspaper, autumn leaves falling around him. Now redraw this exact same subject, pose, and setting in the style of a classical oil painting with visible brushstrokes. Keep the composition, pose, and identity of the subject exactly the same. Only the art style should change."
Frequently Asked Questions
Why does my character's face keep changing between images?
The character description was likely reworded slightly each time. Copy and paste the exact same description every time instead of retyping it.
Can Gemini generate an image with a transparent background?
Describe the background explicitly as plain white or a solid color, then remove it afterward using an image editor, since asking directly for "transparent" in the prompt is not reliably followed by most image models.
Can it edit an image I already have, not just generate a new one?
Yes, this is one of its main strengths. Describe the image you have in detail, then describe the single change you want, and ask it to keep everything else the same.
Why does the output image change size or aspect ratio even though I didn't ask for that?
When you upload more than one reference image, the model sometimes bases the output's aspect ratio on one of the reference images instead of your prompt, and which reference image it picks is not always consistent. If the exact output size matters, state the aspect ratio directly in your prompt in words, like "16:9 widescreen" or "square 1:1," rather than relying on it to infer the ratio from your references.
Why does selecting only part of an image to edit give a stretched or distorted result?
This model tends to handle edits more reliably when it can see the entire image, not just a cropped selection. If a small selected area comes out stretched or misaligned, try feeding it the full image with a written instruction pointing to the area to change, instead of pre-cropping that area yourself.
Why do I get charged or lose credits even when the generation fails?
On some platforms that run this model, a failed generation (an error instead of an image) can still use up your credit or quota, since the cost is tied to the request being processed, not to a successful result. This is a known complaint on third-party tools and creative software that plug into the model, not something you can fix from the prompt side, so check that platform's status page if it happens often.
Can I get a Gemini AI Photo without the watermark?
Free, consumer-facing versions of the tool embed an invisible SynthID watermark plus a visible mark on some outputs, and there is no prompt that removes this. Paid API access or certain third-party platforms provide watermark-free output, since the watermark is tied to which access tier you're using, not to anything in your wording.
How many reference images can I actually combine in one gemini prompt?
Depending on the model version, you can mix multiple reference images (recent versions support well over ten) and use them for different jobs in the same prompt, like one for the character, one for the outfit, and one for the background style. If you're combining several references, say in the prompt which image each one is for, instead of leaving it for the model to guess.
Why does a batch of images from the same gemini prompt come out at different speeds or quality?
The higher-quality version of the model takes noticeably longer per image and costs more, while a faster version trades a bit of quality for speed and much higher throughput. If you're generating a large batch, this speed and quality tradeoff is a setting or model choice on the platform you're using, not something a prompt alone can fix.
Can it keep more than one character consistent in the same image at once?
Yes, but reliability drops as the number of characters goes up. It performs best with a handful of characters described in full detail; beyond that, some characters in a group scene tend to lose their exact details. If a scene has several characters, describe each one fully and separately in the prompt rather than describing the group as a whole.