← Blog

Gemini Image Prompts: Photo Styles, Edits, and Ratios

Gemini Image Prompts: Photo Styles, Edits, and Ratios

Strong Gemini image prompts describe a full scene in sentences: subject, action, setting, style or medium, lighting, camera or lens, and an explicit aspect ratio. Gemini's native image model also supports conversational editing, multi-image blending, consistency across turns in one thread, and direct aspect ratio control. Below are photo-style prompt examples plus a 25-prompt library, with every model capability checked against Google's own documentation as of 2026.

How Gemini Generates Images in 2026

“Nano Banana” is the nickname that stuck to Gemini’s native image generation. As of 2026 there are two tiers you actually need to care about:

  • Nano Banana, the fast Flash-class image model, is the default in the Gemini app. It is the volume workhorse: ad variations, social posts, quick iteration, throwaway concepts.
  • Nano Banana Pro, built on Gemini 3 Pro, reasons before it renders, outputs up to 4K, blends up to 14 reference images, and renders legible text inside the frame. That last capability is the one that changes workflows, because posters, packaging comps and labelled infographics stop being a Photoshop step.

Three access points, same prompt craft: the Gemini app for hands-on work, Google AI Studio for testing prompts side by side, and the Gemini API image generation docs when you want to generate creative variants programmatically.

Every image Google’s models produce carries an invisible SynthID watermark for AI provenance, and Nano Banana Pro can also tell you whether an image it is shown was made by Google AI. Model names churn every few months. The structure of a good prompt does not, which is why the rest of this guide is built around the structure.

Gemini AI Photo Prompts for Photorealistic Results

We promoted this section to the top of the guide on purpose. Over the last 28 days this page pulled 2,891 impressions at an average position of 11.9, but the query “gemini ai photo prompt” sat at position 22.06, still outside the top 10. People arriving on photo-prompt intent were scrolling through stylized art direction before they hit anything about cameras. So the photography half comes first now.

Photorealism comes from photography vocabulary, not from bolting “realistic, 8K, hyper detailed” onto the end of a sentence. Name the gear, the light and the constraint. Our photo template:

Subject and expression + location + lens and aperture + light source and direction + color or film character + aspect ratio.

ShotPrompt
Candid portraitA candid half-length portrait of a woman in her thirties in a linen shirt, laughing and glancing off-camera, in a cafe with a blurred window behind her. Shot on an 85mm lens at f/1.4, natural side light from the left, warm skin tones, fine film grain, no retouching gloss. Vertical 4:5.
Studio headshotA tight editorial headshot of a man with close-cropped grey hair against a mid-grey seamless backdrop, direct gaze, slight smile. One large softbox high and to camera-right with a subtle fill card below, 105mm lens at f/5.6, visible skin texture, neutral white balance. Vertical 4:5.
Product macroA stainless steel watch crown filling the frame, machining marks and micro-scratches visible, single hard key light raking across the metal from the right, black seamless backdrop. Macro lens at f/8, focus stacked, cool neutral white balance. Square 1:1.
LandscapeA wide basalt coastline under breaking storm clouds at first light, long exposure smoothing the surf into soft white bands. Shot on a 24mm lens at f/11, tripod, graduated exposure, muted cold palette, high tonal separation in the mid-greys. Ultrawide 21:9.
InteriorAn empty modern office corner at mid-morning, pale oak floor, one linen armchair, tall window casting hard-edged light shapes across the wall. Shot on a 17mm tilt-shift lens at f/8, verticals corrected, natural light only, low-contrast neutral grade. Horizontal 16:9.
FoodAn overhead flat lay of a bowl of ramen on a dark ceramic tile, steam rising, chopsticks resting on the rim, scattered scallion and chili oil beads. Diffused window light from the top left, 50mm lens, natural shadows, no props clutter. Square 1:1.
StreetA rain-slicked city crosswalk at night, a single figure with an umbrella mid-stride, neon signage reflected in the puddles. Shot on a 35mm lens at f/2, high ISO grain, mixed tungsten and neon color cast, motion blur only on the background traffic. Horizontal 16:9.
App lifestyleA person on a commuter train holding a phone at a natural angle, screen glow lighting their face, blurred window landscape streaking past. 35mm lens at f/2.8, available light only, the phone screen left blank grey for a UI composite. Vertical 9:16.

Four failure modes and their fixes:

  • Plastic skin. Add “visible skin texture and pores, natural blemishes, no beauty retouching” and specify available light instead of studio light.
  • Impossible lighting. Name one light source and its direction. Multiple conflicting cues produce that uncanny composite look.
  • Everything in focus. Real photographs have a focal plane. State the aperture and what is sharp: “sharp on the eyes, background falling off at f/1.8”.
  • Text and hands mangled in the background. Say what to exclude in plain language: “no signage text, no visible logos, hands out of frame”. Gemini has no separate negative prompt field, so exclusions belong inside the sentence.

If those photo assets are heading into paid social or store creative, tie them to measurement before you scale spend. Our breakdowns of the mobile app KPIs worth tracking and how Adjust and AppsFlyer compare cover the attribution side of creative testing.

The Prompt Formula That Works

Google’s own prompting guidance is blunt about this: describe the scene in sentences, do not list disconnected keywords. The model is reading for meaning, not matching tags.

Our working formula has seven slots:

Subject + action + setting + style or medium + lighting + camera and composition + aspect ratio.

Here is the same idea written badly and written properly.

VersionPrompt
Weak keyword listcoffee cup, desk, nice lighting, 8K, professional
Formula-built promptA matte ceramic espresso cup resting on a scratched walnut desk beside an open notebook, steam curling into a shaft of low morning light from a window on the left. Editorial product photography, shallow depth of field, warm neutral palette, shot on a 50mm lens at f/1.8, negative space on the right for a headline. Horizontal 16:9.

The second one gives the model a subject, a light direction, a lens, a layout constraint and a delivery format. That is why it comes back usable instead of generic.

Two habits that raise the hit rate immediately:

  1. Say what the image is for. “Leave the upper third empty for a headline” or “product must be fully inside the safe area for a vertical story” gets respected more often than you expect.
  2. Iterate in one thread. Gemini supports multi-turn editing, so “keep everything, change the cup to a glass tumbler” preserves the composition. Starting a new chat throws the look away.

25 Gemini Image Prompt Examples

Paste these as-is, then refine conversationally. They are grouped by the job the image has to do.

Brand concepts and emotional storytelling

Use these when the product is boring but the promise is not.

PromptWhy it works
A massive ancient stone archway standing perfectly intact in the middle of a swirling misty landscape at sunrise, faint warm glow on the stone. Cinematic, deep focus, volumetric light, muted earth palette. Horizontal 16:9.Stability against chaos, without a single stock handshake. Built for insurance, fintech and security.
A glowing geometric wireframe structure assembling itself mid-air inside an infinite dark void, cool electric blue emission, trailing particles where new edges snap into place. Futuristic digital art, high contrast, 4K.Reads as innovation and self-building systems. Good for launch hero images.
A person on a windy beach with arms outstretched, dissolving into hundreds of translucent colored ribbons blowing out to sea. Dreamlike, artistic motion blur, pastel palette, soft backlight. Vertical 4:5.Freedom and release, abstract enough to avoid casting-model problems.
A single crystal-clear raindrop resting on the tip of a vibrant green leaf after a storm, water beading on the leaf surface. Hyper-realistic macro photography, high contrast, natural window light, black background. Square 1:1.Clarity and precision for minimalist products and design studios.
A tight aerial view at dusk of dozens of small warm window lights glowing through a fog-covered valley, the lights forming an intricate connected pattern. Moody high-angle composition, deep blue and amber palette. Horizontal 16:9.Community and network effects. Works for telecom, marketplaces and nonprofits.

Product and mockup visuals

PromptWhy it works
A vintage hand-stitched leather camera bag floating weightlessly in a zero-gravity studio, surrounded by perfect spheres of colored liquid, rolled paper maps and instant photos drifting nearby. Diffused white light, hyper-realistic render, seamless light-grey backdrop. Square 1:1.Premium and exploratory at once, and the floating props give you crop options.
A rugged hiking boot planted on a moss-covered stalagmite deep inside a crystal-lined cavern, a narrow sunbeam cutting through a hole in the ceiling and hitting the boot. Cinematic adventure photography, volumetric god rays, damp textures. Horizontal 16:9.Durability shown by environment instead of copy.
A glass dome shielding a single steaming cup of herbal tea from a downpour of giant iridescent shards, the tea perfectly still. High-speed photography look, macro lens, sharp focus on the rising steam, dark background. Vertical 9:16.Protection and calm, tailor-made for a story ad.
A high-performance electric scooter parked in a field of tall yellow wheat with minimalist solar arrays on the far horizon at golden hour. Analog film grain, wide cinematic framing, warm naturalistic color. Ultrawide 21:9.Merges green tech and nostalgia for a banner-shaped crop.
A matte black wireless earbud charging case on a sun-drenched Paris cafe table, the whole image rendered in swirling thick impasto brushstrokes in the palette of Vincent van Gogh. Oil on canvas texture, visible knife marks. Square 1:1.Style transfer makes a commodity product memorable in a crowded feed.

Illustration and style transfer

PromptWhy it works
A mesh router rendered as an intricate blue-and-white mid-century modern architectural blueprint, internal components drawn as labelled abstract diagrams. Clean line work, technical illustration style, white background. Horizontal 16:9.Signals engineering rigor for B2B and infrastructure pages.
A bustling 16-bit isometric pixel-art city where each building is a piece of software, a spreadsheet tower, a chat-app block, a database vault, with one tiny avatar climbing the tallest one. Retro limited palette, crisp pixel edges. Square 1:1.Native visual language for developer and gaming audiences.
A minimalist reusable steel water bottle depicted as a detailed Ukiyo-e woodblock print, a towering stylized wave behind it and a silhouetted figure holding it in the foreground. Hand-carved texture, indigo and vermilion Edo palette. Vertical 2:3.Heritage and sustainability without a single leaf icon.
An abstract Bauhaus poster built from circles, triangles and hard rules in primary red, blue and yellow, with one short line of sans-serif copy reading “Function First”. Flat graphic design, high saturation, generous margins. Vertical 4:5.Tests the model’s typography while staying on-brand for design-led products.

Text-in-image, posters and packaging

These are the prompts where Nano Banana Pro earns its keep.

PromptWhy it works
A minimalist event poster on textured off-white stock, large geometric sans-serif headline reading “GROWTH LOOPS”, subhead “A one-day workshop, 14 March”, small logo lockup bottom left, single flame-orange accent shape. Flat print design, subtle paper grain. Vertical 2:3.Legible multi-line text and hierarchy in one pass.
A matte black supplement pouch standing on a dark stone slab, front label reading “DAILY FOCUS” in clean uppercase with a small ingredient list beneath, soft rim light from behind, faint reflection on the stone. Studio packaging photography. Square 1:1.Packaging comps you can show a client before dieline work starts.
A three-step process infographic on a deep navy background, each step a numbered circle with a short label, “Capture”, “Enrich”, “Activate”, connected by a thin flowing line, muted teal and amber accents. Clean vector style, no clutter. Horizontal 16:9.Diagram plus real labels, the classic AI-image failure mode, now solvable.
A dark UI-style comparison card with two columns headed “Manual” and “Automated”, four short bullet rows in each, checkmarks in the right column, thin dividers, soft glow behind the winning column. Modern product-marketing graphic. Horizontal 16:9.Feature comparison graphics for landing pages, no design tool needed.

Social and ad creative

PromptWhy it works
A bold central subject cropped mid-motion so it breaks the top edge of the frame, exaggerated wide-angle perspective, clashing saturated color blocks behind it. Designed for vertical 9:16 with the focal point in the upper third.Built for thumb-stopping in a vertical feed rather than resized afterwards.
A calm professional scene with a single symbolic object placed left of center, wide clean negative space to the right, muted corporate gradient background, restrained cool palette. Horizontal 16:9.The reliable B2B hero for webinars, reports and whitepaper covers.
A wide breathable hero composition with a strong subject on the right third and empty soft-gradient space on the left, gentle top-down lighting, modern UI-friendly color temperature. Ultrawide 21:9.Above-the-fold layouts where copy and a form need room.
One scene split vertically down the middle: left side cluttered, grey and chaotic, right side ordered, bright and spacious, a faint seam running between them. Ultra-wide cinematic composition. Horizontal 16:9.Before and after in a single frame, understandable with the sound off.

If you are generating these for a live ad account, volume and cadence matter as much as the art. Run the free Creative Velocity Audit to see whether enough fresh creative is actually shipping each week to keep the account learning.

Data visuals and conceptual infographics

PromptWhy it works
A polished chrome kinetic sculpture shaped as an infinity loop, its surface etched with faintly glowing lines suggesting data flowing through a continuous process. Dramatic studio lighting, highly reflective, dark seamless background. Square 1:1.Abstract efficiency for process and workflow products.
An illuminated three-dimensional cityscape built entirely from neon bar graphs and line charts, the tallest towers glowing red and green, one small figure watching from a high ledge. Cyberpunk aesthetic, deep black sky, 4K.Analytics and BI imagery that is not another laptop-with-dashboard photo.
A macro view of a topographical map sculpted from thick flowing pastel oil paint, deep blue valleys and bright gold peaks. Heavy impasto texture, macro lens, subtle bokeh at the frame edges. Horizontal 16:9.Turns an emotional or index-style metric into something worth pinning.

Editing Images and Aspect Ratios

Generation is the easy half. Editing is where Gemini beats a prompt-and-pray workflow.

Per the Gemini API image generation documentation, you can:

  • Upload your own photo and edit it conversationally. Add, remove or replace elements, change a background, restyle a whole frame, all in natural language across multiple turns.
  • Blend multiple inputs. Nano Banana Pro accepts up to 14 reference images, which is how you place a real product photo into a generated scene instead of hoping the model recreates your packaging.
  • Hold consistency. Character and product appearance carry across turns in the same conversation, which is the practical substitute for the seed numbers Midjourney and Stable Diffusion expose. Ask for “the same scene, same lighting, swap the product” rather than opening a new thread.
  • Choose an aspect ratio explicitly. State it in the prompt in the app, or set it as a parameter through the API. The Midjourney --ar flag is not Gemini syntax.

Ten aspect ratios are supported in the API, which maps cleanly onto channels:

RatioWhere it goes
1:1Instagram grid, app icons, thumbnails
4:5 and 5:4Feed posts and paid social with maximum vertical real estate
9:16Stories, Reels, TikTok, YouTube Shorts, app store screenshots
16:9Site headers, YouTube thumbnails, presentation slides
21:9Ultrawide banners and above-the-fold hero strips
3:2, 2:3, 3:4, 4:3Print-shaped crops, editorial layouts, posters

Nano Banana Pro also lets you pick output resolution up to 4K, which matters the moment an asset is going to print or onto a retina hero section.

Three edit prompts we reuse constantly

  • Single-element swap: “Keep the composition, lighting and grade exactly as they are. Replace the ceramic mug with a clear glass tumbler of iced coffee, matching the existing light direction and reflections.”
  • Copy space rescue: “Extend the canvas to the left and keep the subject where it is, filling the new area with the same out-of-focus background so a headline fits.”
  • Product placement from a real photo: “Using the uploaded packshot as the exact product reference, place it on the stone slab in the second reference image. Do not redraw the label.”

Access Tiers, Limits and Which Plan You Need

Google gates image volume by plan rather than by feature, and it retunes the caps with demand. Check Google’s AI plans page for today’s numbers before you commit a campaign to a deadline.

TierWhat you get, as of 2026
Free Gemini accountNano Banana image generation with a small daily allowance, around 20 images a day, and limited access to the Pro model. Fine for exploration, not for a campaign sprint.
Google AI ProSubstantially higher daily image limits plus fuller Nano Banana Pro access. This is the tier most marketing teams actually need.
Google AI UltraThe highest limits and earliest access to new image models, aimed at heavy production use.
Gemini API and AI StudioPay-per-use image generation with programmatic aspect ratio and resolution control. See the Gemini API pricing page for current per-image rates.

Two caveats worth planning around. Availability of specific models and features varies by region and by account type, and Google ships changes to this lineup every few months. Never build a client deliverable that depends on a daily cap you have not confirmed inside the product that week.

FAQ: Gemini Image Generation

Is Gemini image generation free?

Yes, with limits. A free Google account can generate images in the Gemini app under a modest daily cap, roughly 20 a day as of 2026, with restricted access to the higher-quality Pro model. Google AI Pro and Ultra raise those caps considerably. The current numbers live on Google’s AI plans page.

What is the best prompt structure for realistic photos?

Describe the subject and what it is doing, then the location, then the lens and aperture, then one light source and its direction, then the color or film character, then the aspect ratio. Photography vocabulary does the work. Words like “8K” and “hyper detailed” do not.

Can Gemini edit my own uploaded photos?

Yes. Upload an image and describe the change in plain language: swap a background, remove an object, restyle the frame, extend the canvas. Edits are multi-turn, so you can keep refining the same asset while the model holds the rest of the composition steady.

Does Gemini watermark AI images?

Every image generated by Google’s models carries an invisible SynthID watermark identifying it as AI-generated, and visible watermarking behavior varies by plan and surface. Assume provenance is detectable, and disclose AI-generated assets wherever the platform, the client contract or local law requires it.

Do these prompts work in ChatGPT or Midjourney?

Mostly, yes. The seven-slot structure of subject, action, setting, style, lighting, camera and aspect ratio transfers to every major generator. Only syntax changes: Midjourney uses flags like --ar and seed numbers, Gemini and ChatGPT both use conversational refinement instead.

Which model should I use for images with text in them?

Nano Banana Pro. Text rendering inside images was the standing weakness of every image model through 2024 and 2025, and the Pro tier is where Google fixed it. For posters, packaging labels and labelled diagrams, take the slower, more expensive model.

What Changed in This Update, and Why

This refresh was driven by our own search data, not a calendar reminder. Over the last 28 days the site recorded 799 clicks from 116,662 impressions, and this page was one of the strongest single URLs: 2,891 impressions, 107 clicks, average position 11.9, and 91 human sessions. The previous refresh moved the page up. The photo query did not move with it. “gemini ai photo prompt” is still stuck at position 22.06, outside the top 10, which tells us the photography material was buried too deep.

So in this pass we:

  • Moved the photorealistic photo-prompt section above the stylized prompt library, so the page answers photo intent in the first screen instead of the fourth.
  • Expanded the photo section from six shots to eight, adding a studio headshot and an interior architecture setup, and added a reusable photo template plus a fourth failure mode about focal planes.
  • Added three named edit prompts we reuse for element swaps, copy space and real-packshot placement, all built on the multi-turn editing behavior documented in the Gemini API image reference.
  • Added a prompt-structure FAQ entry aimed squarely at the photo query.
  • Refreshed every performance number in the article to the current 28-day window.

Earlier passes rebuilt the library into 25 job-based prompts, added the access-tier table, dropped price claims we could not source to Google’s own plans page, and removed the named-brand prompt sets, because generating imagery that mimics another company’s trademarks is a legal problem rather than a creative flex. Those changes stand.

The wider point stands too. Execution is no longer the bottleneck; articulation is. If you are planning creative volume for 2026, our mobile app marketing statistics roundup gives the channel context, and how AEO, GEO and SEO differ covers what happens to your content once AI surfaces start deciding who gets cited.

Save this page. We update it when Google ships a real change to the image models, or when our own data says the page is answering the wrong half of the question.

Want AI doing this for your growth?

We build AI-driven acquisition, content, and automation systems for operators across North America. See your levers in 30 minutes.

Book a growth audit Rated 5.0 on Clutch

Explore AI automation services →