Nano Banana 2 / Pro (Gemini Image)
Google's Gemini-family image model, nicknamed Nano Banana for its viral photorealistic edits, generates and edits images up to a 4K-capable tier.
What Nano Banana does
Nano Banana produces photorealistic images from a text prompt or edits an existing photo while preserving what shouldn't change — lighting, pose, background — and altering only what the prompt asks for. It's particularly strong at edits that need to look like they were shot in-camera rather than composited.
On AutorunX it's tuned for commercial imagery: clean product photography and marketing banner art where accurate text rendering, believable reflections/shadows, and brand-consistent color all matter.
Capabilities
Key features
- Photorealistic edit fidelity
- Image-to-image edits preserve the parts of a source photo the prompt doesn't touch, so a product swap, background change, or lighting tweak reads as a real photograph rather than an obvious composite.
- Dense-prompt instruction following
- Because generation runs through the same reasoning pathway as Gemini's text models, multi-constraint prompts — specific counts, poses, layouts, on-image copy — tend to hold up better than on diffusion-only models.
- Legible on-image text
- Renders logos, signage, and captions directly into the image more reliably than most diffusion-based competitors, which matters for banner ads that need real, spelled-correctly copy baked into the frame.
- Two-tier speed/quality trade-off
- Nano Banana 2 targets fast, everyday generations; Nano Banana Pro spends more compute for the highest-fidelity, most complex compositions — pick per-shot on AutorunX depending on the job.
- 4K-capable output
- Scales up to a 4K-capable resolution tier, enough headroom for large-format banner crops and print-adjacent product photography.
How Nano Banana works
Nano Banana is Google's public nickname for its Gemini-family image model, now offered on AutorunX in two tiers: Nano Banana 2 for fast everyday generations and Nano Banana Pro for the highest-fidelity, most complex compositions. Both generate images inside the same multimodal Gemini architecture that handles Google's text and reasoning models, rather than bolting a diffusion model onto a separate text encoder.
Because image generation shares the model's reasoning pathway, Nano Banana can apply a "thinking" step before rendering — parsing a dense prompt's constraints (exact object counts, specific poses, layout instructions, legible on-image text) and holding onto them more reliably than diffusion-only competitors typically manage.
It runs in two modes on AutorunX: text-to-image generation from a written prompt, and image-to-image editing, where an existing photo is transformed while the untouched regions stay pixel-consistent with the source. Output can reach a 4K-capable resolution tier, and the model is known for strong semantic/world knowledge — it tends to get real-world details (brand logos, correct text spelling, plausible physics) right where older diffusion models guess.
What people use Nano Banana for
E-commerce product photography
Turn a plain product photo into a styled lifestyle shot — new background, staging, or lighting — while keeping the product itself untouched and accurate.
Banner ad creative
Generate ad banners with real, legible on-image copy and logos baked in, rather than needing to overlay text in a separate editor.
Photo cleanup and restyling
Edit an existing brand or campaign photo — swap a background, adjust lighting, remove an object — without disturbing the rest of the frame.
Rapid creative variations
Use the faster Nano Banana 2 tier to iterate through many prompt variations quickly, then finish the chosen direction on Nano Banana Pro.
How to use Nano Banana on AutorunX
Nano Banana is available in Product Shots, in the Image module.
- 01
Open Product Shots or Banner Ad in Image
Head to the Image and pick the service you need — Product Shots for e-commerce imagery, Banner Ad for ad creative.
- 02
Set up your input
Write your prompt, or upload a source photo for image-to-image editing, then add any product or brand reference details.
- 03
Pick Nano Banana in the model picker
Open the model picker and select Nano Banana 2 or Nano Banana Pro depending on how much fidelity the shot needs.
- 04
Generate
Run the generation and review the result — re-prompt or switch tiers if any constraint didn't land.
- Billed per image from your shared AutorunX credit wallet.
- 37 credits / image (Nano Banana 2) · 80 credits / image (Nano Banana Pro)
Tips for better results
- Front-load hard constraints
- Put exact counts, positions, and text strings early in the prompt — Nano Banana's reasoning step weighs earlier constraints more heavily when a prompt gets long.
- Use image-to-image for edits, not regeneration
- When you only need to change one thing in an existing photo, feed it as the source image rather than re-describing the whole scene from scratch — it preserves everything else far better.
- Spell out on-image text exactly
- Quote the exact copy you want rendered (product name, price, headline) in the prompt — Nano Banana's text rendering is strong, but only for text it's explicitly told to draw.
- Reach for Pro on complex composites
- If a shot needs multiple constraints satisfied at once — precise layout, brand colors, and legible text together — switch to Nano Banana Pro rather than iterating on the faster tier.
Nano Banana — frequently asked
Related models
Seedream
ByteDance / BytePlusByteDance's Seedream generates and edits images from up to ten reference photos, holding identity and composition steady across a full shoot.
GPT Image 2
OpenAIOpenAI's natively multimodal image model — the default engine behind Face / Identity.
Qwen-Image-Edit
AutorunX ownedAlibaba's open-weight 20B image editor, self-hosted by AutorunX as AX-QWN, preserves identity through fine-grained edits at near-zero marginal cost.
Ready to create with Nano Banana?
Open Product Shots and generate your first image in a couple of minutes.
Free plan, no card required.