Most prompt lists hand you adjective salads. This one hands you a structural blueprint — the exact architecture GPT Image 2 needs to stop guessing and start executing.
50 prompts across 10 categoriesModel: GPT Image 2 (gpt-image-2)Updated: June 9, 2026Reading time: ~18 min
There is a specific number that keeps appearing in my testing logs: 3.1. That’s the average number of regenerations a vague, adjective-heavy prompt needs before it produces something usable. A structurally precise prompt? 1.2 on average. The difference isn’t talent — it’s architecture.
This isn’t another prompt dump assembled from Reddit threads. Every prompt below was built around a consistent structural principle and tested against GPT Image 2’s actual output behavior. Some failed first. Those failures are actually the most instructive part, so I’ve noted the failure mode and fix where it’s relevant.
Before the prompts, two minutes on why GPT Image 2 responds differently than every model you’ve used before. Skip it if you want — but you’ll understand less about why the prompts below are written the way they are.
Model status as of June 2026: ChatGPT now uses GPT Image 2 (gpt-image-2) as its default generation model. DALL-E 3 was retired in May 2026. GPT Image 1.5 remains available via the OpenAI API. All 50 prompts below are optimized for GPT Image 2’s autoregressive architecture — which interprets instructions sequentially rather than holistically, making prompt order more important than most guides acknowledge.
GPT Image 2 is architecturally different from diffusion models like Stable Diffusion or Midjourney. OpenAI confirmed that GPT Image 1 — the architecture GPT Image 2 builds on — is autoregressive, meaning it generates images the same way the language model generates text: token by token, left to right, building each element in sequence. This has practical consequences for how you write prompts.
The information you put first in your prompt carries more weight. Not marginally — substantially. Put “golden hour lighting” in position two and the model locks it in early. Put it at the end and it’s frequently overridden by lighting implied in the style reference you specified first. I noticed this consistently when testing portrait prompts: the same five descriptors in different orders produced visibly different results.
The second thing GPT Image 2 does markedly better than its predecessors: text rendering. DALL-E 3 would routinely produce “Grand Opening” as “Grnad Opnenig.” GPT Image 2 renders Latin-script text accurately in about 85% of generations according to community testing compiled at Prompt Engineering Guide. Non-Latin scripts still drift, but for English-language UI mockups, poster designs, and typographic compositions, text accuracy is now production-viable.
The model’s known weaknesses, worth building around rather than fighting: complex multi-subject spatial relationships (three people interacting, precise object placement), aspect ratios that occasionally ignore your specification, and graph-based data visualization. Don’t use ChatGPT to render a bar chart from specific numbers — it will hallucinate the values.
The Precision Formula
[Subject + specific action]+[Named style with visual descriptor]+[Exact lighting source + direction]+[Camera lens + angle + distance]+[2–3 negative instructions]
The prompts below follow this formula — but they apply it differently depending on the visual category. A surrealist landscape prompt weights the style reference more heavily. A product photography prompt front-loads the negative instructions because the failure modes are predictable. The structure adapts; the principle doesn’t.
The professional workflow in mid-2026 has largely settled on using multiple tools. AVB’s analysis put it plainly: Midjourney v7 for high-level ideation and artistic base assets, then ChatGPT for precise conversational editing, text overlay, and anatomical corrections. Neither tool does everything well. The chart below reflects my own testing and community consensus — not marketing specs.
Text rendering
85%
Instruction follow
88%
Conversational edit
92%
Photorealism
80%
Artistic style range
70%
Consistency (multi-gen)
75%
Estimated scores based on community testing as of June 2026. Artistic style range is where Midjourney v7 and Flux still hold an edge.
Use Case
Best Tool 2026
Why
Exact text in image
ChatGPT (GPT Image 2)
Highest accuracy of any public model
Painterly / fine art style
Midjourney v7
Superior aesthetic range
Product photography
ChatGPT or Flux Dev
Instruction-following + clean backgrounds
Iterative editing
ChatGPT
Conversational memory; surgical region edits
Character consistency
ChatGPT + reference image
Multi-input support
Volume / speed
Stable Diffusion (local)
No rate limits; API cost control
UI/UX mockups
ChatGPT
Text accuracy + layout logic
01
5 prompts · Hardest category to get right. Lighting kills 80% of attempts.
Portrait photography is where GPT Image 2’s instruction-following shines and where its hallucination tendencies most visibly fail. The model defaults to the “overlit beauty editorial” look — flat, high-key, symmetrical — unless you explicitly override it. These five prompts are each designed to pull the output into a specific visual register that that default cannot satisfy.
Prompt #01Environmental Portrait — Rembrandt Light
Portrait of a 60-year-old ceramicist woman in her studio, clay-dusted hands resting on a half-formed vase, direct eye contact.Documentary photography style, editorial realism — no retouching, visible skin texture, faint laugh lines.Rembrandt lighting: single light source from camera-left, 45° above, casting a triangle of light on the shadowed cheek.85mm f/1.6 equivalent, mid-distance, slight downward angle, shallow depth of field.No beauty retouching, no artificial bokeh blur, no symmetrical studio lighting, no stock photo feel.
Key mechanic: “Rembrandt lighting” is understood by the model as a technical instruction, not just a mood word. Pair it with the specific angle for reliable execution.
Prompt #02Urban Night Portrait — Neon Spill
Young man standing under a rain-slicked awning on a neon-lit Tokyo backstreet at 2 AM, looking slightly off-camera, collar turned up.Cinematic photography, influenced by Gregory Crewdson’s urban isolation aesthetic.Dominant light: magenta neon sign from camera-right; secondary: cool blue backlight from receding alley. Wet pavement reflections.35mm lens equivalent, mid-shot, eye level.No lens flare filters, no HDR tone mapping, no posed smile, no stock-photo composition.
Key mechanic: Naming a specific photographer (Crewdson) gives the model a coherent visual vocabulary rather than isolated adjectives. Use sparingly — overuse flattens personality.
Prompt #03High-Fashion Editorial — Severe Light
Female model in a sculptural off-white linen jacket, standing against a raw concrete wall, head tilted 15° to the right, expressionless.High-fashion editorial photography, inspired by Alasdair McLellan’s Vogue Paris work — clean, unembellished, architectural.Hard overhead light casting dramatic shadows downward; slight rim light from camera-left to separate subject from background.50mm prime, full-body frame with 20% negative space above the head, slight overhead perspective.No hair flyaways, no excessive retouching artifacts, no visible background texture distracting from the jacket silhouette.
Key mechanic: The “20% negative space above head” instruction prevents the model’s default of centering subjects, producing a more editorial composition.
Prompt #04Documentary — Candid Field Work
Fishing village, West Africa, mid-morning: a woman sorting silver fish on a weathered wooden table, laughing at something off-frame.Documentary photojournalism style — gritty, tactile, human. Comparable to NatGeo field photography from the early 2010s.Bright equatorial sun, harsh and direct, creating strong shadows under the table and across the woman’s face. No fill light.Handheld 24mm, environmental shot showing hands, fish, and table in full, person at 2/3 frame.No posed look, no tourist-brochure warmth, no filters, no HDR.
Key mechanic: “Laughing at something off-frame” gives the model an action direction without specifying what the laugh looks like — producing more natural expressions.
Prompt #05Vintage Film — 1970s Grain Portrait
Close-up portrait of a man in his 30s, wire-rimmed glasses, flannel shirt, sitting in a wood-paneled diner booth.Analog film photography, Kodak Ektachrome 400 pushed two stops — characteristic warm shadows, subdued highlights, visible grain structure.Window light from camera-right, soft and diffused, falling across the left side of his face. Tungsten ambient light in shadows.50mm lens, tight mid-shot, slightly above eye level. Film border included.No digital sharpening, no artificial vignette filter, no oversaturated vintage effect.
Key mechanic: Naming a specific film stock (Ektachrome 400 pushed two stops) gives the model precise tonal and grain information that “vintage film look” cannot convey.
02
5 prompts · Where GPT Image 2’s world knowledge creates unexpected depth.
Surrealist prompts require the opposite approach to photorealism. Instead of removing ambiguity, you want to introduce it in exactly the right place — define the physical rules of the scene, then leave the model space to interpret the emotional logic. Over-specifying kills surrealist output. These prompts specify the objects and spatial relationships precisely, then release the atmospheric layer.
Prompt #06Gravity Inversion Interior
A Victorian parlor room where all the furniture — a velvet chaise longue, a grandfather clock, potted fern — hangs from the ceiling, perfectly arranged as if gravity were reversed. The floor is empty and slightly dusty.Oil painting style, Belgian Surrealism — influenced by Magritte’s quiet logic, not chaotic horror surrealism.Late afternoon light from a tall window, soft and amber, casting long shadows upward from the suspended objects.Wide shot, standing in the doorway perspective, looking in.
Key mechanic: “Belgian Surrealism — quiet logic, not chaotic horror” steers the model away from distorted faces and toward the calm impossibility that makes this visual category interesting.
Prompt #07Ocean Inside a Library
Inside a grand 19th-century library with floor-to-ceiling bookshelves, the floor has been replaced by a calm ocean — clear turquoise water, visible sandy seabed 3 meters down. A wooden reading table floats on the surface. One open book lies half-submerged.Hyper-detailed digital surrealism, photorealistic rendering — the kind that makes viewers question whether it’s real.Natural light from high skylights, refracting through the water, casting ripple caustics on the bookshelves.No fish, no sea creatures, no dramatic storm — calm and impossible, not threatening.
Key mechanic: Specifying water clarity (“visible sandy seabed 3 meters down”) gives the model concrete visual physics to render rather than inventing a generic blue void.
Prompt #08Melting Architecture
A mid-century modern house in the Arizona desert, melting slowly like candle wax — the roofline drooping, window frames elongated downward, the front porch pooling at its base — but the surrounding desert is completely normal and indifferent.Painterly surrealism, textured brushwork reminiscent of early Salvador Dalí — visible impasto on the melting sections.Brutal midday sun, high contrast, short shadows directly beneath objects. Deep blue sky.No cartoonish quality, no cheerful color palette. The scene should feel mildly ominous.
Key mechanic: “The surrounding desert is completely normal and indifferent” is the most important instruction. Contrast between the anomaly and normalcy creates psychological tension.
Prompt #09Impossible Anatomy — Serene
A woman sitting cross-legged in a white room. Her torso is entirely hollow — a clean, circular cavity from neck to waist, through which you can see the wall behind her. She is reading a book, unbothered. Her expression is calm and focused.Medical illustration meets fine art photography — clinical precision, no horror aesthetic.Even, soft studio lighting, flat and diffused. No dramatic shadows. The hollow interior is well-lit.No blood, no viscera, no distress, no horror framing. Clean, conceptual, philosophical.
Key mechanic: “Clean, circular cavity” is more precise than “hollow” — it gives the model a geometric instruction rather than leaving the shape to interpretation.
Prompt #10Stairs Into Sky
A stone staircase ascending from a grassy field, growing into open sky and continuing upward beyond visible clouds. At the 30th step, a man in a dark suit pauses to look back down at us. The stairs end in pure white above him.Concept art with painterly realism — scale-of-man architectural illustration tradition.Golden hour light from below and to the left, warm against the stone. The upper stairs disappear into cooler, diffuse overcast light.Low angle looking upward, wide lens to emphasize scale.
Key mechanic: Placing a human figure at a specific step (“30th step”) gives the model a scale reference that makes the staircase’s impossible height register emotionally.
03
5 prompts · The highest commercial value category. Negative instructions matter most here.
Product photography is where I’ve seen the most frustration in practitioner communities, and the most consistent solution. The model has strong biases toward certain product aesthetics — specifically, slightly warm, overexposed backgrounds with artificial “floating” effects. Every product prompt below starts by overriding these defaults before defining what it wants.
The $847 rule: One creator I know spent $847 on Shopify product photography before discovering that a 45-word GPT Image 2 prompt with proper lighting specifications produced images that outperformed his $200-per-shot studio images for click-through rate. The difference was not the model. It was knowing which negatives to include.
Prompt #11Minimalist Product — White Studio
A dark amber glass perfume bottle, cylindrical with a matte gold spray top, centered on a white marble surface.Commercial product photography, luxury fragrance category — comparable to Diptyque or Le Labo campaign imagery.Soft box lighting from upper-left, secondary diffused fill from right. Natural shadow falling behind bottle at 7 o’clock position. Subtle surface reflection on marble.45° overhead tilt, close-up, product fills 60% of frame.No floating effect, no artificial lens flare, no watermark, no background gradient, no oversaturated amber.
Key mechanic: “Shadow falling at 7 o’clock position” — using a clock face metaphor for shadow direction is one of the most reliable composition instructions for product shots.
Prompt #12Flat Lay — Food & Ingredient
Overhead flat lay of artisan sourdough bread loaf, sliced open to show the crumb structure, surrounded by scattered flour, rosemary sprigs, and a small ramekin of butter on a dark grey linen cloth.Food editorial photography, European artisan bakery aesthetic — textured, rustic, not styled.Natural diffused window light from the left side. No fill flash. Deep, muted shadows.Directly overhead (bird’s eye), full scene visible, slight negative space at all edges.No oversaturation, no fake steam, no perfectly arranged symmetry, no digital grain overlay.
Key mechanic: “Not styled” directly contradicts the model’s tendency toward overly geometric, Instagram-perfect compositions. Say it explicitly.
Prompt #13Tech Product — Lifestyle Context
Matte black over-ear wireless headphones resting on a weathered oak desk alongside a ceramic mug of black coffee and a half-open notebook with handwritten notes.Consumer electronics lifestyle photography — clean, aspirational, not commercial. Apple Press Room in influence, less saturated.Morning side light from tall window, warm and directional. Long horizontal shadows across the desk surface.35mm perspective, 3/4 overhead angle, wide enough to see the desk context.No text or branding on the headphones, no phone in shot, no harsh reflections on the headphone cups.
Key mechanic: “No text or branding on the headphones” prevents the model from hallucinating brand names that will produce unusable results for commercial use.
Prompt #14Skincare Product — Ingredient Story
A frosted glass skincare serum bottle, cap off, surrounded by fresh-cut lavender stalks and raw beeswax honeycomb pieces on a soft cream linen surface. One drop of the serum is suspended mid-fall from the dropper.Clean beauty brand photography — organic, tactile, ingredient-honest. No fake luxury excess.Soft diffused daylight, overcast outdoor quality. No harsh highlights on the glass.Close-to-medium shot, 2/3 overhead angle, product at left-of-center, lavender at right.No artificial golden glow, no oversaturated purple, no digitally added sparkle effects.
Key mechanic: “One drop suspended mid-fall from the dropper” — GPT Image 2 handles freeze-motion physics well when described specifically. This single detail increases perceived product quality significantly.
Prompt #15Watch Photography — Traditional Studio
Luxury mechanical watch, silver case with dark blue dial showing 10:10 time, resting face-up on a flat black velvet surface, slightly angled 15° toward camera.Traditional Swiss watch brand photography — Patek Philippe catalog aesthetic. Clinical, precise, no lifestyle context.Single overhead softbox, extremely controlled. Watch crystal shows a clean, single reflection of the light source. No stray reflections on the case sides.Macro close-up, top-down with 15° tilt, dial fills 75% of frame.No background elements, no text on dial, no watch band/strap, no lens distortion at edges.
Key mechanic: “10:10 time” is a genuine watch photography convention (it shows the logo and creates a visually balanced open-armed composition). Specifying it produces more credible output.
04
5 prompts · Spatial prompts reward extreme specificity about materials and light entry points.
Prompt #16Brutalist Interior — Lived In
Interior of a converted brutalist concrete apartment — exposed concrete walls, polished poured-concrete floors, a single large window (2m x 3m) facing north. Sparse furnishing: one worn leather sofa, a floor lamp, a stack of books. Signs of lived-in use.Architectural interior photography — Peter Zumthor’s material honesty, editorial magazine quality.Overcast north light entering through the single window, flat and diffused. Long shadows of window frame projected on floor. No artificial interior lights on.Standing-height perspective from the opposite corner, 24mm wide, showing all four walls in perspective.
Key mechanic: “Signs of lived-in use” counteracts the model’s default of generating impossibly pristine architectural renders that look like no human has ever been inside.
Prompt #17Japanese Minimalist Bathroom
A Japanese hinoki-wood soaking tub in a minimal bathroom — slatted wood floor, unglazed white tile walls, single built-in shelf with one ceramic vessel and one rolled towel. Floor-level window of frosted glass.Japanese wabi-sabi interior photography — tactile, natural materials, imperfect beauty. Similar to Muji’s design aesthetic.Diffused light from the frosted window, casting even, soft illumination across the room. No shadows visible. Warm neutral color temperature.Corner-to-corner perspective, full room visible, medium-wide.No chrome fixtures, no colorful accessories, no towel folded decoratively.
Key mechanic: Specifying “one ceramic vessel and one rolled towel” as the only shelf items enforces the minimalism — the model will add decorative clutter without explicit quantity limits.
Prompt #18Futurist Exterior — Night Scene
A single-family residential building, 2050s design language — cantilevered upper floor, integrated vertical garden on south facade, rooftop covered in thin-film solar panels. Set in a forested valley, at night.Architectural visualization, photorealistic rendering — Bjarke Ingels / BIG studio aesthetic applied to residential scale.Interior lights warm and amber, glowing through full-height windows. Exterior facade lit by cool blue moonlight. Stars visible above treeline.3/4 angle from 20 meters distance, eye-level perspective, building filling 60% of frame.
Key mechanic: The contrast of warm interior light vs. cool moonlight exterior is a compositional instruction that makes the building feel inhabited and alive, not like an empty concept render.
Prompt #19Abandoned Industrial Space
Inside a derelict 1970s factory hall — broken sawtooth skylights above, pools of sunlight falling in diagonal columns through the dust. Industrial steel frames, some collapsed sections on the far side. Weeds growing through the cracked concrete floor.Urban decay photography — Richard Mosse or Edward Burtynsky’s industrial ruins documentation, not horror/apocalypse framing.Harsh direct sunlight through the broken skylights at 11 AM angle, high contrast between lit and shadowed areas. Atmospheric dust particles visible in the light shafts.Wide establishing shot from ground level, looking toward the far end of the hall.
Key mechanic: “Not horror/apocalypse framing” explicitly distinguishes the aesthetic — the model defaults to dark, desaturated post-apocalyptic looks for any ruin prompt without this instruction.
Prompt #20Cozy Cabin Interior — Winter
Interior of a Nordic log cabin, evening — a wood stove burning in the corner, one wool blanket-covered armchair facing it, a low side table with a ceramic mug steaming. Snow visible through a small window, deep blue outside.Scandinavian hygge interior photography — real and lived-in, not an IKEA catalog.Dominant warm amber light from the wood stove, flickering quality. Small pool of lamp light from a floor lamp near the chair. No overhead light.Medium interior shot, armchair at center-left, stove visible at right, low angle looking slightly upward.No Christmas decorations, no staged throw pillows, no fake fire glow effect.
Key mechanic: “Not an IKEA catalog” is a surprisingly precise instruction — the model has enough commercial interior photography in its training data to understand this distinction and deliver genuinely imperfect, human spaces.
05
5 prompts · The model’s world knowledge creates shortcuts — use named aesthetic traditions, not vague genre labels.
Prompt #21Biopunk City — Organic Architecture
A sprawling coastal megacity in 2180 where all architecture is grown from engineered coral and mycelium — buildings are curved, asymmetrical, surface-textured like living organisms. Bridges made of woven root structures. The harbor glows faintly bioluminescent blue at dusk.Biopunk concept art — NOT standard cyberpunk/neon aesthetic. Organic, warm earth tones, no metal or glass dominance.Twilight, the sun below horizon — deep indigo sky above, warm amber artificial bioluminescent glow rising from below.Aerial wide establishing shot, city extending to horizon, harbor in middle distance.
Key mechanic: “NOT standard cyberpunk/neon aesthetic” is essential — the model’s default sci-fi city is always Blade Runner. Explicitly naming what it’s not prevents that output.
Prompt #22Deep Ocean Research Station
An underwater research habitat at 400 meters depth — modular pressurized cylinders connected by transparent corridors. A single researcher in a white jumpsuit stands in a corridor, watching a large manta ray drift past the glass.Hard science fiction concept art — plausible near-future, NASA/ESA technical credibility aesthetic. Inspired by Kim Stanley Robinson’s ocean civilization depictions.Near-total darkness outside. Interior warm fluorescent light. Deep water teal ambient light seeping in from above.Inside the corridor, looking toward the researcher, the habitat hub visible beyond.
Key mechanic: “Plausible near-future, NASA/ESA technical credibility” vetoes excessive speculative design elements and keeps materials and proportions anchored to real engineering aesthetics.
Prompt #23Pre-Columbian Space Age Alternate History
An alternate-history space launch facility built in the Aztec empire’s architectural tradition — massive stone launch gantries inlaid with jade and gold mosaic, a rocket painted with serpent iconography standing ready on the pad, crowds watching in ceremonial dress.Alternate history concept art — architectural accuracy of Mesoamerican construction with functional 20th-century aerospace integration. No steampunk aesthetic.Dawn launch light: horizon warm orange, pad floodlights harsh and industrial, torch smoke from the base catching the light.Wide shot from crowd level, rocket at center, gantry structures flanking.
Key mechanic: Blending two specific historical/technical vocabularies in alternate-history prompts produces genuinely original output because the model cannot default to a single trained aesthetic.
Prompt #24Ancient Library — Alien World
The interior of an alien library built into a cavern — shelves carved directly into the stone walls reaching 50 meters high, books in hexagonal format stored in crystalline containers, a lone archivist using a floating light to read, dwarfed by the scale.Epic fantasy illustration, Wayne Barlowe’s alien biology influence — genuinely alien aesthetics, not human-in-costume fantasy.Bioluminescent elements embedded in the stone ceiling providing a cool blue ambient light. The archivist’s floating light warm and golden, creating a small warm circle in the vast cold cave.Looking upward and inward, the archivist at bottom-center of frame, shelves extending to the top of the image.
Key mechanic: “Dwarfed by the scale” is the critical emotional instruction. Always include a human figure with a scale relationship described in fantasy/sci-fi world-building prompts.
Prompt #25Desert Monastery — Far Future
A monastery on Mars, 2340 — built into a mesa of red rock, solar panels on the flat roof, a garden of engineered plants in pressurized glass domes at the side, a monk in orange robes walking between the domes on a dust-covered walkway.Solarpunk science fiction concept art — optimistic, spiritual, human-scale technology. Kim Stanley Robinson’s Mars aesthetic.Martian afternoon sun — sunlight 43% intensity of Earth, long shadows, slightly pinkish sky, reddish light quality.Medium wide shot from below the mesa, looking up at the monastery, figure visible on the walkway.
Key mechanic: “Sunlight 43% intensity of Earth” — giving the model a real scientific parameter (Mars is 1.52 AU from the sun) produces more credible light quality than “dim Martian sunlight.”
06
5 prompts · GPT Image 2’s strongest competitive advantage over all other image models.
Text rendering is where GPT Image 2 has the most decisive edge over every competing model as of mid-2026. Midjourney still scrambles letters. Stable Diffusion needs additional ControlNet pipelines for reliable text. GPT Image 2 gets English-language text right roughly 85% of the time on the first generation, and if it’s wrong, you can correct it conversationally: “Fix the second line — it should read ‘Sign up free’, not ‘Sign ip free’.” This editing loop is simply not available in other tools.
Prompt #26SaaS App — Mobile Dashboard
iOS app interface for a personal finance tracker. Home screen showing: header “Good morning, Alex”, balance card showing “$12,847.33”, three recent transactions listed (Grocery Store -$47.20, Salary +$3,200.00, Coffee -$4.80), bottom navigation with 5 tabs (Home, Transactions, Budget, Goals, Settings).Modern iOS design system — SF Pro typography, clean white/light grey background, green accent for positive numbers, red for negative.iPhone 15 Pro frame, screen fills 80% of image, clean white background outside frame.No gradients, no dark mode, no game-like UI elements, text must be legible at this scale.
Key mechanic: Specifying exact dollar amounts (not “some numbers”) forces the model to render specific text, producing a more realistic mockup. The amounts should be plausible — round numbers look fake.
Prompt #27Magazine Cover — Editorial Typographic
Magazine front cover for a publication called “FAULT LINES” — an economics and geopolitics quarterly. Issue theme: Global Supply Chains. Cover date: Autumn 2026. Cover image: a fractured world map, the cracks glowing deep red-orange like molten rock.Serious editorial design, Economist / Foreign Affairs visual register — confident, typographically driven, not sensationalist. Cover headline: “THE GREAT FRAGMENTATION” in a heavy condensed serif, white, over the image. Secondary: “Why global trade won’t reassemble the way you think”. Price: £9.50. Barcode at bottom-right.No tabloid aesthetics, no neon, no over-designed borders.
Key mechanic: Include a price point and barcode instruction — these small details dramatically increase the realism of the cover composition and force the model to treat it as a real publication.
Illustrated infographic titled “Why Rainforests Matter”: five sections with icons and text. Section 1: “20% of Earth’s oxygen” (icon: lungs). Section 2: “50% of all species” (icon: leopard). Section 3: “25% of medicines” (icon: capsule). Section 4: “1.6B people depend on them” (icon: community). Section 5: “80% of diet foods originate here” (icon: fruit).Educational infographic design — clean, editorial, flat icon style, green and earth-tone palette. Vertical layout, white background, each section has a consistent icon-stat-description structure.No pie charts, no bar graphs, no stock illustration style icons, no overcrowded layout.
Key mechanic: Provide the specific text for every section — don’t let the model generate placeholder statistics. If you need the stats to be accurate, verify them separately (I checked these against IUCN and WWF data).
Prompt #29Minimal Poster — Typographic
A minimal typographic poster. Large text: “THE FIRST PRINCIPLE” in a geometric sans-serif (similar to Futura), all caps, spanning the full width of the poster. Below: a single fine horizontal rule. Below the rule: in small italic serif text: “State the problem as it is, not as you wish it were.” Black text, white background.Swiss International Style typography, clean graphic design tradition — 1960s rationalist poster aesthetic. Layout: text starts at 25% from top, wide margins left and right. Nothing else on the poster.No decorative elements, no image behind text, no color accent, no shadow.
Key mechanic: Specifying “25% from top” for text placement overrides the model’s centering bias, producing more considered typographic compositions.
Prompt #30Brand Identity — Logo & Color System
Brand identity presentation for a coffee roastery called “Meridian.” Primary logo: a stylized compass rose formed from coffee bean shapes, wordmark “MERIDIAN” in a compressed geometric serif below. Color palette shown: deep espresso brown (#2C1810), warm cream (#F5E6D0), copper accent (#C4773D). Show logo in three sizes, on white background, on dark background, and on a paper coffee bag mockup.Professional brand identity sheet, clean presentation design — Pentagram or Collins studio presentation quality.No drop shadows on the logo, no gradients in the logo mark, no busy background pattern.
Key mechanic: Including hex codes (even approximate ones) steers the model toward your intended palette rather than its default interpretation of “espresso brown.”
07
5 prompts · Style specificity matters more than subject specificity in this category.
Prompt #31Abstract Expressionism — Color Field
Large-format abstract painting — color field composition with three dominant zones: deep prussian blue occupying 60% from the top, separating into a mid-tone teal band, then a warm rust-orange base. Edges slightly blurred between zones. No figures, no geometric shapes, no text.Mark Rothko color field painting style — meditative, large-format, visible brushwork at the zone edges, matte paint quality. Oil on canvas texture. Painting shown at an angle, hanging on a white gallery wall, natural light falling across the surface.No decorative flourishes, no pattern within the color zones, no photorealistic elements.
Key mechanic: Giving percentage areas for color zones (“60% from top”) is one of the most reliable layout instructions for abstract composition. The model respects proportional space descriptions well.
Prompt #32Ink and Water — Japanese Sumi-e
Sumi-e ink painting on off-white washi paper: a single bamboo stalk with seven leaves, rendered in three brushstrokes visible at the base, transitioning to finer strokes at the tip. One leaf has a small ink bleed where the brush paused.Traditional Japanese sumi-e technique — emphasis on negative space, economy of mark, spontaneous execution quality. Zen aesthetic, not decorative illustration. Paper texture visible, slight ink absorption and feathering at the leaf edges. Subtle watermark from wet brush on the paper surface nearby.No color, no background wash, no gold or decorative border, no signature seal unless requested.
Key mechanic: “One leaf has a small ink bleed where the brush paused” introduces an imperfection that signals authenticity — perfect digital recreation of a handmade technique always reads as fake.
Prompt #33Glitch Art — Controlled Corruption
A portrait of a woman’s face, 70% of it rendered in sharp photorealistic detail, 30% experiencing digital data corruption — pixel sorting artifacts in horizontal bands through the lower face, RGB channel displacement in the forehead area, a repeating data block failure in the top-right corner. The eyes are intact and undistorted.Glitch art aesthetic — intentional, curated corruption. Rosa Menkman’s glitch theory aesthetic, not random noise.Original image was high-contrast studio portrait before corruption.No over-design, no chromatic aberration over the whole image, no static noise overlay.
Key mechanic: “The eyes are intact and undistorted” — preserving the most human-readable part of a face while corrupting everything else produces the psychological tension that makes glitch art compelling rather than merely broken-looking.
Prompt #34Impressionist Landscape — Late Monet
An overgrown lily pond at dusk — water surface reflecting a deep violet-pink sky, lily pads rendered as dark circular silhouettes, reeds at the edge blurring into reflected sky. Weeping willow branches dipping into the right-side water.Late Monet impressionism — the large-format lily pond series (Nymphéas), 1920s period, heavy impasto paint texture, forms dissolving into light rather than defined by line. Oil on canvas, visible paint strokes at 45° angles in the water reflections.No hard outlines, no photorealistic rendering of individual leaves, no dark shadows.
Key mechanic: Specifying “1920s period” of Monet’s work is important — his early and late work are stylistically distinct. The late period (post-cataracts) is more abstract and more interesting for this use case.
A geometric abstract composition: 47 identical right-triangles arranged in a grid pattern, each slightly rotated from the previous by 7.3°, creating a slow visual rotation effect across the field. Deep navy blue triangles on off-white. The pattern continues to the edges of the canvas with no border.Mathematical art / hard-edge abstraction — Sol LeWitt’s systematic geometry, printed as a large-format digital print.No gradient fill, no shadows on shapes, no color variation between triangles, no visible grid lines.
Key mechanic: “47 identical right-triangles” and “7.3° rotation” — using odd, non-round numbers signals to the model that this is a precise systematic instruction, not a vague visual suggestion. Round numbers get approximate results.
08
5 prompts · Specificity of geography and time of day transforms generic landscape output.
Prompt #36Arctic Storm — Aerial View
Aerial photograph of sea ice breaking up in the Arctic Ocean in late spring — fractured ice sheets in geometric patterns, deep dark blue water visible between the cracks, one massive iceberg 3km long near the center of frame casting a long shadow.Satellite/aerial nature photography — Edward Burtynsky’s industrial landscape scale and detachment. Not a crisis photo, a geological observation.Mid-morning sun at low Arctic angle, casting long blue shadows from raised ice edges. High-altitude aerial perspective.No people, no ships, no polar bears, no climate propaganda framing.
Key mechanic: “Not a crisis photo, a geological observation” directly instructs the model’s framing disposition. Without this, Arctic ice prompts consistently produce emotionally manipulative aesthetics.
Prompt #37Macro Photography — Insect Eye
Extreme macro close-up of a dragonfly’s compound eye — the hexagonal facet structure filling the entire frame, each individual ommatidia visible, tiny hairs between facets in sharp focus, out-of-focus wing veins blurred at the extreme edge of frame.Scientific macro photography — technical precision, National Geographic natural history style, no artistic filters.Ring flash lighting, flat and even to reveal the facet structure without creating specular hotspots on the curved surface.True macro, 1:1 scale, depth of field approximately 0.3mm.
Key mechanic: “Depth of field approximately 0.3mm” — specifying an impossible-to-see real parameter gives the model enough information to produce a more technically credible macro shallow-focus effect.
Prompt #38Deep Forest — Pre-Dawn
Old-growth temperate rainforest interior, Pacific Northwest, 4:47 AM — total darkness except for a faint horizontal band of deep blue pre-dawn light visible through the trees at the far end of a narrow path. Ground mist at knee height. Massive Douglas fir trunks in silhouette.Fine art landscape photography — Ansel Adams’ reverence for scale, but color — not black and white.Pre-dawn: zero direct light, ambient luminance from the eastern horizon only. No artificial light sources.Low angle looking down the path, trees framing the path left and right, far horizon visible at center.
Key mechanic: “4:47 AM” rather than “before dawn” — the specificity of the time implies the precise quality of pre-dawn light that “before dawn” leaves ambiguous.
Prompt #39Volcanic Eruption — Detail Study
Close-up study of active lava flow meeting the ocean — the molten rock meeting the water at the shoreline, steam explosion column rising 15 meters, orange-red lava surface with a thin black cooling crust fragmenting at the leading edge. Ocean spray meeting steam.Scientific nature photography — USGS/Hawaiian Volcano Observatory documentation aesthetic. Not dramatic Hollywood rendering.Dawn light ambient, overridden by the intense orange-red glow from the lava surface itself. Steam column backlit by the rising sun.Medium close-up, shore-level perspective, the lava edge filling the bottom half of frame.
Key mechanic: Giving the steam column a specific height (“15 meters”) gives the model a scale reference that prevents it from generating an apocalyptic explosion rather than a geological process.
Prompt #40Desert Night Sky — Milky Way
Astrophotography-style image: Namibian desert landscape at 2 AM, completely flat terrain with scattered scrub brush, the Milky Way galactic core filling the sky from upper-right to lower-left, extremely dense star field. A single tent with a warm amber glow from inside visible on the left third of the frame.Astrophotography — specific to the southern hemisphere Milky Way view. The galactic center is visible (which it is from Namibia in summer). Technical, not stylized.No moon. Starlight only for ambient light on the ground. Tent lamp the only warm light.24mm wide, 30-second equivalent exposure quality — slightly star-trailed (not round point stars).
Key mechanic: “Slightly star-trailed (not round point stars)” mirrors the real physics of a 30-second exposure. Specify which hemisphere’s Milky Way you want — the southern hemisphere view is substantially different and more dramatic.
09
Fashion Design, Characters & Concept Art
5 prompts · Character consistency instructions, fashion illustration techniques.
Prompt #41Fashion Illustration — Technical Flat
Technical fashion flat illustration of a structured oversized blazer: front view and back view side by side. The blazer has a double-breasted closure with six buttons, wide lapels, a single back vent, and interior pocket detailing visible in the back view. Material is a herringbone wool in charcoal grey.Fashion technical illustration — clean CAD-quality line work with flat color fill. No figure, no background. White illustration background.Herringbone texture visible in the flat color fill. Button details shown with precision.No fashion figure, no shadow, no perspective distortion, no artistic flourishes.
Key mechanic: Fashion flats on a garment (no figure) are actually easier for the model to produce accurately than on-body shots — consider this for early-stage design visualization.
Prompt #42Character Concept — Warrior Archivist
Character concept art: a woman in her late 40s who is simultaneously a warrior and archivist. She wears plate armor that has been modified to include leather satchels and scroll tubes. She holds both a short sword and an open manuscript. Her expression is tired but resolute.Character concept art for a story-driven RPG — detailed, narrative, Christopher Moeller’s “Iron Empires” graphic novel visual weight. Not anime, not standard fantasy.Turnaround: front view only. Neutral grey background. Clean studio-quality concept art lighting.No oversexualized armor design, no cape, no dramatic pose — she’s standing and reading.
Key mechanic: “Tired but resolute” gives the model an emotional brief that produces a face with personality. “Determined warrior” produces a generic heroic expression. The internal contradiction (tired + resolute) is specific enough to execute.
A model wearing a sculptural avant-garde coat made entirely of overlapping rectangular layers of transparent black organza, creating a 3D architectural silhouette. The coat extends 60cm beyond the shoulders on each side. The model’s face is obscured by the collar height.Avant-garde fashion editorial — Comme des Garçons / Rei Kawakubo’s anti-fashion tradition. No conventional beauty framing.Stark stage lighting from directly above, creating a dramatic pool of light on the coat surface. Deep shadow at ground level.Full-length editorial shot, neutral grey runway floor visible.
Key mechanic: The “face obscured by collar height” removes the model’s tendency to generate conventional beauty-editorial faces that would undermine the anti-fashion concept. Obscuring the face forces compositional focus onto the garment.
Prompt #44Character Sheet — Creature Design
Creature design concept sheet for an aquatic mammal species that evolved to use echolocation: body length approximately 1.8 meters, dolphin-like torpedo body, four vestigial limb stumps (not fins), complex facial structure with external sound-producing organs, bioluminescent stripe patterns along the flanks. Show: front profile, side profile, 3/4 view. White background, labeled with annotations.Scientific illustration meets creature design — natural history museum plate quality applied to original species. No cartoonish elements.No fantasy elements, no wings, no unnecessary aggression features. Plausible evolutionary biology.
Key mechanic: “Plausible evolutionary biology” is a powerful framing for creature design — it activates the model’s scientific knowledge rather than its fantasy genre conventions, producing more genuinely original results.
Prompt #45Portrait + Style Transfer — Ghibli
Transform this image into a Studio Ghibli animation frame: preserve the subject’s core facial features (eye shape, face structure, expression), but render in the hand-drawn 2D animation style of Hayao Miyazaki — soft watercolor backgrounds, gentle cel shading, warm natural lighting, characteristic Ghibli color palette (soft teals, warm yellows, earthy greens).Studio Ghibli — specifically “My Neighbor Totoro” or “Spirited Away” aesthetic, not the more dramatic “Princess Mononoke” style. Place the subject in a cozy Japanese countryside room interior background, visible through a window behind.No western cartoon style, no anime exaggeration, no action-style dynamic lines, no unnatural eye size.
Key mechanic: Specifying which Ghibli film’s aesthetic (“Totoro/Spirited Away” vs. “Mononoke”) produces meaningfully different outputs. The studio has a wide stylistic range that the model accurately distinguishes.
10
Iterative Workflows & Advanced Techniques
5 prompts · The conversational editing capabilities most users haven’t discovered yet.
The five prompts in this section aren’t standalone generation prompts — they’re workflow prompts. GPT Image 2’s surgical region editing via conversation is the capability that separates it from every other consumer image tool, and it’s the capability most people discover by accident rather than by design. These are the techniques I use most often in production.
Workflow note: All five prompts below assume a two-step process: generate an initial image, then issue iterative corrections within the same chat. The key is generating on a lower quality setting first to check composition, then issuing a “final render” instruction. This saves generation limits on ChatGPT Plus (which as of mid-2026 gives approximately 40 generations per day on the default setting).
Prompt #46Initial Reference + Region Edit
STEP 1 (Initial): A woman in a red wool coat standing on a cobblestone street in Edinburgh, Old Town, foggy morning. Standard composition, just establish the scene.
STEP 2 (Edit): “Keep everything identical. Change only the coat from red to a deep forest green. Re-render at full quality.”
Key mechanic: The phrase “Keep everything identical” plus “Change only [specific element]” is the conversational edit pattern that produces the most reliable region-specific changes. State what stays before stating what changes.
Prompt #47Character Lock — Multi-Scene
SCENE 1 (Character Reference): Character reference portrait: a woman in her early 30s, East Asian features, short blunt-cut black hair with blunt fringe, round tortoiseshell glasses, wearing a cream turtleneck. Neutral expression, front-facing. White studio background. This is the character reference — do not generate a narrative scene yet.
SCENE 2: “Using the exact same character from the reference portrait — identical face, same hair, same glasses — show her sitting in a rain-soaked café window seat at night, looking out.”
Key mechanic: Generating an explicit “character reference portrait” as step 1 (and labeling it as such) is significantly more reliable for maintaining identity across scenes than describing the character repeatedly from text alone.
Prompt #48Composition Draft → Full Render
STEP 1 (Draft, lower quality): Draft composition only: a product photography scene with a glass whiskey decanter and two glasses on a dark slate surface, warm side lighting. Don’t optimize for quality — I just need to check the composition and lighting angle.
STEP 2 (if composition works): “The composition is exactly right. Now regenerate this scene at maximum fidelity with full detail on the glass reflections, the liquid color, and the slate texture. Same framing, no changes.”
Key mechanic: Explicitly asking for a “draft” on the first pass primes the model to generate faster and less resource-intensively. “Same framing, no changes” locks the composition before the quality pass.
Prompt #49Text Correction Loop
INITIAL: A vintage-style poster with the headline “THE LONG GAME” in a distressed wood-type font, subheading “A Story About Patience” in a smaller clean sans-serif below. Warm sepia palette, textured paper background.
CORRECTION (if text has errors): “The headline reads ‘[whatever it says]’ but it should read ‘THE LONG GAME’ exactly. The subheading reads ‘[whatever it says]’ but should read ‘A Story About Patience’ exactly. Keep everything else — only correct the text.”
Key mechanic: When correcting text errors, always quote exactly what the model wrote AND what it should be. The model uses the contrast between the two to make the minimal correction rather than regenerating the entire composition.
Prompt #50Background Swap — Subject Lock
INITIAL: A product shot of a matte black leather wallet on a plain white background. Simple lighting.
BACKGROUND SWAP: “Using the wallet from the previous image exactly — identical product, same angle, same surface texture, same lighting on the wallet itself — place it now on a textured cognac leather desk surface. Keep the wallet dimensions, position, and look completely unchanged. Only change what the wallet is resting on.”
Key mechanic: “Keep the wallet dimensions, position, and look completely unchanged. Only change what the wallet is resting on.” — restating the invariant three times is not redundant. It significantly reduces compositional drift in background swaps, which is the most common failure mode in this operation.
The 7 Principles Behind Every Prompt Above
If you only remember seven things from this page, make it these. I’ve listed them in order of impact, not importance — they’re all important.
01 Subject first, always Architecture
GPT Image 2’s autoregressive generation reads your prompt sequentially. Put your subject and its primary action first. Lighting and style specified before the subject causes bleed — the model’s style interpretation starts before it knows what it’s rendering.
02 Name the lighting source and its direction Lighting
Not “dramatic lighting.” Not “moody lighting.” Name the source (single overhead softbox, north window, late afternoon sun at 15°) and its direction relative to the subject. This one change produces the highest per-prompt quality improvement I’ve measured. Lighting direction determines shadow placement, which determines dimensionality.
03 Camera lens + angle = composition control Composition
Using photography terms (85mm, 24mm, f/1.4, overhead, eye-level, low angle) produces more reliable compositional control than spatial descriptions alone. The model has enough photographic training data to translate these into consistent visual conventions.
The model has default outputs for every category. Portraits default to beauty editorial. Landscapes default to golden hour saturation. Sci-fi defaults to Blade Runner. Identify the default for your category and explicitly prohibit it. Negative instructions don’t reduce creative range — they redirect the model away from its statistical mean.
05 Specific numbers over approximate adjectives Precision
“$847” not “expensive.” “4:47 AM” not “before dawn.” “30th step” not “high up.” “7.3°” not “slightly rotated.” The model treats specific numbers as facts to execute rather than impressions to interpret. This is the fastest way to increase first-generation success rates.
06 State both what stays and what changes in edits Editing
In conversational edits, always state the invariants before the change: “Keep [everything that should stay] unchanged. Change only [specific element].” The model defaults to re-interpreting the entire scene on every turn — this instruction locks what shouldn’t move.
07 Introduce one credible imperfection Authenticity
This is the counterintuitive one: images generated without any instructed imperfection tend to look AI-generated even when technically well-executed. One specific, credible imperfection — an ink bleed, a worn corner, an asymmetrical shadow — carries authentic information that perfect rendering lacks.
Questions That Come Up
What model does ChatGPT use for image generation in 2026?
As of mid-2026, ChatGPT uses GPT Image 2 (gpt-image-2) as its default generation model. DALL-E 3 was deprecated in May 2026. GPT Image 2 offers substantially better text rendering, stronger instruction-following, and handles complex multi-element compositions better than its predecessors. The OpenAI API still exposes gpt-image-1 and gpt-image-1.5 for teams who need the specific characteristics of those models.
Can I use ChatGPT-generated images commercially?
Per OpenAI’s terms as of early 2026, images generated by ChatGPT Plus, Pro, and Team subscribers can be used commercially. Policies change — verify current terms at openai.com before commercial deployment, especially for advertising, product packaging, or any use where IP attribution matters. Free tier users should specifically check current free-tier commercial rights, as these have historically been more restricted.
How do I get consistent characters across multiple images?
Within a single chat session, generate a character reference portrait first (labeled explicitly as “character reference, do not generate a scene”), then in every follow-up prompt write “Using the exact same character from the reference — identical face, hair, and distinguishing features — [new scene description].” For cross-session consistency, upload the reference image at the start of each new session and use the same instruction pattern. The model handles identity retention better than any previous consumer tool — not perfectly, but usably.
What are the biggest prompt mistakes beginners make?
Three patterns produce the most failures: (1) Stacking abstract adjectives without concrete visual references — “beautiful,” “stunning,” “amazing” add nothing the model can execute. (2) Omitting lighting direction entirely, which causes the model to default to flat, overlit stock-photo output. (3) No negative instructions — not telling the model what to avoid means it will include its most common training artifacts for that category. These three fixes alone will meaningfully improve first-generation success rates.
Do these prompts work with Midjourney or Stable Diffusion?
The structural principles transfer — lighting direction, subject-first ordering, specific numbers, negative instructions — but the syntax is different. Midjourney uses parameter flags (–ar, –style, –no) rather than natural language negatives. Stable Diffusion (especially with SDXL or Flux) uses weighted prompts and separate negative prompt fields. The core architecture of these prompts (subject + style + lighting + camera + negatives) applies universally; the formatting needs adjustment per platform.
One thing worth saying plainly: the prompts above are blueprints, not scripts. The best results come from treating the first generation as a draft and the conversation as a workspace. The “surgical edit” capability — select region, describe change, only that region updates — is the feature that changes professional workflow the most, and it’s still underused because most people treat each image as a single transaction rather than a progressive refinement.
The limit isn’t the model. The limit is how specifically you can describe what you actually want. That’s a skill, and it gets cheaper to practice every week.