The Ghost in the Machine: Earlier Images Are Haunting Later Ones in GPT Image 2
A few weeks ago, we published a comprehensive roundup of the artifacting issues that have plagued GPT Image 2 since its release — the strange noise clusters that appear in complex textures, the Voronoi-like patterns that creep into fine details, and the workarounds users have cobbled together in the meantime. We thought that covered most of the known issues.
It turns out, we were only telling half the story.
There's a second, distinct class of artifacting that deserves its own investigation — one that has nothing to do with texture complexity or prompting style, and everything to do with how GPT Image 2 handles context within a conversation.
The Discovery: Images Are Bleeding Into Each Other
A Reddit user going by bendyorange posted a finding that stopped a lot of people in their tracks. While generating a sequence of three very different images in the same ChatGPT conversation, they noticed something unsettling: elements from earlier images were showing up — as faint but unmistakable artifacts — in the images that followed.
Let's walk through the experiment, because the evidence is striking.
Image 1: The Concert Hall Setup
The first image was a realistic architectural photo of a concert hall mid-setup — stage crew on the floor, lighting trusses suspended overhead, tiered seating in the background. It came out clean, with no visible artifacts. This is important: the first image in a session almost always generates without issues.

Image 2: The Frog Pilot (and the Ghost of the Concert Hall)
The second image was a sci-fi illustration of a frog in a spacesuit piloting a sleek craft over a swamp — stylistically miles apart from the concert hall.

Yet despite being generated from an entirely different prompt, it carried traces of Image 1:
- Hard lines from the concert hall's lighting truss appeared embedded in the craft's body, aligning spatially with where the truss had appeared in the previous image.

- In the fin area of the spacecraft, a ghosted figure — the stagehand from Image 1 — was visible in the fuselage.

- In the fuselage, a second stagehand from Image 1 also appeared — commenter Echo4Mike noted: "You can clearly see the foremost stagehand right there in the fuselage."

Image 3: The Ballroom Scene (and the Frog's Outline)
The third image was a lavish Victorian-era ballroom rendered in a warm impressionist style.

By this point, the ghosting was compounding across two prior images:
-
The outline of Image 2's spacecraft bled into the ballroom's figures and architecture — the drawn silhouette of the frog's craft faintly but clearly present in the scene.

-
The text "I long John Silwa" — which had appeared as a label in Image 2 — ghosted into Image 3 at the same spatial position, confirmed by a commenter who spotted the identical text carried over.

- The number "22" from the spacecraft's fuselage markings in Image 2 also reappeared in Image 3, noted by another commenter in the thread.

- A stagehand from Image 1 — the same figure that had ghosted into Image 2's fuselage — reappeared again in Image 3, this time overlaid onto a dancing couple at the corresponding position in the frame.

The further into the chat, the more pronounced the ghosting. This is not random compression noise, and it's not a quirk of any particular art style. It is a systematic bleed of visual tokens from one generation to the next.
Independent Confirmation: Reproduced in the Wild
After the Reddit post gained traction, a user in the OpenAI Developer Community set out to reproduce the finding with their own prompts. The result was unambiguous.

They generated two images in the same chat session — Image 1 featuring an airport control tower, Image 2 a KLM aircraft. The control tower from Image 1 was plainly visible, ghosted into the aircraft image. The user's conclusion:
"So it looks like the 2nd image uses the 1st image as a start and puts it as a layer on top. This happens on other images / chats as well: it's a clear bug and not desired. For me, this is a pure mess."
The bug is reproducible, cross-prompt, and cross-style. It is not isolated to a single content type or one user's workflow.
Why Is This Happening? What the Community Thinks
The Reddit thread attracted sharp technical commentary, and a few distinct schools of thought emerged about the root cause.
Theory A: Autoregressive context bleed. User mobcat_40 offered the most technically detailed take. GPT Image 2 appears to be an autoregressive model rather than a pure diffusion model. Unlike diffusion, which starts from fresh random noise on each call, an AR model re-reads every earlier image token in the chat context before generating the next one. The previous image's spatial structure is literally part of the input — and the model can't cleanly separate "context I should reference" from "context I should ignore." This is also why fresh chats fix it and long sessions make it worse. Crucially, the model is trained to support in-chat image editing, which reinforces this cross-turn attention mechanism.
Theory B: Shared noise seed / diffusion blending. User Effective-Cat-1433 proposed that images within the same session might share an initial random noise seed. If two diffusion passes start from the same noise basis, they'll share some structural DNA — even if the prompts are entirely different. A follow-up from the same user suggested this could alternatively stem from the multi-turn architecture itself, where the prior image is in context and its structure gets reused as a foundation.
Theory C: An unintended side effect of editing support. Several users raised the possibility that this is not a pure bug but a side effect of a feature. GPT Image 2 is explicitly designed for in-chat image editing, which requires the model to "remember" what a prior image looked like. User jjp3 noted that during a generation, the model's visible reasoning actually surfaced the thought: "it looks like there's an artifact from the previous image here." The model is aware it's happening — but the mechanism that enables editing is the same one causing unintentional ghosting.
Theory D: Sampling instability under competing signals. Also from mobcat_40: when two conditioning signals are pulling toward different images — your new prompt versus the prior image still in context — the model samples unstably in regions where they conflict. The result "falls out as high-frequency noise." The spotting is not added to hide the bleed; it's what happens when two strong visual attractors fight over the same pixels.
The most credible synthesis: GPT Image 2's architecture treats all prior images in a chat session as input context for every subsequent generation. This is the same mechanism that makes multi-turn editing work, and it's also what causes unintended visual bleed when you simply want a fresh, unrelated image in the same conversation.
Has OpenAI Fixed This? (As of May 7, 2026: No)
As of today, OpenAI has not officially announced a fix for either the general artifacting issues or the cross-image bleed behavior described in this post. The most recent public statement from the team comes from OpenAI core developer Boyuan Chen, who replied on X on May 4, 2026 to a user asking whether the "noise issue in certain backgrounds" was being addressed. His reply: "Working on it, stay tuned." He also confirmed separately that he would post a public announcement when the fix goes live.

Until that announcement lands, the workarounds below are your best options.
What You Can Do Right Now
There's no complete fix available yet. OpenAI has acknowledged that this is a model-level issue — not a fluke or an edge case — and a core developer has publicly confirmed it's being actively worked on. In the meantime, if your project involves any of the following, our recommendation is to hold off on GPT Image 2 for now and consider alternatives:
- Scenes with heavy natural environments — forests, bodies of water, atmospheric weather, dense vegetation
- Highly specific artistic styles — especially those requiring fine, repeating, or organic textures (watercolor grain, impressionist brushwork, anime cel shading, rust and corrosion, scales and chainmail)
- Subjects prone to repeating micro-patterns — fur, feathers, cracked earth, gravel, fabric weaves
- High-fidelity photorealism — any image where you need clean, noise-free surfaces throughout
That said, GPT Image 2 is genuinely exceptional in a number of areas where the current bugs are unlikely to surface, and if your work falls into the categories below, you can use it with confidence:
- Text-in-image rendering — This is GPT Image 2's signature strength and the most dramatic improvement over prior models. Posters, infographics, product labels, UI mockups, signage, and editorial spreads with multi-line copy all come out legible, correctly spelled, and properly placed. OpenAI has also noted stronger understanding of non-Latin text rendering in languages like Japanese, Korean, Hindi, and Bengali.
- Product and commercial photography — Drop in a product image, describe a scene, and get a clean product photo with proper background, lighting, and staging. Bloomberg's coverage of the launch specifically highlighted GPT Image 2's ability to produce accurate, complex visuals for professional workflows. Packaging mockups, e-commerce product shots, and ad creatives with real headlines are all well within its reliable range.
- Complex multi-element compositions — The model's prompt adherence is a headline improvement: when you ask for many specific elements in a scene, GPT Image 2 tends to include all of them — a consistent failure mode of earlier models that GPT Image 2 handles well.
- Structured visuals and diagrams — Complex structured visuals, including infographics, diagrams, and multi-panel compositions are explicitly one of the model's production-grade strengths. The flat, bounded geometry of diagrams also naturally avoids the fine-structure patterns that trigger artifacting.
- UI mockups and app design assets — Clean interfaces, wireframes, and design system components generate with minimal noise, as the artificial flat surfaces inherently sidestep the conditions that produce artifacts.
- Conversational editing within a single image — One of GPT Image 2's practical strengths is conversational editing: generate an image, then follow up with targeted changes like "make the sky more dramatic" or "shift the subject to the left third of the frame." The model applies changes rather than starting from scratch — just be aware this is precisely the workflow where cross-image bleed can compound if you're also switching subjects between edits.
If you're not sure whether your use case falls into the "safe" zone, a simple rule of thumb: the more artificial, flat, and geometrically regular your scene, the cleaner the output will be. The more naturalistic and fine-grained, the higher the risk.
If you still need GPT Image 2 for a project that falls into the risk categories — whether because of deadline, workflow, or capability reasons — the following tips may help reduce the damage.
1. Never generate multiple unrelated images in the same chat
This is the single most important mitigation. The cross-image bleed documented in this article is reliably triggered by generating more than one image in the same ChatGPT conversation — especially when the images are stylistically or thematically different. Always start a fresh chat for each image you want to generate from scratch. The first image in a session is almost always clean.
2. Understand how GPT Image 2 constructs a scene
A sharp observation posted in the OpenAI Developer Community offers useful insight into the model's generation strategy. Analyzing a detailed coffee machine infographic that OpenAI itself generated with GPT Image 2, user EricGT noted that the black background contained layered, precisely bounded rectangular regions — each containing a distinct component like a pipe, valve, or flow indicator. This suggests the model uses a depth-first compositional approach: it defines structured layout regions first, then populates each with relevant visual content.
The practical upshot: write prompts that define clear regions and distinct elements explicitly. Think of your prompt as a layout specification — describe the scene in terms of discrete, bounded parts rather than holistic atmospheric descriptions. This plays to how the model actually builds images and may reduce the instability that leads to artifacts.
3. Avoid prompt elements that reliably trigger artifacting
Community testing has produced a rough but useful taxonomy of what triggers the cluster and noise artifacting (distinct from the cross-image bleed, but related in origin):
Elements that trigger artifacting:
- Microparticles, particle effects, glitter, sparkling
- Natural rough surfaces and organic textures
- Fog, smoke, clouds, and motion
- Scenes with many small plants or dense foliage
- Fantastical, non-real subjects
Elements that suppress artifacting:
- Flat, uniform artificial surfaces
- Artificial uniform backgrounds
- Few details, little-to-no microstructures
- No flying particles, no rough surfaces
- Real existing subjects (photo-realism)
The core pattern: the more fine-grained structure the generator has to produce, the more likely it dissolves into artifact clusters. Keep prompts artificial, flat, and sparse to stay in the clear zone.
4. If you must generate complex subjects: engineer your constraints, not just your descriptions
Sometimes you need smoke, particles, or organic textures — and simplifying the prompt isn't an option. In that case, prompt engineering can help, though it involves real tradeoffs.
User Chain_L documented an experiment comparing two prompts for the same concept: a humanoid figure shattering into obsidian shards.
Prompt v1:
Scene: Abstract dark void, dramatic backlit environment, floating embers and glass shards, intense energetic atmosphere.
Subject: A humanoid form made of molten and shattering obsidian, caught in the exact moment of fragmentation. Sharp, glassy shards suspended in mid-air, catching refractive shafts of cold white light. Deep crimson and orange magma-like glow emanates from the cracks, contrasting violently with the black glass. The pose is dynamic, arms outstretched as if releasing or absorbing energy.
Important details: High-speed fantasy photography, frozen motion aesthetic, refractive glass textures with caustic light patterns, dramatic rim lighting, particle physics detail (shards, embers, vapor), intense chiaroscuro, surreal energy visualization, masterful composition, digital painting meets photorealism style.
Constraints: No messy or chaotic composition, keep the central form recognizable. Balance the internal glow with the black glass texture. Preserve the sharpness of the fragments and the fluidity of the molten sections. No decorative borders or text.
AR 16:9

Result: Heavy artifacting throughout — dense Voronoi cell patterns, webbing, and chaotic micro-textures covering the entire image, with the surface losing any sense of glassy material.
Prompt v2:
Scene: Abstract dark void, dramatic backlit environment, floating embers and suspended glass fragments, intense energetic atmosphere.
Subject: A humanoid form made of smooth, polished black obsidian, captured in the exact moment of fragmentation. The body is breaking apart into large, distinct, sharp geometric shards of glass, NOT small dust. Internal glowing cracks of crimson light emit energy, but the outer surface of the shards remains perfectly smooth and mirror-like.
Important details: Macro glass photography, frozen motion aesthetic, individual distinct shards, sharp geometric edges, refractive light patterns, dramatic rim lighting, particle physics detail (large fragments only), intense chiaroscuro, surreal energy visualization, masterful composition.
Constraints: STRICTLY NO cellular texture, NO webbing, NO neural network patterns, NO repeating Voronoi patterns. The obsidian must look like solid, smooth glass breaking into large pieces, not organic cells or mud cracks. Keep the central form recognizable.
AR 16:9

Result: Dramatically cleaner — large, distinct geometric shards with mirror-smooth obsidian surfaces, dramatic rim lighting, and significantly reduced artifacts.
The community analysis of why this works: Prompt v1 feeds the model fine-grained microstructures — exactly what causes the denoising process to produce cluster artifacts. Prompt v2 suppresses that complexity by demanding flat, large surfaces. The tradeoff is real: you lose dynamic visual complexity in exchange for technical cleanliness.
Important limitation: This approach has a hard ceiling. For certain inherently complex subjects — like a creature made of smoke, or a tornado swirling with thousands of small stones — no amount of constraint language makes the problem fully disappear. As one commenter put it: "The problem is not the prompting, it is the subject." For nature scenes in particular, this optimization doesn't work at all.
5. For batch generation: use the API, not ChatGPT
If you're generating images at scale, the GPT Image 2 API is meaningfully better than the ChatGPT interface for avoiding cross-image bleed. In ChatGPT, each generation in the same chat compounds the context — the model sees every prior image on each turn. With the API, each call is an independent request. You control exactly what context is passed in, and a fresh call with no prior image history won't carry over visual artifacts from previous generations.
Community testing has confirmed that the general artifacting issues (noise clusters, Voronoi patterns) can still appear in API-generated images when the prompt triggers them — the API doesn't eliminate that class of problem. What it does eliminate is the cross-image ghosting behavior specifically documented in this article.
The Takeaway
GPT Image 2 is a remarkable model. Its ability to follow complex compositional prompts, maintain fine text rendering, and produce stylistically rich images puts it in a class of its own in several respects. But it launched with a set of systemic issues that OpenAI is still working through — and the cross-image bleed behavior documented here is arguably the most disorienting of them, because it's invisible until you know what to look for.
The good news: OpenAI has acknowledged the problem, a core developer is actively working on it, and a public fix announcement is promised. The less-good news: it's not here yet.
In the meantime — start a new chat for every image. If you're building anything production-grade, use the API. And if you're generating complex organic or particle-heavy scenes, spend time on your prompt's constraints, not just its descriptions.
We'll update this post — when the official fix lands. To get notified directly from the source, keep an eye on @BoyuanChen0 on X, who has committed to posting the moment it's resolved.
