We evaluated GPT Image 2 and GPT Image 2.5 for human portraits and UGC content across five settings. The tests cover everyday portraits, outfit and expression edits, product placement, mirror selfies and two-person remixes, using identical inputs and prompts for each setting.
GPT Image 2 Low and Medium; GPT Image 2.5 Flare Low and Medium; and GPT Image 2.5 Sunburst Medium. Each row uses identical ordered input images, prompt text and output dimensions. We made one request per setting, with no prompt rewrites or best-of selections. “Low” and “Medium” are provider settings, not equal-price tiers.
The two-person remixes use a scene plus separate Lily and Mila references. Both selected Explore templates were rated Excellent and routed to Nano when selected; this test sends them to the GPT endpoints. The skincare bottle is fictional, and its label is part of the text-rendering test.
Complex edits and two-person remixes
Comparison 1
Layered outfit replacement
Layering, knit texture, belt loops and preserved pose.
Scroll sideways on smaller screens. Open any image to inspect the full-resolution output.
Read the exact prompt
Replace the adult woman's outfit with a fitted cream cable-knit sweater tucked into high-waisted dark jeans, a brown leather belt threaded through the visible belt loops, and an open olive jacket. Make the individual cable-knit stitches, jacket seams, belt buckle and fabric layers realistic. Preserve her exact facial identity, expression, hairstyle, body proportions, pose, hands, yacht railing, ocean wake, lighting, camera angle and framing. Hair and hands must overlap the new clothing naturally. Do not change anything except the outfit.
Comparison 2
Gym mirror consistency
Outfit and phone changes with consistent reflections.
Scroll sideways on smaller screens. Open any image to inspect the full-resolution output.
Read the exact prompt
In this adult woman's gym mirror selfie, change her outfit to navy athletic wear, position her free hand on her hip, and change the phone to a red phone. Keep every visible reflection physically consistent with her pose, outfit and red phone, including correct handedness and reflected geometry. Preserve her exact facial identity, expression, physique, hairstyle, headphones, gym equipment, lighting, background and framing. Do not add another person or an extra phone.
Comparison 3
REMIX CAMERA product placement
Brand spelling, benefit lines, fine print and hand contact.
Scroll sideways on smaller screens. Open any image to inspect the full-resolution output.
Read the exact prompt
Have Lily hold a luxury skincare bottle beside her face: frosted pale-pink glass, a brushed-silver pump, and a label facing the camera. Her fingers wrap naturally around its sides without covering the text. Match the scene's lighting, perspective, contact shadows and reflections.
Print this exact label:
REMIX CAMERA
DAILY SERUM
Same you. Endless possibilities.
FORMULATED WITH
Your face, preserved
Any look, any setting
New photos in seconds
DIRECTIONS
Upload your photos.
Choose a look. Make it yours.
30 mL / 1 fl oz
Use charcoal lettering. Make the brand largest, the product name slightly smaller, the benefit lines medium-sized, and the directions fine but readable. Follow the bottle's curved surface with realistic printed typography. Include every line exactly once, with no additional text.
Preserve Lily's facial identity, expression, hairstyle, outfit, background and framing. Change only the hand and arm position needed to hold the bottle.
Comparison 4
Mila + Lily festival cheek kiss
Two distinct identities, cheek contact and overlapping hands.
See the original scene and identity references
Image 1: source sceneImage 2: Lily identityImage 3: Mila identity
GPT Image 2 Low
Blocked by provider moderation
GPT Image 2.5 Flare LowGPT Image 2 Medium
Blocked by provider moderation
GPT Image 2.5 Flare MediumGPT Image 2.5 Sunburst Medium
Scroll sideways on smaller screens. Open any image to inspect the full-resolution output.
Read the exact prompt
Image 1 is the source scene to remix. Image 2 is the adult Lily identity reference. Image 3 is the adult Mila identity reference.
Replace the foreground woman on the left giving the cheek kiss with Lily from Image 2. Replace the foreground woman on the right receiving the cheek kiss with Mila from Image 3. Preserve the cheek contact, cupping hand, facial expressions, gaze, white tops, gold jewelry and sunglasses. Keep background festival attendees separate from both reference identities.
Use each identity only for its assigned person. Keep the two foreground identities distinct; do not swap faces, blend identities or create twins. Preserve the source scene composition, clothing, lighting, camera angle, interaction and visible person count. Adapt framing to a 2:3 portrait without removing either foreground subject. Photorealistic adults with natural anatomy, coherent hands and realistic skin. The reference photos provide identity only; do not copy their clothing or backgrounds.
Scroll sideways on smaller screens. Open any image to inspect the full-resolution output.
Read the exact prompt
Image 1 is the source scene to remix. Image 2 is the adult Lily identity reference. Image 3 is the adult Mila identity reference.
Replace the woman on the left wearing yellow and holding the phone with Lily from Image 2. Replace the woman on the right wearing red with Mila from Image 3. Preserve their distinct poses, expressions, yellow versus red tops, white versus grey shorts, phone, countertop objects, poster and bathroom reflections.
Use each identity only for its assigned person. Keep the two foreground identities distinct; do not swap faces, blend identities or create twins. Preserve the source scene composition, clothing, lighting, camera angle, interaction and visible person count. Adapt framing to a 2:3 portrait without removing either foreground subject. Photorealistic adults with natural anatomy, coherent hands and realistic skin. The reference photos provide identity only; do not copy their clothing or backgrounds.
Scroll sideways on smaller screens. Open any image to inspect the full-resolution output.
Read the exact prompt
Make the adult woman visibly more athletic and muscular: add clear, natural definition to her deltoids, biceps, triceps and abdominal muscles, with moderately broader shoulders and stronger arms. Keep realistic proportions. Preserve her exact face, expression, hairstyle, grey workout clothing, pose, phone, headphones, mirror, lighting, background and framing.
Comparison 7
Expression: genuine laugh
An open-mouth laugh while preserving hands, rings and hair.
Scroll sideways on smaller screens. Open any image to inspect the full-resolution output.
Read the exact prompt
Change only the adult woman's facial expression from serious to a genuine open-mouth laugh with visible upper teeth, raised cheeks and naturally squinted eyes. Preserve her facial identity, head position, hairstyle including the strands across her face, both hands, rings, black sweater, lighting, background and exact framing.
Scroll sideways on smaller screens. Open any image to inspect the full-resolution output.
Read the exact prompt
Change the adult woman's pose so her torso faces the camera and both arms are raised, with her hands resting behind her head and elbows spread outward. Preserve her facial identity, neutral expression, beige dress, position on the yacht, windblown hairstyle, boat railing, ocean wake, lighting and camera angle. Keep the original crop and do not add people.
Comparison 9
Two people: edit right only
Surprise and a raised hand on one person; the other stays unchanged.
Scroll sideways on smaller screens. Open any image to inspect the full-resolution output.
Read the exact prompt
Change only the dark-haired adult woman on the right: make her look directly at the camera with raised eyebrows, widened eyes and a rounded open mouth expressing surprise; raise her left hand beside her left shoulder with an open palm facing the camera and five naturally separated fingers. Keep the light-haired woman on the left completely unchanged, including her face, laughing expression, gaze, hair, clothing, pose and hands. Preserve both identities, exactly two people, both white shirts, jewelry, coffee glasses, pastries, table, cafe, sunset lighting and framing.
Everyday remixes and simple edits
Comparison 10
Seated In Leather Armchair Heathered Sweat Suit Bare Feet
A seated sweatshirt-and-joggers remix with soft indoor lighting.
Scroll sideways on smaller screens. Open any image to inspect the full-resolution output.
Read the exact prompt
A medium shot captures the person in the uploaded reference image (DO NOT TOUCH FACE) seated in a distressed brown leather armchair. The individual is wearing a light grey heathered cotton sweat suit, consisting of a crewneck sweatshirt and matching jogger pants. Their legs are drawn up towards their chest, with bare feet extended forward, filling the lower portion of the frame. The camera is positioned at a slightly low angle, looking up at the subject. The background features sheer, beige curtains drawn partially closed, and a portion of a textured rug is visible on the floor. The lighting is soft and diffused, coming from the front and slightly to the side, creating gentle shadows. The lens exhibits a shallow depth of field, blurring the background details.
Scroll sideways on smaller screens. Open any image to inspect the full-resolution output.
Read the exact prompt
ultra realistic studio portrait
Create an ultra-realistic, high-end professional studio portrait of one adult subject, styled in a clean, modern, executive fashion aesthetic.
The subject is seated on a simple white cube, body slightly angled toward the camera, posture upright yet relaxed. Hands gently clasped together, resting naturally on the lap. Expression is confident, warm, and composed, with a subtle professional smile.
Overall mood: polished, elegant, modern, trustworthy — pure realism only.
No stylization, no cinematic effects, no beauty filters, no exaggeration.
WARDROBE & STYLING
• Crisp white tailored blazer (matte fabric, no logos)
• Minimal white inner top / high-neck tee
• Neutral beige or tan tailored trousers
• Minimal gold jewelry only (thin chain, bracelet, ring — subtle and realistic)
• Clean, natural grooming
• Hair styled neatly and naturally, exactly matching the uploaded reference face
ENVIRONMENT & LIGHTING
• Seamless neutral grey studio background
• Soft, even studio lighting with natural falloff
• No harsh shadows, no dramatic contrast
• True-to-life skin texture and fabric detail
• Professional fashion/editorial realism
🔐 IDENTITY REFERENCE — ABSOLUTE (NON-NEGOTIABLE)
Use my uploaded face image as the ONLY facial identity reference
Maintain 100% facial accuracy at all times, including:
• Bone structure
• Jawline and chin
• Eye shape, eyelids, iris placement
• Eyebrow thickness and curvature
• Nose bridge, width, and tip
• Lip shape and proportions
• Cheek volume and facial symmetry
• Hairline and facial proportions
Face must remain perfectly consistent from every angle
(front view, ¾ view, side profile, seated pose, slight head tilt, full-body framing)
Identity preservation in the original bakery scene.
See the original scene and identity references
Image 1: source sceneImage 2: Lily identity
GPT Image 2 Low
Blocked by provider moderation
GPT Image 2.5 Flare Low
Blocked by provider moderation
GPT Image 2 Medium
Blocked by provider moderation
GPT Image 2.5 Flare MediumGPT Image 2.5 Sunburst Medium
Scroll sideways on smaller screens. Open any image to inspect the full-resolution output.
Read the exact prompt
A scene featuring the person in the uploaded reference image (DO NOT TOUCH FACE).
Image 1 is the frozen source scene. Image 2 is the target identity reference.
Replace only the primary visible adult subject in Image 1 with the person from Image 2.
Preserve the exact visible person count.
Replace every secondary person with a separate, distinct new adult identity that roughly preserves that secondary person's visible hairstyle silhouette, styling, clothing, pose, scale, and scene role from Image 1.
Use Image 2 exclusively for the primary subject. Never copy or duplicate the person from Image 2 onto a secondary person. Do not create twins or clones.
FACE EXPRESSION IS THE HIGHEST PRIORITY for the primary subject. Reproduce this literal facial performance: The head is turned over the shoulder toward the viewer's left. Both eyes are open and directed toward the camera, and the eyebrows are relaxed. The mouth is closed with neutral corners and low facial tension.
Remove every bright-green mannequin placeholder, green color, synthetic material, plastic texture, blank face, and placeholder artifact.
Preserve the exact person count, head and body poses, crop, clothing, props, lighting, shadows, background, camera angle, and composition.
Do not add or remove people. Do not preserve any original source identity.
Photorealistic adults, natural skin texture, anatomically coherent faces, hands, and bodies.
Keep one primary foreground reference subject visually dominant, with companions or background people clearly secondary.
Scroll sideways on smaller screens. Open any image to inspect the full-resolution output.
Read the exact prompt
a hyper-realistic, ultra-detailed portrait of the woman in the uploaded reference image (DO NOT TOUCH FACE) sitting casually on a step indoors against a soft gray wall. Her exact facial features are preserved without any alteration. She wears a fitted off-shoulder top and pink leggings, paired with Nike Air Force 1 sneakers and Nike socks. She holds a plastic cup of iced coffee with a straw in her hands. To her left lies a brown plush handbag, and to her right — a small stack of magazines and papers. The pose is relaxed, her gaze directed toward the camera, under gentle natural light. The scene conveys a casual sporty-chic aesthetic with a minimalist interior. Captured with a Canon EOS R5, full-frame lens, 8K ultra-high resolution, realistic textures, crystal-clear focus, sharp image details, cinematic depth of field, professional lighting and color grading. Strictly preserve the reference person's face and overall likeness from the uploaded reference image; change only the requested scene, pose, wardrobe, props, lighting, camera angle, and background. Render as a polished photorealistic image with natural skin texture, realistic hands, believable fabric, coherent anatomy, balanced exposure, and no duplicate faces, extra fingers, unreadable text, plastic skin, or overprocessed HDR.
Scroll sideways on smaller screens. Open any image to inspect the full-resolution output.
Read the exact prompt
Change only the white blazer to deep navy blue. Preserve its cut, buttons, seams and fabric texture. Keep the same adult woman, exact face, hair, white inner top, beige trousers, jewelry, hands, seated pose, white cube, lighting, background and framing unchanged.
Comparison 15
Studio portrait: background only
Replace the background while preserving the subject and lighting.
Scroll sideways on smaller screens. Open any image to inspect the full-resolution output.
Read the exact prompt
Replace only the grey studio background with a softly blurred modern office with tall windows and green plants. Keep the same adult woman, exact face, hair, white blazer and inner top, beige trousers, jewelry, hands, seated pose, white cube, camera angle, framing and lighting on the subject unchanged.
Original Explore remixes
Comparison 16
Taupe Cutout Dress at Garden Restaurant
Identity replacement in the original garden-restaurant scene.
Scroll sideways on smaller screens. Open any image to inspect the full-resolution output.
Read the exact prompt
A scene featuring the person in the uploaded reference image (DO NOT TOUCH FACE).
Image 1 is the frozen source scene. Image 2 is the target identity reference.
Replace only the primary visible adult subject in Image 1 with the person from Image 2.
Preserve the exact visible person count.
Replace every secondary person with a separate, distinct new adult identity that roughly preserves that secondary person's visible hairstyle silhouette, styling, clothing, pose, scale, and scene role from Image 1.
Use Image 2 exclusively for the primary subject. Never copy or duplicate the person from Image 2 onto a secondary person. Do not create twins or clones.
FACE EXPRESSION IS THE HIGHEST PRIORITY for the primary subject. Reproduce this literal facial performance: Her head is upright and nearly frontal with a slight tilt toward her right. Both eyes are open toward the camera, brows relaxed, mouth closed with a small soft smile, and cheeks are lightly lifted.
Remove every bright-green mannequin placeholder, green color, synthetic material, plastic texture, blank face, and placeholder artifact.
Preserve the exact person count, head and body poses, crop, clothing, props, lighting, shadows, background, camera angle, and composition.
Do not add or remove people. Do not preserve any original source identity.
Photorealistic adults, natural skin texture, anatomically coherent faces, hands, and bodies.
Keep one primary foreground reference subject visually dominant, with companions or background people clearly secondary.
Scroll sideways on smaller screens. Open any image to inspect the full-resolution output.
Read the exact prompt
A scene featuring the person in the uploaded reference image (DO NOT TOUCH FACE).
Image 1 is the frozen source scene. Image 2 is the target identity reference.
Replace only the primary visible adult subject in Image 1 with the person from Image 2.
Preserve the exact visible person count.
Replace every secondary person with a separate, distinct new adult identity that roughly preserves that secondary person's visible hairstyle silhouette, styling, clothing, pose, scale, and scene role from Image 1.
Use Image 2 exclusively for the primary subject. Never copy or duplicate the person from Image 2 onto a secondary person. Do not create twins or clones.
FACE EXPRESSION IS THE HIGHEST PRIORITY for the primary subject. Reproduce this literal facial performance: The head is nearly frontal with a very slight downward pitch. Both eyes are open and directed toward the camera, and the eyebrows are neutral. The lips are gently parted with no visible teeth and low facial tension.
Remove every bright-green mannequin placeholder, green color, synthetic material, plastic texture, blank face, and placeholder artifact.
Preserve the exact person count, head and body poses, crop, clothing, props, lighting, shadows, background, camera angle, and composition.
Do not add or remove people. Do not preserve any original source identity.
Photorealistic adults, natural skin texture, anatomically coherent faces, hands, and bodies.
Keep one primary foreground reference subject visually dominant, with companions or background people clearly secondary.
An emerald-dress portrait from the same frozen reference.
See the original photo
Original photo
GPT Image 2 Low
Blocked by provider moderation
GPT Image 2.5 Flare Low
Blocked by provider moderation
GPT Image 2 Medium
Blocked by provider moderation
GPT Image 2.5 Flare Medium
Blocked by provider moderation
GPT Image 2.5 Sunburst Medium
Scroll sideways on smaller screens. Open any image to inspect the full-resolution output.
Read the exact prompt
Reference image: uploaded photo. Recreate the image with ultra-realistic, high-fashion editorial accuracy. Preserve pose, angle, lighting, color palette, and mood exactly as in the reference. A woman crouched in a three-quarter turn, body angled away from the camera, head turned back over her shoulder. She holds a single white rose close to her face, fingers relaxed, gesture soft and controlled. Expression calm, intense, slightly distant, direct eye contact with the camera. Lips softly parted, gaze sharp and elegant. Outfit: deep emerald green satin slip dress with a low open back and thin spaghetti straps. Fabric smooth, fluid, with subtle sheen catching the light. Dress hugs the body naturally without exaggeration, mid-thigh length, minimalistic and refined. Hair: sleek, center-parted, pulled into a low polished bun. Hair glossy, tight, elegant, no loose strands. Makeup: sculpted editorial glam. Warm bronzed skin with strong highlight on cheekbones and shoulders, sharp elongated winged eyeliner, defined brows, nude glossy lips. Skin texture realistic, luminous, high-end beauty finish. Accessories: no visible jewelry. Focus remains on the dress, skin, and rose. Environment: neutral warm-toned studio background with soft shadows. Floor visible with scattered white rose petals near the subject, adding subtle narrative detail. No props besides the rose. Lighting: warm directional studio light from the side, emphasizing skin glow and muscle definition, soft shadows, cinematic contrast. No harsh highlights, painterly depth. Camera & style: medium-full portrait framing, shallow depth of field, sharp focus on face and upper body, luxury fashion editorial, ultra-realistic, Vogue-level aesthetics. No stylization, no reinterpretation, no added elements, no color shifts.
Price: Flare and Sunburst have the same listed rates
At these sizes, GPT Image 2.5 has lower listed prices than GPT Image 2. These examples include one input image. Longer prompts, extra references and request complexity can change the final bill; this is not our measured batch spend.
Output size
GPT 2 Low
GPT 2 Medium
2.5 Low
2.5 Medium
1024 × 1024
$0.015
$0.061
$0.00588
$0.01317
1024 × 1536
$0.018
$0.054
$0.00474
$0.01029
Listed rates checked September 9, 2026.
Completion and elapsed time
These totals cover all 21 tested cases, including the three fully moderated cases omitted from the image comparisons. Elapsed time includes provider queueing, polling and output download. Medians use completed requests only, so blocked rows change the sample. These timings are not an inference-speed guarantee. A generated image may still miss instructions.
Setting
Generated
Moderated
Median elapsed
GPT Image 2 Low
18/21
3
22.2 s
GPT Image 2.5 Flare Low
13/21
8
17.6 s
GPT Image 2 Medium
13/21
8
65.9 s
GPT Image 2.5 Flare Medium
14/21
7
22.8 s
GPT Image 2.5 Sunburst Medium
16/21
5
27.4 s
Earlier runs: moderation depends on the inputs
The initial six Explore tests generated 6 of 30 outputs; 24 requests were moderated. A second batch of everyday remixes and simple edits generated 27 of 30, with three moderation blocks on one bakery scene. These were different prompts and inputs, so the change does not measure a model upgrade or a general moderation rate. The totals include every attempt. Cases blocked by all five settings are omitted from the image comparisons.
Let the votes decide
We have not assigned editorial winners. The nine harder-edit cases form 76 blind pairwise comparisons for Arena. Vote for the image you prefer before revealing the model names. Arena records visual preferences; it does not separately score edit accuracy. Use the original photos and exact prompts above to inspect fidelity.