An image creation model balancing precise image editing and text layout
GPT Image 1.5 is OpenAI's image generation and editing model. It can create new visuals from text and make targeted modifications to existing images. It is especially suited for tasks that need to preserve people's appearances, product features, brand marks, and original compositions, while improving complex instruction following and dense text rendering to help marketing, e-commerce, and design teams continuously refine the same set of visual assets.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and interface features
Creation method
Text-to-image generation; image editing with text instructions
Reference image input
Platform editing endpoint: a single image URL or an array of up to 16 URLs
Output files
Platform endpoint: png, jpeg, webp
Result format
Platform endpoint: url or b64_json; supports image results and the task_id response structure
Background control
Platform endpoint: transparent, opaque, auto
Editing fidelity control
Platform editing endpoint: input_fidelity can be high or low
Image creation and detail preservation are model capabilities; the number of reference images, file formats, and result formats are invocation specifications for this platform's endpoint.
Core capabilities
Modify targets while preserving the core of the image
GPT Image 1.5 excels at adding, removing, combining, or replacing elements in existing images while preserving lighting, composition, and people's appearances as much as possible. When changing clothing, backgrounds, or product colors, clearly specify what must remain unchanged so the edit focuses on the target rather than redesigning the entire image.
Complex compositions follow instructions more closely
For creative requirements involving multiple objects, clear positional relationships, or segmented layouts, the model places greater emphasis on relationships between elements. Describe subjects, spatial positions, visual hierarchy, and style separately for advertising concept images, composite visuals, and informational posters; complex tasks should still be checked for accurate object counts and placement.
Create text and brand visuals together
The model improves rendering of smaller, denser text and enhances preservation of brand marks and key visuals during editing. It is suitable for developing headlines, product images, and layouts together, reducing drafts where text and visuals feel disconnected. Before formal publication, proofread all text item by item and check whether mark details have changed.
Applicable Scenarios
E-commerce Product Scene Expansion
Provide a product image and a description of the target scene, specifying that the shape, color scheme, and branding must be preserved while changing only the environment, props, or display angle. This can be used to create candidate images for product catalogs, holiday scenes, and presentation assets in different styles for team selection; product details should be compared item by item with the physical product to avoid treating generated variations as actual specifications.
Marketing Posters and Campaign Assets
Provide the campaign theme, headline copy, brand images, and layout requirements, and have the model generate advertising visuals containing text or modify the background and decorative elements of an existing poster. The deliverables can serve as social media assets, campaign visual drafts, and design proposals, followed by review for text accuracy, information hierarchy, and brand consistency.
Character Styling and Style Exploration
Based on a portrait photo, describe changes to clothing, hairstyle, environment, or artistic style, and clearly require that facial features and pose be preserved. Suitable for creating styling proposals, character visual drafts, and candidate creative portraits. When making successive edits, keep the original image for comparison to help identify deviations in appearance or local details.
How to Choose This Model
How It Compares with GPT Image 1
If the task focuses on precise edits to existing images, preserving brand elements, or visual creation containing substantial text, GPT Image 1.5 is more worthwhile to test first than GPT Image 1. Its improvements focus on instruction following, image preservation, and editing quality; if the existing workflow already meets your needs, compare the results using the same set of assets before deciding whether to make adjustments.
Creating from Scratch or Editing an Original Image
When there are no visual assets yet and you want to explore the overall direction, choose text-to-image generation; when the product, person, or composition has already been determined and only local content needs to be changed, choose image editing. For detail-sensitive tasks, try the high input fidelity setting and write the changes and elements to preserve separately, avoiding the assumption that high fidelity means pixel-level invariance.
Getting Started
Choose text-to-image or image editing
Provide a prompt when generating; provide both image and editing instructions when editing. Clearly specify the original text, subject preservation requirements, and target aspect ratio.
Specify the model and parameter format
Use the image generation or image editing endpoint, explicitly specifying model=gpt-image-1.5; use auto or WIDTHxHEIGHT for size, and generate one image first before evaluating. Configure masks, quality, and file format according to the guide for this endpoint.
Check images and usage records
Read the URL or Base64 image according to the response format, save the task_id asynchronously before querying the result; check text, reference details, and the alpha channel, and record usage according to the current Pricing rules.
Trial recommendation: precise product editing
Input and goal
Use the original image as the basis, preserving the packaging, logo, perspective, and all text; change only the tabletop background to off-white, and add soft side lighting while keeping the shadows natural.
Acceptance and next steps
Choose input_fidelity=high in supported editing requests, but still compare the logo, small packaging text, and outlines item by item; fidelity control is not pixel-perfect locking.
Limitations
Preservation of human appearance has improved, but images with multiple people may still show deviations in facial details, and consecutive edits do not guarantee that identity features remain completely unchanged. When a consistent person is needed, compare against the original photo and check the eyes, contours, and accessories; avoid checking only the overall atmosphere.
Improved text rendering does not mean layout and spelling are completely accurate, especially for multilingual text, dense small print, and complex marks. Dates, prices, names, and body text in posters should be proofread separately; for content requiring precise typography, use design tools to complete the final layout after the image is finalized.
Generated images may contain scientific or factual errors and may not fully achieve a certain style. The model is suitable for visual creation and asset iteration, but should not be used directly as an accurate product structure diagram or scientific illustration; structure, proportions, and professional information require human review.
Frequently Asked Questions
How does GPT Image 1.5 generate new images?
Submit model=gpt-image-1.5 and prompt to /openai/images/generations. The prompt should describe the subject, background, lighting, layout, and any text that needs to appear; when there are no constraints from an original image, this method is better suited for exploring the complete image and visual direction.
How do I use GPT Image 1.5 to edit an existing image?
Submit image and prompt to /openai/images/edits, and explicitly specify model=gpt-image-1.5. image can use a single image URL or an array of URLs; the instructions should separately specify the elements to change and the parts that must be preserved, such as people, products, and composition.
Can I create with multiple reference images?
The editing endpoint allows up to 16 image URLs. You can provide references for people, products, or environments, and explain in the prompt the role of each image, which elements need to be combined, and which cannot be changed. When there are many reference images, avoid conflicting style and appearance requirements.
Is it suitable for generating posters with Chinese text?
You can try creating posters with Chinese titles and copy. The model has improved dense and small-font text rendering, but multilingual performance still has limitations. Clearly provide the original text, text positions, and hierarchy, and check every character after generation; long body text or strict brand fonts should be typeset separately.
How can generated results be integrated into an application?
You can choose url to obtain the image address, or choose b64_json to obtain encoded image data, and select PNG, JPEG, or WebP according to your use case. Responses may contain created/data or may return task_id; applications should handle completed images and task responses separately, and should not treat a task ID as an image result.