All models

qwen-image-3.0-pro

AlibabaImage
Get your API key
qwen-image-3.0-pro

An Image Model for Complex Typography and Refined Commercial Visuals

qwen-image-3.0-pro is Alibaba's Qwen Image 3.0 Pro image model, designed to turn detailed creative requirements into visually rich, clearly structured finished images. It balances multi-region layouts, multilingual text, and subtle material rendering. On this platform, you can use a single entry point for text-to-image generation, reference image editing, and batch delivery, making it especially suitable for posters, menus, storyboards, and product concept design.

AlibabaModel Brand
ImageModel Type
Generate · EditCreation Method
STANDARD APIs · QUICK SETUP

Bring this model into your workflow

Submit requests to the public API at api.acedata.cloud using the documented parameters, then use the results in your application.

API hostapi.acedata.cloud
modelqwen-image-3.0-pro

Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.

Specifications and API Features

Native Input and Output
Text and image input; image output
Native Long-Instruction Capability
Up to 4.5k tokens of input; platform prompt limit is 18,000 characters
Native Text Rendering
Supports 12 languages, multiple fonts, and text rendering as small as 10px
Reference Images and Batching
The platform supports 1–3 reference images and generates 1–6 images per request
Output Dimensions
1K / 2K output; size is set by width*height, with aspect ratios from 1:8 to 8:1
Creation and Delivery Controls
Prompt expansion, negative prompts, watermark toggle; synchronous and asynchronous tasks with completion callbacks

Native capabilities describe the model's creative scope; use the platform entry-point instructions for reference image count, dimension settings, and task delivery methods. Token and character limits are calculated separately.

Core Capabilities

Bring text and visuals together in one composition

The focus of Pro is not just generating backgrounds, but organizing titles, body text, illustrations, and multiple content areas within the same image. You can specify layout sections, text hierarchy, and exact copy separately in the prompt to create menus, newspaper-style posters, or storyboard pages, reducing the work of arranging all the text again after creating the image.

Balance detail with realistic texture

The model emphasizes details such as micro-expressions, pores, and strands of hair, making it suitable for precise visual requirements for people and products. When creating, describe the subject, materials, lighting, and composition separately instead of relying only on abstract terms like “premium feel,” so the direction of the final image is clearer; important details should still be checked at original size.

Extend designs from reference materials

After adding reference images, you can modify the background, colors, style, or local elements around an existing subject. For multi-image tasks, clearly specify whether each image is responsible for the subject, composition, or style, and list the features that must be retained. Generation and editing share the same interface, with results delivered through image URLs for continued display, download, and distribution.

Use Cases

Multilingual menus and marketing posters

Enter approved dish names, campaign copy, brand colors, and regional layouts to generate menus or promotional posters with images and text. When product appearance needs to be retained, add reference images and then explore different compositions in batches. Before delivery, check prices, spelling, and text alignment item by item, and keep suitable versions for further refinement.

Educational diagrams and storyboard drafts

Organize key knowledge points, question text, or shot sequences into section-based prompts to generate exam-style pages, diagrams, or multi-panel storyboards. Clearly defining the subject, explanation, and sequence of each panel makes complex content easier to review. Generated images can be used as discussion drafts, but formulas, labels, and shot continuity should be verified separately.

Product and interface visual proposals

Provide product reference images, target audiences, and interface content to create visual concepts for websites, games, or livestream interfaces. Use specific sections to describe navigation, display areas, and button copy, creating full-page images for discussion; it delivers visual proposals, not functional websites or applications with real interactions.

How to choose this model

Choose Pro when output requirements are high

When tasks involve complex layouts, dense text, or business details that require careful viewing, Pro is better suited for high-precision final output. If the main goal is frequently trying styles, exploring compositions, and iterating in batches, consider qwen-image-3.0 from the same series. Both can use similar workflows, but you should explicitly specify the corresponding model ID to avoid treating Standard as Pro.

Plan creative steps by deliverable

If the core task is to generate complete copy and visuals together, first use a structured prompt to determine the layout, then refine the details; if an existing image needs to preserve the subject, add a reference image and clearly specify what must remain unchanged. For projects requiring editable text, precise vector graphics, or functional interfaces, you can use the Pro result as a visual draft, then complete the final production with design or development tools.

Get started

Prepare copy and reference materials

Organize the title, body text, and layout; when editing is needed, prepare 1–3 publicly accessible images and separately describe the subject, background, and style references.

Clearly choose Standard or Pro

Send model=qwen-image-3.0-pro and prompt to /qwen-image/images; start with size=2048*2048 and n=1. When editing, use image_urls; do not write dimensions separated by x.

Retrieve generated images and continue revising

Read image results synchronously, or use async=true to save task_id and then query /qwen-image/tasks; verify the text and subject, then use the selected image for the next editing round.

Trial suggestion: bilingual Chinese-English coffee menu

Input and goal

Create a one-page coffee menu, with “Espresso 浓缩” and “Latte 拿铁” arranged by category on the left and corresponding prices on the right; use only the provided copy, with off-white and dark brown as the main colors, plus a small cup illustration.

Acceptance criteria and next steps

Check the Chinese and English, numbers, and row-column alignment item by item, then check whether the cup and text are crowded; dense small text still needs to be reviewed in design software.

Usage Limits

  • Rendering text as small as 10px is a model capability, but does not mean that small text in every output can be used directly for publishing. Dense pages may still have missing characters, typos, or confused hierarchy; key copy should be written explicitly and checked at the final size, rather than only viewing thumbnails.
  • Knowledge-based images and interface simulations do not equal fact verification or online search. This model does not support Web Search; formulas, data, and labels in images should be based on the input materials. Attractive diagrams cannot replace content validation, and interface images will not automatically implement interactions.
  • Reference image editing does not guarantee that all unmentioned areas remain completely unchanged. Requirements for subjects, styles, and compositions across multiple images should be stated separately; the agent prompt extension is for text-to-image only and should not be relied on for editing tasks. Clearly specifying elements to preserve helps reduce unnecessary visual changes.

Frequently Asked Questions

How can I ensure I am calling Pro?

When submitting a request to POST /qwen-image/images, explicitly set model to qwen-image-3.0-pro and provide a prompt. Do not omit the model name or use qwen-image-3.0 instead; the latter is a Standard model in the same series, suited to different creative trade-offs.

How should reference images be organized?

Text-to-image generation is performed when image_urls is not provided; for editing, you can provide 1–3 publicly accessible images. Specify the purpose of each image in the prompt, for example, preserve the subject from the first image and reference the composition of the second, and state which features must not change to avoid attributes from different images being mixed together.

How do I set the aspect ratio and batch quantity?

Use a width*height pixel format for size, for example 1024*1024; the aspect ratio must be between 1:8 and 8:1. n can be set to 1–6, and one image is generated by default. Text-dense designs should be checked for readability at the delivery size rather than simply increasing the batch quantity.

How can long prompts and multilingual text be written more effectively?

First describe the work type, then list the exact copy, font direction, and visual requirements by area. Native long-instruction capability and platform character limits use different units and cannot be directly converted. Multilingual content should provide approved text; language support also does not mean that all font and small-text combinations can be rendered reliably.

How can generated results be integrated into business workflows?

You can receive results synchronously, or set async to true to obtain a task_id, then query through /qwen-image/tasks or receive a completion callback. Successful results include image_url and are stored on this platform's CDN, making them convenient for business systems to display, download, and distribute later.