Can GPT-4o mini view images and generate images?
It can combine images and text questions to generate text answers, making it suitable for image descriptions and visible information extraction, but it is not an image generation model. In Chat Completions, you can combine text and image_url blocks in the content array; if you need an actual finished image, use a dedicated image generation or editing model.
How should I choose between Responses and Chat Completions?
Applications that already use a messages conversation structure can use Chat Completions and read answers from choices; responsive workflows can use Responses and submit input. Both use model: gpt-4o-mini, but their input organization and response parsing differ.
Do I need to resend the conversation history for every turn?
When using Chat Completions, put relevant history in messages; when using Responses, organize input and related conversation content according to the documentation. Provide the latest material, revision goals, and key constraints each turn; for longer tasks, retain interim summaries and a final version that can be reviewed independently.
Will function calling directly execute my business code?
Defining a function alone does not automatically execute code. The model can return a function name and arguments, while the application is responsible for validating the arguments, performing the corresponding operation, and returning the result. Tasks such as querying orders or sending emails should have separate permissions and execution logic; in particular, do not omit confirmation steps for write operations.
Can I send a PDF directly as an image?
PDF files and image inputs are not the same type of request. For GPT-4o mini, clear images and already extracted text are the main input methods described here. When you need to process a PDF, first extract the body text or convert relevant pages to images, then submit the content according to goals such as summarization, field extraction, or question answering.