Is gpt-transcribe the same as gpt-4o-transcribe?
When using it, treat gpt-transcribe as an independent call ID, and do not replace it yourself with gpt-4o-transcribe or gpt-4o-mini-transcribe. Similar names do not mean identical versions; model selection should focus more on your own recording samples and required delivery format.
How do I submit audio when calling it directly?
Submit the binary file to /v1/audio/transcriptions, and explicitly set model=gpt-transcribe. Do not rely on the default selection when the model is omitted. If submitting a URL through an MCP audio transcription tool, follow that tool's input method; do not treat the URL string directly as an uploaded file.
Can it generate SRT or VTT subtitles directly?
The endpoint provides srt and vtt request options. You can first test a short recording to see whether the subtitle result meets your needs. Before formal delivery, also check sentence segmentation, time alignment, and proper nouns; if the workflow primarily uses plain text, you can transcribe first and then arrange the timeline during subtitle editing.
How should I prepare requests when there is a lot of specialized vocabulary?
You can prepare concise language information and terminology context around language, prompt, or keywords[], and first test common names and easily confused terms. Prompts are not dictionaries that guarantee correct recognition, nor are they suitable for embedding rewriting instructions; key terms should still be checked against the original recording one by one.
Can I get real-time responses while speaking?
The main workflow of gpt-transcribe is audio file transcription, and its output is recognized text rather than conversational responses. Even when using streaming configuration, it should not be treated as a real-time two-way voice conversation; to listen and respond at the same time, you must separately design audio capture, conversation processing, and speech playback steps.