High-precision text embedding model for multilingual semantic retrieval
text-embedding-3-large is the high-precision variant among OpenAI's third-generation text embedding models. It converts natural language and code into numerical vectors for semantic search, knowledge base retrieval, and text clustering. It provides full representations of up to 3072 dimensions and also supports reduced dimensions, making it suitable for targeted trade-offs between retrieval quality and index size, rather than directly generating chat answers.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and interface features
Clarify capacity, inputs and outputs, and invocation methods before selecting a model.
Native full dimensions
Up to 3072 dimensions; outputs full dimensions by default when dimensions is not set
Dimension control
Supports using dimensions to shorten vectors; official examples include 1024 and 256 dimensions
Input formats
Non-empty text, text arrays, token integer arrays, or batches of token sequences
Batch item count
Text arrays or batches of token sequences support up to 2048 items
Output encoding
float numeric arrays or base64 strings; defaults to float
Invocation endpoint
POST /v1/embeddings; model uses text-embedding-3-large
Returned information
Per-item embedding and index, along with model, usage, and created
Vector dimensions are native model specifications; input organization, batch item count, and return encoding fall within the scope of this platform's embedding interface.
Core capabilities
Learn what text-embedding-3-large can bring to your work.
Build retrieval connections around semantics
The model converts text content into computable conceptual representations, so retrieval does not need to rely solely on whether keywords match exactly. Questions, passages, and code descriptions can each be converted into vectors, then compared for similarity by a retrieval system. It provides the representations, while document ranking, filtering, and final answers are handled by the application workflow.
Better suited for tasks that prioritize multilingual quality
In multilingual retrieval and English-task evaluations at release, large achieved higher average scores than the same-generation small and the previous-generation ada-002. For multilingual knowledge bases, queries with diverse phrasing, or content that requires distinguishing between closely related topics, it can be considered when quality is the priority, with results then validated using real queries.
Adjust index size with dimensions
Full vectors provide representations of up to 3072 dimensions, and output can also be shortened through dimensions. Shorter vectors can reduce the burden of vector storage and similarity computation, but may sacrifice some accuracy. When an existing vector database has dimensional constraints, test a shortened configuration first instead of immediately replacing the entire embedding model.
Use cases
Start with specific tasks to find where the model can make an impact.
Retrieval layer for enterprise knowledge bases
Split product manuals, operating procedures, and frequently asked questions into text passages, generate vectors in batches, and save their corresponding document locations. When users ask questions, generate vectors for the questions using the same model and dimensions, retrieve relevant passages, and pass them to an answer model. The deliverable is a searchable content index, not answers written directly by the embedding API.
Unified search for multilingual materials
For help centers or research repositories containing materials in different languages, convert titles, summaries, and body excerpts into vectors to establish a foundation for semantic retrieval. Before launch, prepare representative queries for each language and check whether relevant content can enter the candidate results; then combine filters for language, product, and time to create a usable search list.
Feedback categorization and similar-content discovery
Input ticket descriptions, user feedback, or code explanations to obtain a corresponding vector for each item, for subsequent clustering and similarity analysis. This can help identify feedback with similar topics, potential duplicate content, and related technical explanations. Final category labels and duplicate determinations should be confirmed by business rules or manual spot checks, rather than by directly reading vector values.
How to choose this model
Choose based on task complexity, input materials, and expected results.
Trade-offs with small: define the quality target first
text-embedding-3-small is more geared toward efficient applications, while large is better suited for tasks that prioritize retrieval quality. Do not replace all indexes simply because the model is larger: first fix the corpus, queries, and retrieval workflow, compare relevant-document recall, then evaluate vector size and computational burden. If small already meets requirements, keep it; if difficult cases remain common, then test large.
Migrating from ada-002: redesign vector configuration
Compared with text-embedding-ada-002, large offers stronger published benchmark performance and native dimension-shortening capability. The official documentation also provides an example where large's 256-dimensional representation exceeds full ada-002 on MTEB. When migrating, regenerate document and query vectors, and standardize the model and dimensions; do not merely change the query-side model while continuing to use the old index.
Prepare the inputs first, then connect them to the corresponding application workflow.
Prepare inputs
Select Chinese and English document excerpts and real queries, retaining a document ID, title, and access permissions for each excerpt.
Organize calls and downstream workflow
Use /v1/embeddings to select text-embedding-3-large, and manage the text and its document ID separately. Write vectors to an index with a fixed configuration, and use the same encoding for queries; first confirm that candidate source text is traceable, then use retrieval results in the application.
Design the task directly from the inputs and acceptance priorities below.
Recommended task
Use the same model and dimensions to generate vectors for documents and queries, then perform similarity retrieval; compare full vectors with shortened-dimension configurations.
Key checks
Calculate recall for manually labeled relevant passages, and check cross-lingual and synonymous expressions; retain original sources in the results, while a downstream model is responsible for generating answers.
Usage boundaries
Before formal use, understand the output quality and capability scope.
It is an embedding model, not a chat, translation, or summarization generation model. It returns numerical representations and cannot be used directly as readable answers or classification rationales; knowledge-base Q&A still requires retrieval logic and a response model, and document search is not automatically completed simply because vectors have been generated.
Reducing dimensions is not lossless compression, and full vectors do not mean that all tasks can achieve the best retrieval results. When choosing dimensions, evaluate both relevant-content recall and index size; the document side and query side must maintain consistent configurations, and vectors generated by different models should not be mixed for comparison.
A maximum batch size of 2048 items is an input-entry constraint and does not mean that each text can be arbitrarily long. Long documents should first be split into semantically complete paragraphs, and empty text should be removed; PDFs, images, or audio must first be extracted into text and cannot be submitted directly as these text input formats.
Frequently Asked Questions
Answers to common questions about using text-embedding-3-large.
Can text-embedding-3-large directly answer knowledge base questions?
It cannot directly generate answers. It converts questions and document chunks into vectors, helping the retrieval system find relevant content. A complete question-answering workflow typically retrieves relevant passages first, then uses a generation model to compose the answer; the embedding returned by the embeddings API is not itself an answer and does not include an explanation.
Should I use the full 3072 dimensions or reduce the dimensionality?
When quality is the priority, you can start with the full dimensionality; if your vector database has dimension limits or the index is large, you can test a reduced-dimension configuration. Official examples provide 1024- and 256-dimension options, but the specific trade-off should be determined using your own query set after comparing retrieval performance and storage overhead.
Why choose large instead of text-embedding-3-small?
At release, large performed better on average in multilingual retrieval and English task evaluations, making it suitable for applications with higher retrieval quality requirements. small is oriented toward efficiency. It is recommended to compare them using the same corpus and real queries, focusing on difficult cases rather than deciding based solely on the model name.
After batch input, how do I match each result to its input?
You can place multiple non-empty texts in the input array to generate multiple vectors in a single request. In the response data, each item includes an index and embedding, which can be used to associate it with the original input. When saving vectors, you should also save the document identifier, chunk position, and dimensions used to facilitate later retrieval.
How should I choose between float and base64?
float returns an array of numbers by default, suitable for direct use with programs or vector databases that require numeric vectors. base64 returns an encoded string, suitable for systems that already have a corresponding decoding workflow. They change the return representation, not the embedding task, and do not replace dimensionality control through dimensions.