Section
Embeddings

Embeddings

Vector representations for retrieval and similarity.

Embeddings turn text into vectors you can compare numerically. Use them for semantic search, retrieval-augmented generation, clustering, and similarity.

Use a model from your catalog

The examples below call nomic-ai/nomic-embed-text, the embeddings model these docs use throughout. Browse the Embeddings tab in Model Library for the embedding models available to your workspace, and substitute the id if you want a different one.

#Minimal example

from openai import OpenAI
import os

client = OpenAI(
  api_key=os.environ["OMNIVA_API_KEY"],
  base_url="https://api.tokenfactory.omniva.com/v1",
)

resp = client.embeddings.create(
  model="nomic-ai/nomic-embed-text",
  input="The quick brown fox jumps over the lazy dog.",
)
vector = resp.data[0].embedding
print(len(vector))

#Batching

Pass input as an array of strings. The response is an array of embedding objects in the same order as the input.

resp = client.embeddings.create(
  model="nomic-ai/nomic-embed-text",
  input=[
      "The quick brown fox jumps over the lazy dog.",
      "Pack my box with five dozen liquor jugs.",
      "How vexingly quick daft zebras jump.",
  ],
)

for item in resp.data:
  print(item.index, len(item.embedding))

Batch up to roughly 100 inputs per call for the best throughput-to-latency balance. Beyond that, parallelize across calls rather than packing more into one.

#Picking a model

Each embedding model has its own output dimensionality and trade-off between quality, latency, and cost. Open the Embeddings tab in Model Library and pick by:

  • Quality vs. cost — larger models usually score better on retrieval benchmarks but cost more per token and produce larger vectors to store.
  • Dimensionality — typical embedding dimensions are 256, 768, 1024, 1536, or 4096 — smaller dims are faster and cheaper, larger dims often more accurate. The model used throughout this guide (nomic-ai/nomic-embed-text) returns 768-dimensional vectors. Pick the smallest dimension that meets your retrieval quality bar.
  • Context window — long-document workloads need a model that can swallow the full document in one call rather than chunking pre-emptively.

Don't mix dimensions in the same index — pick one model and stick with it.

#Chunking long inputs

Documents longer than the model's context window need chunking. Reasonable defaults: 256–512 token chunks with 50–100 token overlap for long-form prose; sentence-aware splitting for natural-language content; structural splitting (per heading or per code block) for technical content. Embed each chunk separately and search at the chunk level — then re-rank or stitch results back into the original document at the application layer.

#Pick a similarity metric

  • Cosine similarity is the safe default for most embedding models, including nomic-ai/nomic-embed-text. Normalize vectors to unit length first — most embedding APIs return normalized vectors already (verify with numpy.linalg.norm(vec) == 1.0 or equivalent).
  • Dot product equals cosine for already-normalized vectors and is slightly faster — useful when your index supports it natively.
  • L2 / Euclidean distance is rarely what you want for semantic similarity. Avoid unless your model's documentation explicitly suggests it.

#Storing vectors

Most teams store vectors in a purpose-built index. Common options: pgvector if you already run Postgres, or managed services like Pinecone, Weaviate, or Qdrant. We return vectors as number[] — float32 is sufficient resolution for cosine and dot-product retrieval, and storing as float64 just doubles your bytes for no gain.

#Common mistakes

  • Mixing models across writes — embeddings from different models live in different vector spaces and aren't comparable. Re-embed the whole index when you switch.
  • Inconsistent normalization — if you cosine-similarity, normalize all vectors the same way, or rely on the model's native scale for every read and write.
  • Storing as float64 when float32 is enough — wastes space and bandwidth.
  • Re-embedding on every query when you could cache. Repeated identical inputs should hit a cache, not the API.
  • Sending too-long inputs. Every embedding model caps its input, and the cap varies widely — nomic-ai/nomic-embed-text accepts 8,192 tokens, while smaller models stop at 512. Hitting the cap typically returns a 400 error or silently truncates — chunk before embedding (see Chunking long inputs).
  • Asymmetric model used symmetrically. Some retrieval-tuned models expect a task prefix that marks a query as a query and a document as a document. Check your model's instruction card before reusing the same call shape for both sides.
  • Re-normalizing already-normalized embeddings. Most embedding APIs return unit-length vectors; re-normalizing is a no-op at best and lossy at worst.
  • Embedding whole documents without chunking. Long-form embeddings dilute meaning — a single 4000-token vector averages over too many topics. Chunk first, embed each chunk, search at the chunk level.
RAG: embed once, query many

For retrieval-augmented generation, generate embeddings once at ingest time and store them in your index. At request time, only embed the query. This avoids paying repeatedly to re-embed the same documents.

#What next

Was this page helpful?