Glossary · AI Engineering

What are embeddings?

Short answer

Embeddings are lists of numbers (vectors) that represent the meaning of a piece of text, an image or another item, produced by an embedding model. Items with similar meaning get vectors that are close together, so software can find related content by comparing numbers instead of matching exact words. They power semantic search, RAG and recommendations.

How embeddings work

An embedding model reads an input, say “How do I reset my password?”, and returns a vector of a few hundred to a few thousand numbers. “I forgot my login” produces a vector pointing in almost the same direction, even though the two sentences share no words. Similarity is measured with cosine similarity or distance between vectors.

What they are used for

  • Semantic search: find documents by meaning rather than keywords.
  • RAG: retrieve the right passages to give an LLM.
  • Recommendations and duplicates: “similar articles”, de-duplicating support tickets.
  • Clustering and classification: group feedback by theme.

Practical tips

  • Use the same embedding model for indexing and for queries; vectors from different models are not comparable.
  • Store the model name with each vector so you can re-embed when you change models.
  • Combine semantic and keyword search (hybrid search) for product names, codes and exact phrases.
  • For moderate data sizes, PostgreSQL with pgvector is often enough; you may not need a separate vector database.

Published · Updated · By · All terms

Go deeper