ToolNest

Month 1 ยท LLM Applications

Day 3 โ€” Embeddings and Vector Similarity

Published September 14, 2026

Embeddings are the quiet foundation of everything from day 20 onward โ€” semantic search, RAG, memory, deduplication all reduce to one operation: turn text into a vector, compare vectors by distance. Today builds that operation from nothing and verifies it experimentally, because the fastest way to trust embeddings is to watch them cluster meaning without being told to.

The day's deliverable is a twenty-line script with a clear verdict: for a handful of sentence pairs, related sentences must land closer together than unrelated ones. If that sanity check fails, nothing built on top later can work.

The short answer

An embedding maps text to a dense vector where geometric closeness encodes semantic similarity โ€” measured for text with cosine similarity, which compares direction not length. Near-meaning sentences land near each other regardless of wording. This single property is the foundation of semantic search, RAG retrieval and agent memory.

What an embedding is, in plain terms

An embedding model is a neural network with a narrow output: a fixed-length vector โ€” 1,536 floats, or 1,024, depending on the model โ€” that summarizes the meaning of its input. The model was trained so that texts appearing in similar contexts produce nearby vectors. Nobody labels the dimensions; meaning is distributed across all of them.

The property that matters is contrastive: "a cat slept on the sofa" and "the cat napped on the couch" land close together even though they share almost no words, while "a cat slept on the sofa" and "quarterly revenue fell" land far apart. Keyword search cannot see that first relationship; embeddings exist precisely for it.

Cosine similarity versus Euclidean distance

For text, cosine similarity is the default: it compares the angle between vectors, ignoring magnitude โ€” and magnitude often encodes artifacts like text length rather than meaning. Euclidean distance measures absolute separation and is sensitive to those artifacts.

There is a practical shortcut worth knowing: with normalized vectors (length one), ranking by cosine and ranking by Euclidean distance produce identical orderings, because the two measures become monotonic transforms of each other. Many vector libraries normalize on write precisely to make this cheap.

Choosing an embedding model

Selection criteria in rough priority: does it support your languages (multilingual models matter the moment you mix scripts); what is the maximum input length against your expected chunk size; what are the dimensions (storage and index cost scale with them); how does it rank on retrieval benchmarks such as MTEB; and what does a million tokens cost. The famous proprietary model and a strong open-source alternative are both acceptable answers โ€” what is not acceptable is picking one and never re-running the sanity check.

  • One model per index, forever: vectors from different models live in different spaces and must never be mixed.
  • Sanity-check script belongs in the repo โ€” rerun it when swapping models.
  • Record dimensions and max input length next to the index; future-you needs both.

Today's hands-on task

Write the similarity script. Take three sentence pairs: two paraphrases (expected close), one related-topic pair (moderately close), and one unrelated pair (expected far). Embed all sentences, print the cosine similarity for each pair, and assert the ordering: paraphrase > related > unrelated. Twenty lines, no vector database โ€” the point is to see the geometry work before adding infrastructure.

How an interviewer asks about today

Day 3 interview questions
QuestionWhat a strong answer covers
What problem do embeddings solve?They turn semantic similarity into geometry โ€” enabling search, clustering and retrieval by meaning instead of keyword overlap.
Cosine or Euclidean โ€” and why?Cosine for text (direction = meaning, magnitude = noise); with normalization the orderings coincide, which is why libraries normalize.
How would you choose an embedding model?Language coverage, input length vs chunk size, dimensions vs storage cost, MTEB ranking, price โ€” then validate on your own data, not the leaderboard.

Common mistakes on day three

  • Mixing vectors from different models in one index โ€” the distances are meaningless across spaces; rebuild the index when the model changes.
  • Testing similarity only on unrelated pairs, where everything scores low and any model looks fine. Contrasting pairs are what reveal quality.
  • Skipping normalization, then debugging why scores look inconsistent โ€” cheap to do at write time, expensive to diagnose later.

Frequently asked questions

Do embeddings understand meaning or just statistics?
They encode statistical co-occurrence patterns at a scale that behaves like meaning โ€” paraphrases land close together without sharing words. Whether that counts as understanding is philosophy; that it enables retrieval is engineering.
Can I use embeddings instead of a full RAG pipeline?
Embeddings plus a vector index is the retrieval heart of RAG, but RAG adds chunking, re-ranking, grounding and citations. A raw similarity search is the right day-3 tool and the wrong production answer.
How long can the input text be?
Each model has a maximum input length โ€” often 512 to 8,192 tokens. Longer text must be split and embedded per chunk, which is exactly what the chunking strategies later in this series formalize.

Part of a public learning journal โ€” general educational content, not professional advice. See our disclaimer.