Vector Databases, Embeddings and Reranking: Three Different Parts of Retrieval

Embeddings, vector databases and rerankers are three different parts of retrieval. An embedding model converts text or other data into numerical representations; a vector database or vector index stores and searches those representations to retrieve candidate items; a reranker takes a smaller candidate set and reorders it using a more expensive relevance model or scoring method. They often appear together in RAG, but none of them is the same thing as RAG, and none is mandatory in every retrieval system.
What this really means
Search systems have two competing goals: find enough potentially relevant material and put the best material near the top. Fast first-stage retrieval usually optimizes candidate generation. A stronger second-stage model can then spend more computation distinguishing the best candidates.
Embeddings, vector indexes and rerankers occupy different positions in that process. Treating them as one feature hides important design choices about recall, precision, latency, storage, metadata filtering and model cost.
The distinction also prevents a common RAG mistake: assuming that storing document embeddings in a vector database automatically creates high-quality retrieval. Retrieval quality depends on the embedding model, chunking, metadata, query construction, index configuration, candidate count, hybrid retrieval, reranking and the authority of the underlying sources.
The simplest example
Suppose a knowledge base contains 100,000 document chunks. A user asks: “How do I revoke an API token?”
First, an embedding model can encode the query into a vector. Document chunks may already have their own stored embeddings. A vector search then compares the query vector to the indexed document vectors and returns, for example, 30 likely candidates.
Those 30 candidates can then be passed to a reranker. The reranker compares the query more directly with each candidate and produces a new relevance ordering. The application might keep the best five for the model context.
A basic two-stage semantic retrieval pipeline
Where the simple example stops
Real retrieval systems do not have to use dense embeddings at all. Keyword search such as BM25 can be the first-stage retriever. Sparse learned retrieval, SQL filters, graph traversal or application APIs can also generate candidates.
A reranker also does not care that the candidates came from a vector database. It can rerank BM25 results, hybrid results, hand-selected documents or candidates from multiple retrievers.
Likewise, embeddings do not require a specialized vector database. Small datasets can be compared in memory or with general-purpose databases and vector extensions. Specialized vector systems become useful when indexing, approximate nearest-neighbor search, filtering, scale, update behavior or operational requirements justify them.
Three different retrieval components
| Embedding | Vector database / index | Reranker | |
|---|---|---|---|
| Primary job | |||
| Typical input | |||
| Typical output | |||
| Cost profile | |||
| Typical failure |
Embeddings: representation, not retrieval
An embedding is a numerical representation produced by a model. For semantic retrieval, texts with related meaning are intended to occupy useful positions in a vector space so that a similarity or distance function can compare them.
Sentence-BERT was an influential step in making sentence-level semantic similarity practical with bi-encoder-style representations that can be computed independently and compared efficiently. The general idea remains central to modern dense retrieval: precompute document representations, compute the query representation at search time, then compare them.
The embedding itself does not search a corpus. It is data produced by an embedding model. Retrieval begins when the system compares the query representation against stored candidates.
The embedding model defines the representation space
Document and query vectors must be compatible with the model and configuration used to create them. Replacing an embedding model can change dimensionality, similarity behavior, language coverage and domain performance.
That is why an embedding-model migration is not merely an API-name change. Existing documents may need to be re-embedded and the index rebuilt or versioned.
Dense and sparse representations are different
Dense embeddings usually contain many non-zero dimensions and are commonly used for semantic similarity. Sparse representations contain many zeros and can preserve stronger token- or term-like structure.
Both can support semantic retrieval, and modern search systems can combine dense, sparse and lexical signals. “Vector search” therefore does not always mean one dense cosine-similarity pipeline.
Similarity functions are part of the representation contract
Cosine similarity, dot product and Euclidean distance do not mean the same thing. The correct metric depends on how the embedding model was trained and normalized.
Current Qdrant documentation, for example, requires a distance metric as part of vector configuration and documents cosine, dot-product and Euclidean-style choices. The important architectural rule is to treat the metric as part of the embedding/index contract rather than choose one arbitrarily.
Vector databases and indexes: candidate retrieval
A vector database or vector-capable search system organizes vector representations so the application can retrieve nearby candidates efficiently. Practical systems usually associate vectors with IDs and payload metadata such as source, language, tenant, document type, timestamp or access scope.
Qdrant, for example, organizes data into collections of points where a point contains a vector and optional payload metadata. Its documentation describes HNSW-based similarity search and metadata filtering as separate capabilities of the retrieval layer.
That distinction matters: the vector index answers a nearest-neighbor problem, while payload filters enforce structural constraints such as tenant, document class or language.
Approximate nearest-neighbor search trades exactness for efficiency
Comparing one query vector against every vector can be practical for small collections but expensive at large scale. Approximate nearest-neighbor indexes such as HNSW reduce search cost by navigating an index structure instead of exhaustively scanning every vector.
Approximate search introduces a recall/latency trade-off. Faster search can miss candidates that exact search would return. Index parameters therefore affect retrieval quality, not just infrastructure performance.
Qdrant exposes both HNSW-related parameters and an exact-search option, illustrating that vector storage and approximate retrieval policy are separate decisions.
Metadata filtering belongs before or during candidate retrieval
If the user may only access tenant A, retrieving semantically similar chunks from tenant B and attempting to remove them later is the wrong security boundary. Authorization and hard eligibility filters should constrain the candidate space before those candidates can influence downstream processing.
The same principle applies to locale, document status, source class, date, product version and other deterministic constraints. Similarity should rank eligible candidates; it should not override eligibility.
A vector database is optional
For a small corpus, brute-force cosine comparison may be simple and sufficient. A relational database with vector support may also be adequate. A dedicated vector database becomes valuable when its indexing, filtering, distributed storage, update behavior or operational features solve a real requirement.
Choosing a vector database because “RAG needs one” reverses the architecture process. Start with retrieval requirements and scale, then select the storage/index technology.
Reranking: second-stage relevance refinement
A reranker receives a query and a smaller set of already retrieved candidates, then assigns stronger relevance scores or a new ordering. It is normally more computationally expensive than first-stage retrieval, which is why it is applied after candidate generation rather than to the entire corpus.
Current Elastic guidance describes semantic reranking as a final-stage technique over a small top-k set and notes that it can refine lexical, semantic or hybrid retrieval. Cohere documents the same architecture: first-stage lexical or semantic search followed by a reranking stage.
A common implementation uses a cross-encoder-like model that examines the query and each candidate together. That richer interaction can distinguish relevance more precisely than independent embedding similarity, but it is much more expensive at corpus scale.
Bi-encoder retrieval and cross-encoder reranking solve different cost problems
| Property | Bi-encoder / embedding retrieval | Cross-encoder-style reranking |
|---|---|---|
| Encoding | Query and documents represented independently | Query and candidate processed jointly |
| Document computation | Can be precomputed at ingest | Normally recomputed per query-candidate pair |
| Corpus-scale search | Suitable with vector indexes | Usually too expensive across the entire corpus |
| Typical role | High-recall candidate generation | High-precision ordering of a small candidate set |
| Main trade-off | Fast and scalable but relevance interaction is compressed into vectors | Richer relevance judgment but higher latency/cost |
A reranker cannot recover what retrieval missed
If the relevant document is absent from the candidate set, reranking has nothing to promote. This is the central reason to evaluate retrieval and reranking separately.
A pipeline can have excellent reranker precision and still fail because first-stage recall is poor. Increasing reranker quality will not repair missing source coverage, bad chunking, restrictive filters or a weak candidate retriever.
Hybrid retrieval is a separate design choice
Dense semantic retrieval is strong when query and document use different wording but express related meaning. Lexical retrieval is strong when exact terms, identifiers, names, codes or rare phrases matter.
Hybrid retrieval combines multiple candidate signals, often lexical BM25 and vector similarity, then merges rankings using a method such as Reciprocal Rank Fusion or a weighted score combination.
Reranking can then operate on the fused candidate set. Hybrid retrieval and reranking are therefore complementary but distinct stages.
BM25 is not obsolete because embeddings exist
Keyword search can outperform dense retrieval for exact identifiers, version numbers, error messages, product codes and specialized vocabulary. SQLite FTS5, for example, includes a BM25 ranking function for full-text search.
A strong retrieval architecture can use lexical retrieval as the only first stage, vector retrieval as the only first stage, or combine both depending on the corpus and query distribution.
Chunking changes what embeddings and rerankers can see
If a document is split poorly, no later retrieval component can fully reconstruct the missing semantic unit. A chunk that cuts a condition away from its exception may embed misleadingly and may also be reranked incorrectly because the candidate text is incomplete.
Chunk size, overlap, structural boundaries and metadata therefore affect both candidate recall and reranker judgment. Retrieval evaluation should test the complete ingestion-to-ranking pipeline, not only the embedding model.
Do not compare retrieval scores as if they were universal probabilities
Cosine similarity, BM25 scores, sparse-vector scores, RRF ranks and reranker scores have different meanings. A score of 0.82 from one embedding model is not automatically comparable with 0.82 from another model or with a reranker score.
Thresholds should be calibrated for the actual model, corpus and task. Current Elastic guidance also notes that embedding similarity scores can be query-dependent, which makes universal cutoffs risky.
Evaluate retrieval stages separately
| Layer | Useful question | Example metric or test |
|---|---|---|
| Source coverage | Does the corpus contain the needed information? | Coverage audit / known-answer source set |
| Chunking | Is the needed evidence retrievable as a coherent unit? | Chunk-level support review |
| First-stage retrieval | Does the relevant item enter the candidate set? | Recall@k |
| Ranking | How high does relevant evidence appear? | MRR, nDCG, precision@k |
| Reranking | Does second-stage scoring improve ordering? | Delta nDCG / MRR / precision |
| Context selection | Do the final selected passages contain sufficient support? | Context relevance / coverage |
| Answer stage | Does the model use the selected evidence correctly? | Faithfulness / claim-evidence evaluation |
This separation is operationally important. If Recall@50 is poor, the reranker is not the first component to fix. If Recall@50 is strong but the best passage remains at rank 38, reranking or ranking fusion becomes a plausible target.
Which layer actually failed?
Symptoms and likely retrieval layer
| Observed symptom | Likely layer | First diagnostic | |
|---|---|---|---|
| Relevant document never appears | |||
| Relevant document appears too low | |||
| Semantically good but forbidden result | |||
| Relevant but outdated result | |||
| Correct result retrieved but omitted from prompt |
Relevance and Source of Truth are different
A reranker can make a stale document look extremely relevant. A vector index can retrieve a secondary summary that is semantically closer than the primary source. Retrieval quality therefore cannot replace authority rules.
Where source authority matters, metadata filters, source classes, version rules and provenance should constrain retrieval before the result becomes model context.
Original implementation evidence
Source of Truth Research Engine: lexical and semantic retrieval are separate
The Source of Truth Research Engine contains a local lexical retrieval path using SQLite FTS5/BM25 and a separate optional semantic retrieval path using locally generated embeddings.
Its semantic search implementation computes a query vector and compares it with stored chunk vectors using cosine similarity. The project deliberately treats semantic similarity as a discovery signal rather than evidence: a candidate must still be traced back to a concrete source and locator before it supports a claim.
This is useful implementation evidence for R01 because the same corpus can support lexical ranking and vector similarity without confusing either mechanism with evidentiary authority.
Aaasaasa AI Client: Qdrant is a vector infrastructure component
Aaasaasa AI Client includes Qdrant/vector infrastructure as a separate local resource. The Electron architecture exposes Qdrant services from the trusted main-process side rather than treating vector search as part of the model itself.
The repository contains a Qdrant client adapter, Qdrant service configuration and Docker-based Qdrant infrastructure. This demonstrates the architectural separation between AI provider/model execution and vector storage/search.
The existence of Qdrant support should not be overstated as a complete production RAG pipeline. The evidence here is narrower: vector infrastructure is implemented as its own component boundary.
| Implementation evidence | What it demonstrates |
|---|---|
| SQLite FTS5/BM25 in Source of Truth Research Engine | Lexical retrieval can exist independently of embeddings. |
| Local Ollama embeddings | Representation generation is its own stage. |
| Stored semantic vectors + cosine comparison | Semantic retrieval consumes embeddings after they have been produced. |
| Qdrant support in Aaasaasa AI Client | Vector storage/search is an infrastructure capability separate from the model provider. |
| Evidence/provenance rules in Source of Truth Research Engine | Retrieved similarity does not equal authority or proof. |
| No claimed custom reranker in these implementations | Reranking is explained as an architectural stage, not falsely claimed as already implemented evidence. |
When do you need each component?
| Need | Likely component |
|---|---|
| Semantic similarity across different wording | Embedding model + vector similarity search |
| Efficient search over a large vector corpus | Vector index/database or vector-capable search engine |
| Exact identifiers, error codes or rare terms | Lexical/full-text retrieval such as BM25 |
| Both exact terminology and semantic meaning | Hybrid lexical + semantic retrieval |
| Candidate set is good but ordering is weak | Reranker |
| Relevant items are absent from candidate set | Improve source coverage, chunking, retriever, filters or candidate count before reranking |
| Hard tenant/source/version constraints | Deterministic metadata/authorization filtering |
| Small corpus | Potentially simple brute-force similarity or general-purpose database rather than dedicated vector DB |
A practical retrieval design sequence
Design retrieval from requirements, not from product names
Common misconceptions
| Misconception | Correction |
|---|---|
| “An embedding is a vector database.” | An embedding is a representation; the database/index stores and searches representations. |
| “A vector database creates semantic meaning.” | The embedding model creates the representation; the vector system indexes and compares it. |
| “RAG requires a vector database.” | RAG requires retrieval, not a specific retrieval technology. |
| “Reranking is the same as vector search.” | Vector search generates candidates; reranking reorders a candidate set. |
| “Rerankers fix poor recall.” | They cannot promote a document that was never retrieved. |
| “Dense search replaces BM25.” | Lexical search remains valuable for exact terms, identifiers and specialized vocabulary. |
| “Higher similarity means more authoritative.” | Similarity and source authority are different dimensions. |
| “More top-k always improves RAG.” | Larger candidate sets can improve recall but add latency, noise and context-selection burden. |
| “One score threshold works everywhere.” | Scores depend on model, query, corpus and retrieval method and must be calibrated. |
| “A dedicated vector DB is always more advanced.” | It is only justified when its operational and retrieval capabilities match the requirements. |
Edge cases and limitations
Some applications do not need semantic search. Exact database lookup or structured SQL can be more correct, faster and easier to audit than embedding retrieval.
Some corpora are so small that a full vector scan is acceptable. Approximate indexing adds complexity without meaningful benefit.
Some queries require high recall before any precision optimization. Legal discovery, research and compliance review may prefer broad candidate retrieval followed by transparent filtering and human review.
Multilingual and domain-specific retrieval can behave very differently across embedding models. Benchmark claims from public datasets should not be treated as proof for a private corpus.
Reranking latency grows with the number and length of candidates. Candidate size should therefore be tuned as an accuracy/cost/latency variable rather than copied from a tutorial.
What would change this answer?
The component boundaries would not change if a vendor packages embedding generation, vector indexing and reranking behind one API. The product may hide the stages, but they remain conceptually different responsibilities with different failure modes.
Future embedding or retrieval models may reduce the need for separate reranking in some workloads, while stronger late-interaction or learned sparse methods can blur traditional dense/lexical categories. The architecture should still ask which stage produces representations, which stage generates candidates and which stage refines ranking.
The best design also changes with corpus size, query mix, language, domain terminology, update frequency, source authority, latency budget and evaluation results.
Related canonical knowledge
R01 assumes the basic RAG concept is already understood. RAG is the wider pattern in which retrieved external information is supplied to a model; embeddings, vector search and reranking are optional retrieval components inside that pattern.
When retrieval fails, diagnose source coverage, retrieval, ranking, context assembly and generation separately rather than treating the whole system as one “RAG failure.”
Source-of-Truth architecture is the authority layer around retrieval: it decides which source can establish a claim, while embeddings and ranking only decide which candidates appear relevant.
Frequently asked questions
Embeddings, vector databases and reranking
What is the difference between embeddings and a vector database?
What does a reranker do?
Does RAG require a vector database?
Why not use the reranker on the whole corpus?
Can reranking fix a missing document?
Is cosine similarity a relevance probability?
Should I use BM25 and vector search together?
When do I need a dedicated vector database?
Glossary
Key retrieval terms
- Embedding
- A numerical representation of content produced by an embedding model for similarity, clustering, retrieval or related tasks.
- Dense vector
- A vector representation in which many dimensions carry non-zero values, commonly used in semantic retrieval.
- Sparse vector
- A high-dimensional representation in which most dimensions are zero, often preserving stronger token- or term-like structure.
- Vector index
- A data structure that organizes vectors for efficient similarity or nearest-neighbor retrieval.
- Vector database
- A storage/search system designed to manage vectors, associated metadata and vector retrieval workloads.
- ANN
- Approximate nearest-neighbor search, which trades exact exhaustive comparison for faster retrieval at scale.
- HNSW
- Hierarchical Navigable Small World, a graph-based approximate nearest-neighbor indexing approach widely used for vector retrieval.
- BM25
- A lexical relevance-ranking method based on term occurrence and corpus statistics, widely used in full-text search.
- Hybrid search
- Retrieval that combines results or scores from multiple retrieval methods such as lexical and vector search.
- Reranking
- A later retrieval stage that re-scores and reorders an already generated candidate set.
- Bi-encoder
- An architecture that encodes query and candidate independently, enabling precomputation and scalable similarity search.
- Cross-encoder
- A model that jointly processes a query and candidate text, often improving relevance judgment at higher computational cost.
- Recall@k
- The fraction of relevant items recovered within the top k retrieved candidates.
- nDCG
- Normalized Discounted Cumulative Gain, a ranking metric that rewards relevant results appearing higher in an ordered list.
Conclusion
The clean retrieval model is simple: embeddings represent meaning, vector search retrieves candidates, and rerankers refine candidate ordering.
Once those boundaries are explicit, architecture decisions become easier to diagnose. Missing candidates point toward source coverage, chunking, embeddings, filters or first-stage retrieval. Poor ordering points toward ranking, fusion or reranking. Incorrect final answers can then be investigated separately at context and generation layers.
The most important result is not choosing the most fashionable retrieval component. It is building a retrieval pipeline whose stages, authority boundaries, metrics and failure modes can be measured independently.
Primary sources and implementation evidence
The external references below document the representation, vector-search and reranking mechanisms used in this article. Project-specific sections are original implementation evidence and are intentionally narrower than claims about complete production RAG maturity.
Sentence-BERT: Sentence Embeddings using Siamese BERT-NetworksFoundational paper demonstrating independently computable sentence embeddings for efficient semantic similarity search.
Qdrant — Architecture and data structure overviewOfficial documentation describing collections, points, vectors, payload metadata and HNSW-based similarity indexing.
Qdrant — SearchOfficial vector-search documentation covering similarity queries, filtering, exact versus approximate search and dense/sparse behavior.
Elastic — Vector searchCurrent documentation on dense/sparse vector retrieval, lexical/vector combinations and multi-stage search pipelines.
Elastic — Semantic rerankingCurrent guidance defining semantic reranking as a later-stage relevance operation over a smaller candidate set.
Cohere — Reranking with CohereCurrent documentation showing reranking as a second-stage improvement over lexical or semantic first-stage retrieval.
SQLite FTS5Official SQLite documentation for full-text search and the built-in BM25 ranking function used as lexical retrieval evidence.
Related Articles

What Is Context Engineering? What the Model Receives Before It Answers
Context engineering designs what information an AI model receives before inference, including prompts, retrieval, memory, application state, tool results and conversation history.

The Answer Validity Boundary: The Missing Layer Between Relevance and Reliable AI Answers
A source can be relevant, authoritative and still be wrong for the question being asked. The missing layer is applicability: the conditions under which an answer holds, and the changes that force it to be reconsidered. This article introduces the Answer Validity Boundary as a source-design pattern for humans, AI search and RAG systems.

MCP Explained: What It Connects, What It Does Not Do and Where It Fits
Model Context Protocol connects AI applications to external tools, resources and prompts through a standard client-server boundary. Learn what MCP does, what it does not do, and where it fits in agent architecture.

What Is RAG? The Simplest Explanation of How It Works
RAG sounds complicated, but the idea is simple: before an AI answers, it first looks up useful information from a knowledge source and gives that information to the language model. This guide explains RAG, LLMs, state, memory and tools using one simple mental model.

Front- and Backend Development
Front-end and back-end development is an essential part of web development and involves the creation of web applications and websites. Front-end development focuses on the user interface, while back-end development is responsible for programming and managing the server side.

When Should an AI Stop Trusting Its Own Knowledge? — The Retrieval Trigger
An AI model does not need retrieval for every question. The important problem is knowing when its internal knowledge is no longer enough. The Retrieval Trigger is a practical decision boundary that determines when an AI system should stop relying solely on model knowledge and obtain external evidence before answering.

Why More Context Can Make AI Answers Worse
A larger context window does not guarantee a better answer. This article explains how signal dilution, conflicting evidence, stale state, position sensitivity, and lossy compression can reduce AI reliability—and introduces a practical Context Pressure Test.

Agentic AI Explained: When an AI System Can Plan, Use Tools and Act
Agentic AI uses models inside multi-step execution loops where they can choose tools, observe results, update state and adapt their next action within explicit runtime and permission boundaries.

Source of Truth in AI Systems: Where Reliable Knowledge Actually Comes From
A Source of Truth defines which source is authoritative for a specific fact or state. Learn how it differs from RAG, provenance, memory, context, vector databases and systems of record.

Mastering the SEO Workflow: Essential Optimization Strategies for Organic Growth
A structured SEO workflow is crucial for sustainable organic growth. Learn the ten foundational strategies, from keyword research and technical optimization to content quality and performance analysis.

Where Does an LLM Get Its Data? RAG Data Sources in Python
An LLM does not magically know your files, databases or APIs. This practical continuation of the RAG series shows, with simple Python, how external data becomes retrievable evidence: from text files and SQL to full-text search, embeddings, context assembly and the final LLM call.

The GPU Is Not the Product: Future-Proof Private AI Architecture
Private AI infrastructure should not be designed around one GPU or one model. A more resilient approach combines fast inference GPUs, memory-rich AI systems, physical-AI nodes and optional frontier cloud models behind a capability-aware routing layer.