Vector Databases for AI Apps: When You Actually Need One
A flat FAISS index is the right answer for a surprising number of RAG apps — and the wrong answer for a few specific ones. Where the boundary sits, what a flat index costs you, and the four signals that mean it's time to move.
Vector Databases for AI Apps: When You Actually Need One
There are two camps in the vector-database conversation and both are selling something. One says you can build RAG in twenty lines with a local index and never think about it again. The other says you need a managed vector database or your app will fall over. The truth is that a flat index is genuinely the right answer for a large class of apps, and genuinely the wrong answer for a smaller, specific class — and the useful skill is knowing which class you're in.
This article is about where that boundary actually sits, using a plain faiss.IndexFlatL2 scaffold as the reference point, because it's the simplest thing that works and therefore the clearest thing to reason about.
What a Flat Index Actually Is
A flat index stores every vector and, on a query, compares the query vector against all of them. No approximation, no clustering, no graph. It's exact nearest-neighbor search, and its cost is linear in the number of vectors.
The scaffold looks like this:
self.text_splitter = RecursiveCharacterTextSplitter(
chunk_size=1000,
chunk_overlap=200,
length_function=len,
separators=["\n\n", "\n", ". ", " ", ""],
)
self.index = faiss.IndexFlatL2(self.dimension) # 1536 for text-embedding-ada-002
Two details in there matter more than they look.
chunk_size=1000 is in characters, not tokens. That's roughly 250 tokens for English text. It's a reasonable default, but it's a character count, which means the same setting produces very different token counts for code, for German, or for CJK text. If your retrieval quality is inconsistent across content types, this is one of the first places to look.
chunk_overlap=200 is what keeps sentences from being cut in half. Without overlap, a chunk boundary can land in the middle of the sentence that answers the question, and neither chunk retrieves well. The overlap is cheap insurance and it's the setting people most often set to zero to "save space."
The metadata attached to each chunk is the other half of the design:
all_metadatas.append({
**meta,
"_chunk_index": j,
"_chunk_total": len(chunks),
"_source_doc_index": i,
})
_source_doc_index is what lets you answer "which document did this come from" — the difference between a RAG app that cites sources and one that just asserts things. If you're building anything where the user needs to verify the answer, that field is not optional.
The Sentence That Matters More Than the Vector Math
The most important line in a RAG pipeline is not in the retrieval code at all. It's in the system prompt:
You are a helpful assistant. Use the following context to answer the user's question.
If the context doesn't contain relevant information, say so and provide a general answer.
Without that second sentence, the model answers from its parametric memory whenever retrieval comes back empty or irrelevant — and it does so fluently, with no signal that the context didn't help. Your retrieval evals then measure nothing, because the model was never actually constrained by retrieval in the first place.
This is the single highest-leverage line in the whole pipeline, and it's the one most often missing from tutorials that focus on the vector math.
Where a Flat Index Is the Right Answer
A flat index is the correct choice when all of these hold:
- Single tenant. One index, one owner. No per-customer isolation to enforce.
- Under roughly fifty thousand chunks. At 1000 characters per chunk, that's on the order of fifty million characters of source material — a lot of documentation, a lot of internal knowledge base, more than most apps ever ingest.
- One process. A single service instance owns the index. No concurrent writers.
- Exact recall matters more than latency. Flat search is exact. Approximate indexes trade recall for speed, and if your use case is "find the one clause in the contract," exact is worth the milliseconds.
If that describes your app, a managed vector database is adding a network hop, a vendor, and a bill in exchange for capabilities you're not using. The flat index is not a toy — it's the right tool.
Where It Stops Being Enough
Four signals, in the order they usually show up:
1. Latency becomes the wall, not correctness. Flat search is O(n) per query. At tens of thousands of vectors it's still fast. At hundreds of thousands, p95 latency climbs and the answer is still correct — you're just waiting for it. This is the point where an approximate index (IVF, HNSW) starts to pay for itself, because you're trading a little recall for a lot of speed. Note the order: correctness first, then latency. Don't reach for an approximate index before you have a latency problem.
2. You need multi-tenancy. A single global index with no namespaces is fine for one owner and wrong for many. The moment two customers' documents share an index, isolation becomes a query discipline — every query has to filter by tenant, and every place that forgets is a leak. This bites before scale, not after: it's a correctness problem from the second tenant onward, not a performance problem at the hundred-thousandth vector.
3. You need more than one process. A flat index loaded from disk into one process is a single-writer design. Two instances mean two divergent copies, and whichever one wrote last wins. If you're running more than one replica, you need either a shared store or a coordination layer — and that's a different architecture, not a bigger index.
4. You need retrieval features the flat index doesn't have. A fixed k integer, no hybrid BM25+dense search, no re-ranking, no MMR for diversity. When retrieval quality plateaus and you've already tuned chunk size and overlap, these are the levers — and a flat index has none of them.
Deletion Is the Hidden Cost
One thing that's easy to miss until you need it: IndexFlatL2 doesn't support deletion. Removing a document means rebuilding the index without it — and rebuilding means re-embedding every remaining chunk, because the vectors aren't stored separately from the index.
# Rebuild index from remaining entries
self._create_new_index()
if remaining:
texts = [m["_system_text"] for m in remaining]
embeddings = await self._generate_embeddings(texts) # re-embed everything
self.index.add(embeddings_array)
For a small index this is fine. For a large one it's a full re-embedding bill every time a user deletes a document — which is a real cost, and a real latency spike, and a real reason to store embeddings alongside the index if deletion is a common operation in your app.
The Embedding Model Is Part of the Index
One constraint that isn't obvious from the API: similarity search only means anything if every vector in the index came from the same embedding model, at the same version. Mix two models in one index and the similarity scores between them are meaningless — the index will still return a number between 0 and 1, it just won't mean anything.
This matters for two practical reasons. First, upgrading your embedding model means re-embedding the entire corpus, not just new documents — old and new vectors can't coexist. Second, if you ever want per-tenant embedding models (say, a customer requires a locally-hosted model), you need separate indexes per model, because there's no way to mix them.
Conclusion
The vector-database decision isn't "flat vs. managed." It's a set of specific questions, and the answers tell you which side you're on:
- Single tenant, under ~50k chunks, one process, exact recall wanted? A flat index is correct, and a managed database is overhead.
- Multi-tenant? You have an isolation problem before you have a scale problem — solve that first, and it may push you toward namespaces or per-tenant collections regardless of size.
- Latency-bound at high vector counts? Now an approximate index earns its keep.
- Need hybrid search, re-ranking, or diversity? Those are features a flat index simply doesn't have.
The mistake in both directions is the same: choosing the architecture before you know which of these you actually are. Start flat, measure, and let the four signals above tell you when to move.
Further reading: FAISS – Guidelines to choose an index for the flat/IVF/HNSW trade-offs; Pinecone – Namespaces for what per-tenant isolation looks like in a managed store.