AI Architecture & Engineering
Deep dives into multi-tenancy, RAG systems, vector databases, and SaaS engineering.
AI Agent Platforms: Framework, Platform, or Just Build It Yourself
The three ways to get an agent into production — a framework you assemble, a platform you configure, or code you write — and the specific questions that decide which one fits. Including the one that matters most: who owns the failure when the agent does something wrong.
AI Agents in Production: What the ReAct Loop Doesn't Tell You
The ReAct pattern is a dozen lines in every tutorial. Here's what actually breaks once an agent runs against real tools, real users, and a real bill — loop limits, tool-result size, streaming intermediate steps, and why 'the agent decided to' is not an error message.
AI App Builders: What They Generate, and What They Leave to You
Scaffolding tools are good at the boring 80% and quiet about the 20% that decides whether your app survives contact with users. A look at what a generator actually produces, where the seams are, and the questions to ask before you commit to one.
AI Development Tools: The Integration Problem Nobody Warns You About
The hard part of building with AI isn't any single tool — it's making auth, providers, streaming, retrieval, and cost tracking agree on the same request. A look at the integration seams where AI development tools actually break.
Guardrails as Tenant Configuration: Beyond One-Size-Fits-All Security
Why globally hardcoded guardrails fail in multi-tenant AI products, how to build input and output checks as a configurable pipeline per tenant, and why guardrail decisions need an audit trail.
Multi-Tenancy for AI & RAG Systems: The Landscape Behind the Hype
Why classical SaaS multi-tenancy breaks at four new frontiers in RAG products, and why caching—semantically and technically—is the most easily-overlooked place for cross-tenant leaks.
Token Costs per Tenant: Why Rate Limiting Works Differently for LLM-SaaS
Why 'requests per minute' is the wrong metric for LLM products, how multi-provider pricing complicates cost tracking, and how the reservation pattern prevents tenants from busting their budget undetected.
Vector Databases for AI Apps: When You Actually Need One
A flat FAISS index is the right answer for a surprising number of RAG apps — and the wrong answer for a few specific ones. Where the boundary sits, what a flat index costs you, and the four signals that mean it's time to move.
Vector-Store Isolation in Multi-Tenant Systems: Where RAG Architectures Really Break
Why the classic tenant_id column doesn't work for vector stores, which isolation models exist (and what they mean concretely in Pinecone, Weaviate, Qdrant, or pgvector), and how to structurally prevent cross-tenant leaks in RAG retrieval instead of hoping no one forgets a filter.