Hybrid RAG and reranking instead of naive vector search
Character-count chunking and vector-only search make hallucinations expected. What hybrid search and reranking change, and what the older RAG article taught wrongly.
Articles on cloud architecture, AI engineering and software development. From projects, not from a textbook.
Character-count chunking and vector-only search make hallucinations expected. What hybrid search and reranking change, and what the older RAG article taught wrongly.
A prompt or model change makes quality worse. Nobody notices if evaluation is missing. What I measure in regulated LLM projects.
Without observability nobody knows cost per request and nobody can reconstruct a bad answer. What production LLM systems must log.
A one-customer pilot is not a product for twenty. What tenant isolation means in RAG and agent systems.
When a model invents citations, legal stops the project. LLM traceability needs source grounding, not a better prompt.
A healthcare project with LangChain, OpenAI and Milvus. The pipeline was fast. As a recipe it is dangerous: fixed chunks, vectors only, no reranking.
A nationwide government application moved from J2EE onto Kubernetes. What I learned on that project.
An MCP server for a healthcare platform. What the Model Context Protocol changes in practice, and where it breaks.
Cookies & tracking
This website uses Google Ads conversion tracking to measure ad effectiveness. This involves setting cookies and transferring data to Google. Reach measurement with Matomo is cookieless. Details in the privacy policy.