There is a moment every developer hits when building with LLMs: the demo works brilliantly, then falls apart the moment you add real data, real conversations, or real complexity. The model starts forgetting things, hallucinating, going off-track, or simply failing to perform the task you know it's capable of.
You try a different model. You rewrite the prompt. You tweak the temperature up, then down. Nothing really sticks.
The diagnosis is almost always the same: a context problem.
Not a model problem. Not a prompt problem. A context problem, i.e., the model simply didn't have the right information at the right time.
This is what context engineering is about, and mastering it is the difference between a toy prototype and a system you'd trust in production.
What Is Context Engineering?
The term has been independently described by several influential figures, each adding a layer to the concept:
"The art of providing all the context for the task to be plausibly solvable by the LLM." — Tobi Lutke, CEO of Shopify
"The discipline of designing and building dynamic systems that provide the right information and tools, in the right format, at the right time." — Philipp Schmid, Hugging Face
Anthropic offers the most engineering-precise formulation: "the set of strategies for curating and maintaining the optimal set of tokens (information) during LLM inference."
What unites all these definitions is the same insight: the model's behavior at any given moment is entirely determined by what's in its context window. The weights provide potential; the context window determines what gets activated. If an LLM-powered system fails, the root cause is almost always that the model had the wrong information, too much noise, or critical context missing at inference time.
Why "Context Engineering" and Not "Prompt Engineering"?
Prompt engineering focuses on crafting effective instructions: clever phrasing, chain-of-thought, roles, and response format guidance. It's a valuable skill, but it addresses one dimension of a much larger problem.
Context engineering operates at a higher level of abstraction:

The shift matters because modern AI applications are not single-turn chatbots. They're agentic systems that operate across multiple inference calls, retrieve information from external sources, maintain memory across sessions, and use tools that return structured results back into the context. Managing all of that is an engineering problem, not a prompt-writing problem.
The Core Problem: The Context Window as a Finite Resource
A language model's context window is its working memory. Think of it as a desk: the model can only pay attention to what's on the desk right now. Everything not on it, for example past conversations, unretrieved documents, facts the model doesn't know - they simply don't exist at that moment.
And here's the most important technical principle in context engineering:
More context is not automatically better.
Transformers work by computing relationships between every pair of tokens simultaneously. With 1,000 tokens, that's ~1 million relationships. With 100,000 tokens, it's 10 billion. The more you pile on the desk, the harder it becomes for the model to keep track of what actually matters. Anthropic calls this phenomenon "context rot": even within a technically supported context length, response quality degrades progressively as the token count grows.
The "Lost in the Middle" paper (Liu et al., 2023) confirmed this empirically: "performance is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts."
The goal, then, is not to fill the context window - it's to curate it. As Anthropic puts it: "Find the smallest set of high-signal tokens that maximize the likelihood of your desired outcome."
This single principle drives every technique covered in this guide.
The Four Pillars of Context Engineering
A useful model comes from LangChain's framework, which organizes context engineering into four categories:
1. Write: Saving information outside the active window (scratchpads, memory files, databases)
2. Select: Retrieving relevant information into the window (RAG, memory retrieval, tool descriptions)
3. Compress: Reducing token consumption (summarization, trimming, compaction)
4. Isolate: Separating concerns across components (multi-agent systems, sandboxed environments)
Everything that follows falls into one of these categories.
1. Retrieval-Augmented Generation (RAG)
RAG is the most widely deployed context engineering technique. The core idea is simple: instead of trying to cram all relevant knowledge into the model's weights (impossible) or the system prompt (expensive and unwieldy), you keep that knowledge in an external store and retrieve only the relevant pieces when a question arrives.
It's the difference between trying to memorize an entire encyclopedia and simply knowing where to look things up.
This solves three problems at once:
- Hallucination — responses are grounded in actual documents you provided, not the model's guesses
- Stale knowledge — update the knowledge base without retraining anything
- Transparent reasoning — you can show exactly which source backed each answer
The Three RAG Paradigms
Naive RAG: Split documents into chunks, embed them, and when a question arrives, fetch the most similar chunks and inject them into the prompt. Simple, effective, and a perfectly reasonable place to start.
Advanced RAG: Adds steps before and after retrieval. Before: rewrite the query for better matching, or generate a "hypothetical ideal document" and use that as the query vector. After: rerank retrieved results, compress the less relevant ones. Noticeably better results on complex or ambiguous questions.
Modular RAG: For production systems — flexible, rearrangeable pipeline components (routers, iterative retrievers, fusion mechanisms) you compose based on your use case. What large enterprises reach for when Advanced RAG still isn't enough.
Retrieval Techniques
There are two fundamentally different families of retrieval, and they have completely different strengths.
Embedding-based (dense) retrieval converts text into numerical vectors — mathematical representations of meaning. When you search for "how to treat fever in infants," it finds documents about "high temperature in newborns" even with zero words in common. It understands intent.
BM25 (sparse) works completely differently — and it's worth understanding well, because it's still indispensable in 2025.
What BM25 is and why it still matters
BM25 stands for Best Match 25 — the 25th iteration of a retrieval algorithm family dating back to the 1970s. The core idea is straightforward: a document is relevant to a query if it contains the same words, especially if those words are rare across the rest of the collection.
Two factors drive the score:
- TF (Term Frequency): how many times the query word appears in the document. A document mentioning "PostgreSQL" ten times probably covers it more deeply than one that mentions it once. But BM25 caps the contribution — past a certain point, more repetitions add diminishing returns (unlike classic TF-IDF, which grew without bound).
- IDF (Inverse Document Frequency): words that appear everywhere are worth less. If 95% of your documents contain the word "system," it barely helps distinguish what's relevant. But "pg_trgm" likely appears in very few documents — if it shows up in both the query and a document, that's a strong relevance signal.
In practice, BM25 is an extremely well-calibrated keyword search. And that's a huge advantage in many real scenarios:

The bottom line: when the exact word matters more than the concept behind it, BM25 still beats embeddings. That's why the hybrid approach consistently outperforms either one alone — you need both perspectives.
Hybrid Search: Ensembling Dense and Sparse
In practice, you rarely want to choose between embeddings and BM25 - you want both. Hybrid search runs both retrieval methods in parallel on the same query, then fuses their ranked results into a single list.
The most common fusion strategy is Reciprocal Rank Fusion (RRF): each result gets a score based on its position in each list (1/rank), and the scores are summed across lists. A document that ranks #2 in semantic search and #5 in BM25 scores higher than one that ranks #1 in only one of them. This simple heuristic is surprisingly robust, it requires no learned weights and works across any pair of rankers.
Other fusion approaches include:
- Weighted linear combination: normalize scores from each retriever to [0,1], then combine with tunable weights (e.g., 0.7 × dense + 0.3 × sparse). Lets you bias toward whichever method works better for your domain.
- Cross-encoder reranking: retrieve a broad candidate set from both methods, then pass the top-N candidates through a cross-encoder model that scores each (query, document) pair jointly. More expensive but significantly more accurate.
Most vector databases (Pinecone, Weaviate, Qdrant, Milvus) and search platforms (Elasticsearch, OpenSearch) now support hybrid search natively. The typical implementation:
- Index documents with both dense vectors and a sparse (keyword) index
- At query time, run both retrievers
- Fuse results using RRF or weighted combination
- Optionally rerank the fused list with a cross-encoder
Anthropic's Contextual Retrieval experiments confirmed this: adding BM25 to contextual embeddings reduced retrieval failures by an additional 14 percentage points beyond embeddings alone. Hybrid search is the default recommendation for any production RAG system.
Contextual Retrieval
Anthropic published a significant improvement to standard RAG in 2024. The problem: when documents are chunked, individual chunks lose their broader document context. A chunk that says "The revenue declined 10% in Q3" has no meaning without knowing which company and which year.
The fix: before embedding each chunk, prepend a short (50-100 token) LLM-generated summary explaining where it fits within the document:
Results from Anthropic's experiments:
- Contextual Embeddings alone: 35% reduction in retrieval failures
- Combined with BM25 hybrid: 49% reduction
- Adding a reranker on top: 67% reduction
The cost with Claude and prompt caching: approximately $1.02 per million document tokens.
Chunking Strategy
The chunk size is a hyperparameter worth tuning:
- Small chunks (128–256 tokens): better retrieval precision, less context per chunk
- Large chunks (512–1024 tokens): more context, noisier retrieval
- Hierarchical indexing: small chunks for retrieval, their parent chunks for generation. Best of both worlds.
A widely validated rule: measure retrieval quality (e.g., recall@k) with different chunk sizes on your actual data. Grid-search this before optimizing anything else.
2. Memory Systems
Imagine you hired a brilliant assistant. On Monday they help you work through a complex architecture decision. On Tuesday, when you pick up where you left off, they remember nothing. You have to re-explain everything from scratch.
That's how most AI agents behave today without a well-designed memory system. For agents operating across multiple turns or sessions, memory isn't a nice-to-have, it's the central context engineering problem.
The Four Memory Types in Practice
Episodic memory: Records of past experiences and interactions. "Last time I helped this user, they preferred concise answers and used Python 3.11." Stored as key-value pairs or vector embeddings, retrieved by similarity to current task.
Semantic memory: Facts about the world or domain. Company knowledge base, product catalog, technical documentation. This is standard RAG territory.
Procedural memory: How to do things. System prompts, CLAUDE.md files, instructions that define agent behavior. Updated rarely, referenced constantly.
Working memory: What the agent is thinking about right now. The active context window, scratch notes, current task state.
Memory Storage Patterns
LangChain's memory types illustrate the trade-offs:

For most production agents, the summary + buffer hybrid is the right default.
File-Based Persistent Memory
Anthropic's Claude uses a simple but powerful pattern: a /memories directory where the agent reads memory at session start, writes updates mid-session, and saves a summary before ending. This mirrors how Claude Code uses CLAUDE.md files.
The pattern:
This is procedural and episodic memory combined — lightweight, human-readable, and easy to edit manually when needed.
3. System Prompt Design
The system prompt is the foundation of everything. It's the text the model reads before any user interaction — establishing who it is, what it can do, how it should behave, and what format to respond in. Because it's present in every inference call, a well-crafted system prompt multiplies the impact of everything that comes after it.
Design Principles
Right altitude: Aim for the right level of specificity. Too prescriptive → brittle behavior when edge cases arise. Too vague → the model fills in the gaps in unpredictable ways.
Motivation over commands: Anthropic's prompt engineering documentation explicitly recommends providing motivation behind instructions, giving this example: instead of "NEVER use ellipses," write "Your response will be read aloud by a text-to-speech engine, so never use ellipses since the text-to-speech engine will not know how to pronounce them." Their reasoning: "Claude is smart enough to generalize from the explanation." The model can then handle novel situations the instruction writer didn't anticipate - because it understands the constraint's purpose, not just its letter.
XML structure for complex prompts: Claude and other models parse XML tags reliably. For prompts with multiple distinct sections:
The golden rule: Show your system prompt to a colleague without explaining it. If they'd be confused about what the agent is supposed to do, the model will be too.
Long-Context Prompting Patterns
The "lost in the middle" finding has direct implications for how you structure prompts with large documents:
1. Put longform data above instructions — the model attends to instructions more reliably when they come after the content
2. Put the query at the end — can improve response quality by up to 30% on complex multi-document tasks
3. Use document tags — wrap documents consistently:


