Why This Piece Exists
This is the opening piece in SelfMinds AI's Living Intelligence series — a set of long-form technical reviews tracing how AI systems progress from a single model call to a fully self-improving platform. Each piece exists to document, with full citation, how a specific layer of that stack is evolving in the published literature and in production practice. This one is not a description of a SelfMinds AI product; it is a synthesis of where the research on context engineering currently stands, written because the term entered wide use in 2025 faster than the underlying literature got organized, and a rigorous map of it is genuinely useful — to us as engineers building on top of it, and hopefully to anyone else doing the same.
The 2025 Survey of Context Engineering for Large Language Models by Mei et al. is a useful anchor for that claim: it reviews more than 1,400 papers in an attempt to unify research areas — retrieval, in-context learning, memory, long-context inference, compression, tool interaction, planning, state management, and agent orchestration — that had been developing largely independently for several years under other names. What follows draws on that survey, on the foundational papers it and related literature trace back to, and on practitioner sources from model developers and applied teams, organized into a single working framework with formal notation where the underlying claims support it.
1. Defining Context Engineering
Context engineering can be defined, more precisely than the phrase is often used, as: the design and optimization of the inference-time information state through which a foundation model perceives a task, its environment, its history, available knowledge, available actions, and intermediate computational state.
This is substantially broader than prompt engineering. A prompt is one component of context. The full context an agent sees at inference time is assembled, not authored, and may contain:
- system policy and task instruction
- user state and conversation history
- retrieved knowledge and memories
- tool schemas and tool observations
- plans and intermediate results
- few-shot examples
- environment state and multimodal observations
- output constraints
Context engineering is therefore fundamentally an information-allocation problem under bounded computational resources. It can be stated as an optimization:
where It is all information potentially available at inference step t, Ct is the subset or representation actually supplied to the model, M is the model, Tt is the current task, and U measures downstream utility. Equation 1 is a framing device, not a metric anyone computes directly — its value is in making explicit that every context-construction decision trades utility against cost, noise, and risk, rather than simply asking "does more context help."
Liu et al.'s Lost in the Middle (TACL, 2024) is the paper most directly responsible for this insight becoming common knowledge: it demonstrated that models do not use all positions in a long input with equal reliability, and that accuracy on a task can degrade measurably depending on where in the context the relevant information sits — even when that information is present and the model's stated context limit is nowhere near exceeded.
1.1 From Prompt Engineering to Context Engineering
Prompt engineering asks "how should I phrase the instruction?" Context engineering asks a different question: what should the model know at this particular inference step, where should that information come from, how should it be represented, and what should be excluded? That reframing produces several genuinely different engineering problems, summarized in Table 1.
| Prompt Engineering | Context Engineering |
|---|---|
| Write instructions | Construct inference state |
| Mostly static | Frequently dynamic |
| Human-authored | Machine-assembled |
| Single model call | Multi-step lifecycle |
| Optimize wording | Optimize information |
| Prompt template | Context pipeline |
| Few-shot examples | Retrieval policy |
| Context is input | Context is managed state |
| Usually stateless | Persistent state possible |
| Tokens treated uniformly | Tokens assigned different utility |
The 2025 context-engineering survey organizes the field around three broad phases — context retrieval and generation, context processing, and context management — then studies how those integrate into RAG, memory, tool-based reasoning, and multi-agent systems. A finer-grained, implementation-oriented decomposition, used as the organizing spine for the rest of this piece, is:
Source → Retrieve → Rank → Transform → Compress → Structure → Place → Consume → Evaluate → Persist/Forget
The last two stages are easy to skip and are where most production systems leave value on the table: context engineering should not end when the prompt is assembled.
2. The Seven Layers, and Where Context Comes From
A useful way to organize the field is as seven layers: sources, retrieval, transformation, compression, memory, tools/multi-agent acquisition, and lifecycle management. This piece works through each. The first — sources — is deceptively simple: potential context can come from the current user request, conversation history, system or developer instructions, documents, databases, knowledge graphs, vector indexes, web search, APIs, environment observations, multimodal input, execution traces, user profiles, episodic memory, semantic memory, procedural memory, other agents, previous model outputs, code repositories, and runtime state.
That list immediately differentiates context engineering from RAG. RAG is one context-acquisition mechanism among many, not context engineering itself — a distinction worth stating plainly because the two terms are frequently used interchangeably in practice. Lewis et al.'s original RAG formulation (NeurIPS 2020) explicitly combined a model's parametric knowledge with external non-parametric memory retrieved at inference time — a specific architectural pattern, not a synonym for "giving the model outside information" in general.
3. Context Retrieval: Beyond Top-k Similarity
Retrieval, in its narrowest form, asks:
But modern context engineering requires substantially more than top-k semantic similarity over a query q against a corpus K. A well-designed system determines whether retrieval is necessary at all, what should be searched, which data source to search, how many results are required, how diverse the evidence should be, whether results conflict with each other, and whether another retrieval iteration is warranted before generation proceeds.
The conventional architecture — query → embedding → similarity search → top-k chunks → LLM — works surprisingly well in practice, but it introduces a specific, well-documented failure mode: wrong query leads to wrong documents leads to an apparently plausible answer built on the wrong evidence. The retriever can become the hidden bottleneck of an otherwise capable model, silently, because a fluent wrong answer gives no external signal that retrieval — not generation — is where the system failed.
3.1 HyDE, Self-RAG, and Corrective RAG
HyDE (Gao et al., ACL 2023) changes retrieval from embedding the question directly to first generating a hypothetical answer document, embedding that, and using it to find real documents — the generated hypothetical acts as a semantic bridge into the document space, often closer in embedding space to a real answer than the bare question is.
Self-RAG (Asai et al., ICLR 2024) introduced a concept that is now foundational to context engineering more broadly: retrieval should itself be conditional. The model is trained to decide whether retrieval is required at all, and to critique both the retrieved evidence and its own generation, rather than blindly retrieving a fixed number of passages for every query regardless of whether the query needs external evidence.
Corrective RAG (CRAG) (Yan et al., ICLR 2025) goes a step further by explicitly evaluating retrieval quality and activating corrective behavior when the retrieved evidence appears weak:
The branching structure in Equation 3 is closer to what "context engineering" means in practice than simple RAG is: retrieval as a conditional, self-correcting subroutine rather than a fixed lookup step.
3.2 Hierarchical Context and RAPTOR
Flat chunk retrieval often destroys document structure. A 200-page report has a real hierarchy — document, chapter, section, subsection, paragraph, sentence — and flattening it into independent, uniformly-sized chunks loses much of that structure. RAPTOR (Sarthi et al., ICLR 2024) addresses part of this by recursively clustering and summarizing chunks into a tree that holds information at multiple levels of abstraction simultaneously, so a query can be answered from a document-level summary, a section-level synthesis, or an exact paragraph, depending on what it actually needs. Context, in other words, does not have to have a single resolution — context engineering increasingly resembles multiresolution information processing rather than flat lookup.
4. Context Transformation, Compression, and the Context-Window Fallacy
Retrieved information usually should not enter the model's context unmodified. The path is better thought of as raw information → relevant information → model context, with a transformation layer in between doing extraction, deduplication, summarization, canonicalization, entity resolution, contradiction detection, reranking, metadata augmentation, chunk merging, query-focused summarization, compression, structured conversion, and graph expansion as needed. Many production RAG systems still effectively do retriever → concatenate documents → model; a stronger architecture is retrieve → filter → rerank → reconcile → compress → structure → model. This middle layer is underappreciated relative to how much failure-mode reduction it buys.
Context also has a real inference cost: for Transformer models, longer input sequences affect memory, latency, and attention computation. That makes a central context-engineering question genuinely quantitative: what is the minimum representation that preserves task-relevant information? LLMLingua (Jiang et al., EMNLP 2023) frames prompt compression explicitly as a method for reducing inference cost while preserving useful semantic content, reporting substantial compression ratios under the conditions it tested. LongLLMLingua (Jiang et al., ACL 2024) extends the same approach to long-context settings, explicitly modeling both information density and positional effects together.
More context ≠ better context.
It is sometimes true that Accuracy(Ccompressed) > Accuracy(Craw) — because compression, done well, increases the signal-to-noise ratio the model has to reason over.
4.1 The Context-Window Fallacy
One of the more important conclusions in this literature is that maximum context length and effective context length are different quantities. A model advertised as accepting hundreds of thousands or millions of tokens does not necessarily reason equally well over every one of them — this is exactly what Lost in the Middle demonstrated with positional sensitivity. That motivates a concept of context utilization (CU):
A model with a 1M-token window and poor utilization can underperform a model with a 32K-token window supplied with carefully selected evidence — which is why the metric worth optimizing is not raw token count but effective information density:
5. Context Placement as a Layout Problem
Information ordering itself can influence performance, which means assembling context is partly a layout problem, not only a selection problem. A context ordered SYSTEM → TASK → GLOBAL CONSTRAINTS → RELEVANT MEMORY → RETRIEVED EVIDENCE → CURRENT OBSERVATION → IMMEDIATE QUESTION may behave differently from the same components concatenated in a different or random order. That creates a largely unexplored optimization problem:
where π is a permutation or structured arrangement of the available context components. The long-context positional effects documented by Lost in the Middle make this more than a stylistic formatting preference — it is a genuine, currently under-studied lever on model performance.
6. Memory Is Context Across Time
Memory becomes relevant the moment context engineering spans more than one interaction. A useful decomposition distinguishes five kinds: working memory (information actively required by the current reasoning episode), episodic memory (specific historical experiences — "the user requested architecture A," "the experiment failed because of X," "we changed parameter Y and the result improved by Z"), semantic memory (generalized knowledge extracted from experience — "for this codebase, component X should never directly call Y"), procedural memory (how to perform recurring operations — "when deploying: run tests → build → validate → deploy"), and profile memory (stable facts about users, organizations, environments, or projects). Each of these plausibly needs a different retention policy, retrieval algorithm, storage representation, decay rate, and confidence estimate — treating "memory" as one undifferentiated store tends to blur those distinctions in practice.
6.1 MemGPT and Virtual Context
MemGPT (Packer et al., 2023 — a research preprint from the Letta/MemGPT project, not a peer-reviewed venue) introduced one of the clearest conceptual analogies available for this problem: LLM context can be managed somewhat like memory in an operating system. Instead of putting everything into a limited context window, the architecture manages memory tiers and decides, dynamically, what should occupy the active context at any given moment. The analogy extends further than the original paper spells out, as shown in Table 2.
| Operating System | Context-Engineered Agent |
|---|---|
| RAM | Context window |
| Disk | Persistent memory |
| Cache | Working memory |
| Page retrieval | Memory retrieval |
| Eviction | Forgetting |
| Scheduler | Context controller |
| Process state | Agent state |
| Filesystem | Knowledge store |
This is a stronger conceptual foundation for long-running agents than the common alternative of simply appending every new conversation turn to a growing context indefinitely — a pattern that runs directly into the context-window fallacy discussed above.
6.2 Reflection as Context Generation
Not all useful context exists externally — a system can generate new context from context it already has. Generative Agents (Park et al., UIST 2023) demonstrated this through a loop of observation, memory retrieval, reflection, and planning. The distinction between external context (obtained from outside the model — a document, a tool result) and derived context (produced by processing existing information — a summary, a plan, a hypothesis, a lesson learned, a synthetic query) is a useful one to keep explicit in system design, since the two carry different provenance and reliability characteristics (see §10). Formally:
where Ot is an observation, At an action, and Rt an outcome at step t. Under this formulation, context is a dynamical state that evolves with the system, not a static prompt assembled once.
7. Tools, Coding Agents, and Knowledge Graphs as Acquisition Mechanisms
Tool use changes a model from a passive text generator into an information-seeking process. ReAct (Yao et al., ICLR 2023) established the now-standard reason → act → observe → reason pattern, where actions obtain information and observations become new context for the next reasoning step. Toolformer (Schick et al., NeurIPS 2023) approached the same problem from the training side: teaching a model which APIs to call, when to call them, how to parameterize them, and how to incorporate the returned result. Under this view, tools are context operators — a search call injects knowledge context, a calculator injects deterministic computation, a database query injects structured state, a vision call injects perceptual information, and a filesystem read injects implementation context. The agent does not merely "use tools"; it uses tools to alter its own epistemic state.
7.1 Context Engineering for Coding Agents
This becomes particularly concrete for large codebases, where the naïve strategy — read the repository, put a large amount of code into context — is inefficient and often infeasible. A more scalable strategy is progressive: repository map → module map → class signatures → function signatures → selected implementation → dependent tests, expanding only along task-relevant branches:
The agent initially needs structural context — what exists and how it's organized — not full implementation detail, expanding deeper only when a specific task branch requires it. This minimizes token expenditure while preserving navigational awareness of the whole codebase.
7.2 Knowledge Graphs as a Context Channel
Knowledge graphs provide a complementary retrieval channel to vector search. Vector retrieval answers, approximately, "what text looks semantically similar?" Graph retrieval answers a different question: "what entities and relationships are structurally connected to this query?" A sophisticated context system can draw on both, plus keyword search, memory, and tool results, as a unified candidate pool before a selection stage narrows it down:
This hybrid-retrieval framing is particularly attractive for complex, multi-hop reasoning tasks where a single retrieval channel systematically misses one type of relevant evidence. This series' Knowledge Fusion Engineering and Graph Engineering pieces cover this territory — vector versus graph retrieval, hop limits, provenance through the graph — in dedicated depth.
8. Multimodal and Multi-Agent Context Engineering
The field becomes considerably more interesting once models are multimodal. Context can consist of text combined with several other modalities and derived signal types at once:
Simply converting everything into text can discard important structure. For video specifically, the context problem is not "can my LLM accept a long transcript?" It is: which moments, frames, speakers, visual entities, OCR strings, actions, temporal relationships, and retrieved prior events should actually be available for answering this particular question? That requires modality-aware context selection, not uniform flattening to text — a theme this series' Beyond OCR piece develops in more depth for the document-specific case.
8.1 Multi-Agent Context and Isolation
Multi-agent systems introduce a distinct challenge: the context universe becomes a union of private, shared, message, and environment context —
— and each agent should probably not receive everything known by every other agent. Context can instead be partitioned by role, authority, relevance, privacy, task stage, and expertise, giving a per-agent selection function:
This creates an opportunity for context routing analogous to message routing in distributed systems. An underappreciated design principle follows directly from it: information one subsystem knows does not automatically need to become context for every other subsystem. Deliberate context isolation buys lower token usage, reduced distraction, lower information leakage, easier debugging, and a smaller prompt-injection surface. In a typical research-agent pattern, a writer agent does not need the researcher's complete reasoning trajectory — it may only need validated findings, with the researcher's raw exploration kept private. This series' Graph Engineering piece — on coordinating multiple loops — treats multi-agent context routing as one of its central topics.
9. Provenance, Freshness, Conflict, and Security
Four related problems become important once context is assembled from many sources rather than authored by a human: knowing where it came from, how long it stays valid, what to do when sources disagree, and what to do when a source is actively adversarial.
9.1 Context Provenance
Production context systems should store more than the bare content of a context unit. Storing only {"text": "X happened."} loses information a system needs to reason about reliability later. A richer record looks closer to:
Context without provenance is a real risk, not a theoretical one, because generated summaries and retrieved facts otherwise become indistinguishable in the model's input once concatenated as plain text — the model has no way to weight a fabricated-sounding but actually well-sourced claim differently from a fluent hallucination unless the system tracks that distinction upstream.
9.2 Context Freshness
Context relevance changes over time, and at very different rates depending on what the content is. For some information, value decays roughly exponentially:
with λ varying enormously by content type, as shown in Table 3.
| Context | Decay Rate |
|---|---|
| Mathematical theorem | Almost zero |
| User preference | Slow |
| Company org chart | Moderate |
| Stock price | Extremely fast |
| Current webpage | Potentially fast |
| Task execution state | Immediate |
A strong context system needs explicit freshness semantics per content type rather than one global expiry policy.
9.3 Context Conflict
Contradictory context is a largely underdeveloped problem in production systems. If retrieved context contains both "Product price = $99" and "Product price = $129," naïve RAG may simply expose both claims to the model and let it guess. A context-engineered system should instead determine source reliability, temporal ordering, whether both claims can genuinely be true simultaneously, and whether one supersedes the other.
Context Selection ≠ Similarity Search.
It includes epistemic reconciliation — deciding which of two conflicting, similarly-retrievable claims is actually true right now.
9.4 Context Poisoning and Security
As tool-using agents ingest arbitrary external content, context becomes an attack surface. Potential attacks include direct and indirect prompt injection, poisoned retrieved documents, malicious tool descriptions, corrupted memory, and instruction/data confusion. Every context unit should carry both a semantic role (what this information means) and an authority level (what this information is permitted to influence) — which argues for keeping instruction, observation, evidence, untrusted content, memory, and tool result as distinct, taggable categories rather than concatenating everything into undifferentiated text where a retrieved document can silently masquerade as a system instruction.
10. Context Caching and the Context Lifecycle
Many inference requests share context — system instructions, a repository map, API documentation, a company ontology often remain unchanged across many calls. That intersects context engineering directly with prefix caching, KV-cache reuse, semantic caching, retrieval caching, and prompt caching as system-level mechanisms. StreamingLLM (Xiao et al., ICLR 2024) illustrates the model-level counterpart of the same problem: managing attention state efficiently across very long streaming interactions. Infini-attention (Munkhdalai et al., 2024 — an arXiv preprint from Google Research, not yet peer-reviewed at time of writing) explores compressive memory built directly into the attention mechanism itself, rather than managed externally.
Pulling the pipeline stages introduced throughout this piece together, context is best understood as passing through a full lifecycle, not a one-shot assembly step:
This lifecycle view — not "prompt building," a one-shot metaphor — is the more useful mental model for engineering a context system: every stage after INJECT still has work to do, and the branch at the bottom, back into memory or into deletion, is what makes the system capable of learning across calls rather than starting from zero each time.
11. Measuring Context Quality
Context quality is worth evaluating along at least ten independent dimensions, since a system can score well on one and poorly on another — high recall with poor signal-to-noise is a common and specific failure pattern, not a contradiction:
where R is Relevance, C is Completeness, A is Accuracy, F is Freshness, D is Diversity, P is Provenance, S is Signal-to-noise, T is Temporal consistency, U is Utility, and K is Cost. Several of these dimensions have narrower, more classically IR-flavored formalizations worth stating explicitly. Context Recall asks how much of the required information actually made it in:
Context Precision asks the complementary question — how much of what made it in was actually useful:
Most naïve systems optimize recall alone, by adding more context — but a high-performing system needs both together. A third metric distinguishes a failure of retrieval from a failure of the model to use what it was given, which the earlier Lost in the Middle discussion (§4.1) shows is not the same failure at all:
And two efficiency measures turn context construction into an explicit resource-allocation decision rather than an unconstrained "add more" default:
12. Context Rot, Compilation, and Model-Awareness
As conversations or agent traces grow, context accumulates outdated assumptions, duplicate facts, irrelevant intermediate reasoning, superseded plans, contradictory state, and tool noise. Performance can deteriorate even though the model technically still has capacity available — a system-level phenomenon worth naming explicitly as context entropy accumulation. Good context engineering performs active entropy reduction through summarization, consolidation, deletion, canonicalization, supersession, and deduplication, rather than letting context grow monotonically.
A second useful analogy treats context engineering like compilation: raw information → parsing → intermediate representation → optimization → dead-context elimination → compression → context scheduling → model-specific serialization → inference. Under this framing, raw information plays the role source code plays in a compiler, structured knowledge plays the role of an intermediate representation, and the assembled prompt/context plays the role of an executable representation — and just as different compilation targets need different code generation, different target models plausibly should not receive identical context. That motivates a model-aware framing of the core optimization:
since different models genuinely differ in effective context length, instruction-following behavior, positional biases, tool ability, domain knowledge, and reasoning strength — an optimal context compiler, in this framing, has to be model-aware rather than task-only.
13. Adaptive and Control-Theoretic Context Engineering
Static context pipelines typically fix parameters like chunk size, top-k, and summary length in advance. A better system determines these dynamically: a simple factual question might warrant retrieving one or two facts; a comparative question might warrant multiple independent sources; long-form synthesis might warrant a full retrieve → cluster → summarize → recursively retrieve loop; uncertain reasoning might warrant retrieve → reason → identify missing evidence → retrieve again. Context construction, under this view, becomes a policy:
where atC is a context-management action taken given the current state st. This creates a genuine connection to control theory: at every inference step, a system observes its current state, estimates an information deficit, acquires context, reasons, observes the outcome, and updates state —
— which makes context engineering, at this level of maturity, a closed-loop controller rather than a one-shot template fill. That is considerably closer to agent architecture than to prompting.
13.1 Agentic Context Engineering
This direction became explicit with Agentic Context Engineering (ACE) (Zhang et al., ICLR 2026), which treats context as an evolving playbook that accumulates and refines strategies through generation, reflection, and curation, rather than remaining a static, hand-authored prompt. Microsoft Research's own summary of the same work (see References — company R&D page, not a substitute for the ICLR paper itself) and independent coverage in VentureBeat both frame it around a specific failure mode the paper names "context collapse." That work marks a useful staged framing for the field's own trajectory:
Context Engineering 1.0 — humans design the context pipeline.
Context Engineering 2.0 — agents participate in managing their own context.
Context Engineering 3.0 — the context-management policy itself learns from outcomes.
That last stage begins to look like self-improving inference infrastructure — the territory this series' later Evolution Engineering piece is aimed at.
14. What Context Engineering Is Not, and a Reference Architecture
Context engineering should not become a fashionable replacement name for every LLM technique that touches information flow. It is not synonymous with prompt engineering, RAG, vector databases, memory, agents, long-context models, or knowledge graphs individually — each of those is a component or mechanism within it. A more precise statement:
Pulling every mechanism discussed in this piece together into one architecture gives something close to what a genuine context engine looks like in production:
15. Open Research Gaps
Several problems in this literature remain genuinely open rather than merely under-optimized. Learned context selection would replace handcrafted top-k rules with a learned P (ci | task, state, model) based on eventual task success, rather than a fixed heuristic. Counterfactual context evaluation would ask, for each included item, whether removing it would change the answer — enabling causal attribution of context utility rather than correlational guesses. Context attribution would determine which specific context units actually influenced each generated claim, which matters directly for auditability in regulated domains. Multi-resolution memory would maintain raw episodes, compressed episodes, semantic abstractions, and procedural knowledge simultaneously rather than collapsing to one representation. Context uncertainty would store confidence not only in model outputs but in the context itself, propagating uncertainty rather than discarding it at retrieval time. A context graph would represent relationships between facts, evidence, agents, tasks, decisions, and outcomes explicitly, instead of keeping a flat memory store. Context garbage collection — agents need genuine, principled forgetting mechanisms, not just growing stores with occasional manual pruning. Security-aware context routing would ensure untrusted information does not inherit authority merely by having entered the context window (see §9.4). Cross-agent context optimization would determine, formally, what information one agent should transmit to another rather than relying on ad hoc message design. And multimodal context compression — compressing video or multimodal history without destroying cross-modal relationships — remains substantially harder than text summarization and is comparatively under-studied relative to text compression.
16. A Research Hierarchy, and Relationship to Adjacent Disciplines
The field organizes reasonably cleanly into three top-level concerns — acquisition, processing, and management — feeding a common assembly and evaluation loop:
Context engineering also relates to, but is distinct from, several adjacent named disciplines: prompt engineering optimizes instructions; RAG engineering optimizes external knowledge retrieval specifically; memory engineering optimizes information persistence across time; tool engineering optimizes external actions and observations; agent engineering optimizes decision loops; harness engineering optimizes the execution infrastructure surrounding the model; and context engineering, in this framing, optimizes what information reaches inference, and when. These areas overlap substantially in practice but operate at genuinely different levels of abstraction — which is also why this series treats several of them as separate, dedicated pieces rather than folding everything into one.
17. The Core Principle, and Where the Field Is Going
Across this literature, the operating principle worth keeping is stated plainly at the top of this piece: the objective is not to give the model more information — it is to give the model the right information, in the right representation, at the right time, and to remove it once it stops being useful. Figure 4 shows how SelfMinds AI organizes this principle into a concrete stack.
The trajectory this body of work points toward is a progression from prompt engineering, through RAG, through memory, to context engineering, and — per the ACE line of work discussed in §13.1 — toward context policy learning. In that end state, a system would not have a developer hand-specifying top_k = 5, history = last_10_messages, include_summary = True; instead, a context controller would learn, from outcomes, what to acquire, retain, transform, retrieve, expose, hide, compress, or forget before the next inference step, given the current model, task, history, token budget, uncertainty, and available information. At that point, context engineering stops being infrastructure that merely surrounds an LLM and becomes an adaptive information-management layer that materially governs the model's effective cognition — which is also the reason two systems built on exactly the same underlying model can differ dramatically in reliability, cost, memory, and autonomy: the difference is their context architecture, not their model weights.
18. Limitations of This Synthesis
In the interest of the rigor this piece is trying to model, three limitations are worth stating explicitly rather than leaving implicit. First, several of the formalizations above — the ten context-quality dimensions in Equation 14, the efficiency metrics in Equations 18–19, "context entropy accumulation," and the reference architecture in §14 — are proposed synthesis frameworks, not results independently validated in a peer-reviewed benchmark. They are included because they are useful ways to reason about the problem, and each is flagged in-line where it appears; they should not be cited as established metrics with agreed measurement protocols. Second, evidentiary weight varies substantially across the works cited: peer-reviewed venues (NeurIPS, ICLR, ACL, EMNLP, and similar) carry more weight than arXiv preprints, which in turn carry more weight than a company blog post or news article — status is labeled explicitly for every entry in the References section below specifically so this distinction stays visible rather than getting flattened by uniform-looking citation links. Third, this is a synthesis of a fast-moving literature as of mid-2026; several of the papers cited here (ACE, several long-context and compression preprints) are recent enough that independent replication and follow-up critique are still ongoing, and some of the field's own terminology — "context engineering" as an umbrella term — is itself less than two years old at time of writing.
References
Peer-reviewed venue, company/institutional research publication, or unreviewed preprint — these carry meaningfully different evidentiary weight, so each entry is labeled. Links go directly to the primary source. No content beyond what was supplied for this piece's drafting was introduced during writing — this list matches the source bibliography in full.
Context Engineering & Surveys
- Mei et al. (2025) — A Survey of Context Engineering for Large Language ModelsarXiv preprint
- Gao et al. (2023) — Retrieval-Augmented Generation for LLMs: A SurveyarXiv preprint
- Zhang et al. (2025) — Agentic Context Engineering (ACE)ICLR 2026, peer-reviewed
- Hu et al. (2025) — Memory in the Age of AI AgentsarXiv preprint
- Singh et al. (2025) — Agentic RAG: A SurveyarXiv preprint
RAG & Retrieval Foundations
- Lewis et al. (2020) — Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksNeurIPS 2020, peer-reviewed
- Karpukhin et al. (2020) — Dense Passage RetrievalEMNLP 2020, peer-reviewed
- Izacard & Grave (2021) — Fusion-in-DecoderEACL 2021, peer-reviewed
- Guu et al. (2020) — REALMICML 2020, peer-reviewed
- Borgeaud et al. (2022) — RETROICML 2022, peer-reviewed
- Izacard et al. (2022) — AtlasJMLR 24(251), peer-reviewed
- Gao et al. (2023) — HyDEACL 2023, peer-reviewed
- Asai et al. (2024) — Self-RAGICLR 2024, peer-reviewed
- Yan et al. (2024) — Corrective RAG (CRAG)ICLR 2025, peer-reviewed
- Sarthi et al. (2024) — RAPTORICLR 2024, peer-reviewed
- Yang et al. (2024) — CRAG BenchmarkNeurIPS 2024 D&B Track, peer-reviewed
- Khattab & Zaharia (2020) — ColBERTSIGIR 2020, peer-reviewed
- Nogueira & Cho (2019) — Passage Re-ranking with BERTarXiv preprint
- Thakur et al. (2021) — BEIRNeurIPS 2021 D&B Track, peer-reviewed
- Xiong et al. (2020) — ANCEICLR 2021, peer-reviewed
In-Context Learning & Prompting
- Brown et al. (2020) — Language Models are Few-Shot LearnersNeurIPS 2020, peer-reviewed
- Wei et al. (2022) — Chain-of-Thought PromptingNeurIPS 2022, peer-reviewed
- Kojima et al. (2022) — Zero-Shot ReasonersNeurIPS 2022, peer-reviewed
- Wang et al. (2023) — Self-ConsistencyICLR 2023, peer-reviewed
- Min et al. (2022) — What Makes In-Context Learning Work?EMNLP 2022, peer-reviewed
- Liu et al. (2022) — Good In-Context Examples for GPT-3DeeLIO / ACL ecosystem, peer-reviewed
- Dong et al. (2024) — A Survey on In-Context LearningACM TIST, peer-reviewed
- Rubin et al. (2022) — Learning to Retrieve PromptsNAACL 2022, peer-reviewed
- Lu et al. (2022) — Fantastically Ordered PromptsACL 2022, peer-reviewed
- Zhao et al. (2021) — Calibrate Before UseICML 2021, peer-reviewed
Long-Context Modeling
- Liu et al. (2024) — Lost in the MiddleTACL Vol. 12, peer-reviewed
- Beltagy et al. (2020) — LongformerarXiv preprint
- Zaheer et al. (2020) — Big BirdNeurIPS 2020, peer-reviewed
- Dai et al. (2019) — Transformer-XLACL 2019, peer-reviewed
- Press et al. (2022) — ALiBiICLR 2022, peer-reviewed
- Su et al. (2021) — RoFormerNeurocomputing 568, peer-reviewed
- Chen et al. (2023) — Positional InterpolationICML 2023, peer-reviewed
- Peng et al. (2023) — YaRNICLR 2024, peer-reviewed
- Xiao et al. (2023) — StreamingLLMICLR 2024, peer-reviewed
- Munkhdalai et al. (2024) — Infini-attentionarXiv preprint, Google Research
- Zhang et al. (2024) — Found in the MiddleNeurIPS 2024, peer-reviewed
Context & Prompt Compression
- Jiang et al. (2023) — LLMLinguaEMNLP 2023, peer-reviewed
- Jiang et al. (2024) — LongLLMLinguaACL 2024, peer-reviewed
- Jiang et al. (2024) — LLMLingua-2ACL 2024, peer-reviewed
- Li et al. (2023) — Compressing Context to Enhance Inference EfficiencyEMNLP 2023 Findings, peer-reviewed
- Chevalier et al. (2023) — AutoCompressorsEMNLP 2023, peer-reviewed
- Mu et al. (2023) — Gist TokensNeurIPS 2023, peer-reviewed
- Ge et al. (2023) — Extensible PromptsEMNLP 2023, peer-reviewed
Memory
- Packer et al. (2023) — MemGPTresearch preprint, Letta/MemGPT
- Park et al. (2023) — Generative AgentsUIST 2023, peer-reviewed
- Zhong et al. (2023) — MemoryBankAAAI 2024, peer-reviewed
- Shinn et al. (2023) — ReflexionNeurIPS 2023, peer-reviewed
- Wang et al. (2023) — Voyagerresearch preprint, NVIDIA
- Sumers et al. (2023) — Cognitive Architectures for Language AgentsTMLR, peer-reviewed
Tools & Agents
- Yao et al. (2023) — ReActICLR 2023, peer-reviewed
- Schick et al. (2023) — ToolformerNeurIPS 2023, peer-reviewed
- Mialon et al. (2023) — Augmented Language Models: A SurveyTMLR, peer-reviewed
- Patil et al. (2023) — Gorillaresearch preprint, Berkeley/Microsoft Research
- Qin et al. (2023) — ToolLLMICLR 2024, peer-reviewed
- Paranjape et al. (2023) — ARTarXiv preprint
- Gao et al. (2023) — PALICML 2023, peer-reviewed
- Chen et al. (2023) — Program of Thoughts PromptingTMLR, peer-reviewed
- Zhou et al. (2023) — Language Agent Tree Search (LATS)ICML 2024, peer-reviewed
Agent Planning & Self-Improvement
- Yao et al. (2023) — Tree of ThoughtsNeurIPS 2023, peer-reviewed
- Madaan et al. (2023) — Self-RefineNeurIPS 2023, peer-reviewed
- Luo et al. (2025) — Self-Training LLMs for Tool-Use Without DemonstrationsFindings of NAACL 2025, peer-reviewed
Industry & Practitioner Sources
- Chase — "The Rise of Context Engineering"LangChain Blog, 2025-06-23
- LangChain Team — "Context Engineering"LangChain Blog, 2025-07-02
- Anthropic Applied AI — "Effective Context Engineering for AI Agents"Anthropic Engineering, 2025-09-29
- Microsoft Research — "Efficient AI Applications: Context Engineering and Agents"Microsoft Research
- Microsoft Research — ACE publication pageMicrosoft Research, 2025-10
- Google Cloud — "What Is AI Context Engineering?"Google Cloud, 2026-04-23
- OpenAI — Consensus customer case studyOpenAI, 2025
- Dickson — "ACE Prevents Context Collapse"VentureBeat, 2025-10-16
- WIRED — "Do Large Language Models Dream of AI Agents?"WIRED, 2025
- Paniego & Gosthipaty — "Harness, Scaffold, and the AI Agent Terms Worth Getting Right"Hugging Face Blog, 2026-05-25
- Chen — "Context Engineering & Reuse Pattern Under the Hood of Claude Code"Hugging Face Blog, community article, 2025-12-22
- Tavily Team — "Building Deep Research: How We Achieved State of the Art"Hugging Face Blog, 2025-11-24
- Polak — "Beyond Prompting: Context Engineering and Memory Management at Scale"InfoQ, 2026
- Ng / DeepLearning.AI — The Batch, Issue 328The Batch, 2025-11-19
SelfMinds AI publishes this kind of literature review because we build directly on top of it — our retrieval, memory, and multi-agent architectures are informed by the same research this piece surveys. If you're evaluating how a context-engineered system should be architected for your data, we're glad to talk through it.
Discuss Your AI Architecture →