Literature Review Living Intelligence Series · Part 1

Context Engineering for Large Language Models: Foundations, Architectures, and Research Directions

A research synthesis, not a product pitch: what context engineering is, how it differs from prompt engineering and RAG, and what the literature — peer-reviewed and industry — currently says about retrieval, compression, memory, tools, multi-agent context, and the open problems nobody has solved yet.

Context engineering for large language models — information converging into a structured core

The objective is not to give the model more information. It is to give the model the right information, in the right representation, at the right time — and to remove it once it stops being useful.

Why This Piece Exists

This is the opening piece in SelfMinds AI's Living Intelligence series — a set of long-form technical reviews tracing how AI systems progress from a single model call to a fully self-improving platform. Each piece exists to document, with full citation, how a specific layer of that stack is evolving in the published literature and in production practice. This one is not a description of a SelfMinds AI product; it is a synthesis of where the research on context engineering currently stands, written because the term entered wide use in 2025 faster than the underlying literature got organized, and a rigorous map of it is genuinely useful — to us as engineers building on top of it, and hopefully to anyone else doing the same.

The 2025 Survey of Context Engineering for Large Language Models by Mei et al. is a useful anchor for that claim: it reviews more than 1,400 papers in an attempt to unify research areas — retrieval, in-context learning, memory, long-context inference, compression, tool interaction, planning, state management, and agent orchestration — that had been developing largely independently for several years under other names. What follows draws on that survey, on the foundational papers it and related literature trace back to, and on practitioner sources from model developers and applied teams, organized into a single working framework with formal notation where the underlying claims support it.

1. Defining Context Engineering

Context engineering can be defined, more precisely than the phrase is often used, as: the design and optimization of the inference-time information state through which a foundation model perceives a task, its environment, its history, available knowledge, available actions, and intermediate computational state.

This is substantially broader than prompt engineering. A prompt is one component of context. The full context an agent sees at inference time is assembled, not authored, and may contain:

  • system policy and task instruction
  • user state and conversation history
  • retrieved knowledge and memories
  • tool schemas and tool observations
  • plans and intermediate results
  • few-shot examples
  • environment state and multimodal observations
  • output constraints

Context engineering is therefore fundamentally an information-allocation problem under bounded computational resources. It can be stated as an optimization:

Ct* = arg maxCt⊆It U (M, Ct, Tt) − λ1Cost (Ct) − λ2Noise (Ct) − λ3Risk (Ct)
(1)

where It is all information potentially available at inference step t, Ct is the subset or representation actually supplied to the model, M is the model, Tt is the current task, and U measures downstream utility. Equation 1 is a framing device, not a metric anyone computes directly — its value is in making explicit that every context-construction decision trades utility against cost, noise, and risk, rather than simply asking "does more context help."

This framing explains why larger context windows alone do not solve the problem.

Liu et al.'s Lost in the Middle (TACL, 2024) is the paper most directly responsible for this insight becoming common knowledge: it demonstrated that models do not use all positions in a long input with equal reliability, and that accuracy on a task can degrade measurably depending on where in the context the relevant information sits — even when that information is present and the model's stated context limit is nowhere near exceeded.

1.1 From Prompt Engineering to Context Engineering

Prompt engineering asks "how should I phrase the instruction?" Context engineering asks a different question: what should the model know at this particular inference step, where should that information come from, how should it be represented, and what should be excluded? That reframing produces several genuinely different engineering problems, summarized in Table 1.

Prompt EngineeringContext Engineering
Write instructionsConstruct inference state
Mostly staticFrequently dynamic
Human-authoredMachine-assembled
Single model callMulti-step lifecycle
Optimize wordingOptimize information
Prompt templateContext pipeline
Few-shot examplesRetrieval policy
Context is inputContext is managed state
Usually statelessPersistent state possible
Tokens treated uniformlyTokens assigned different utility
Table 1. Prompt engineering versus context engineering, contrasted across ten dimensions.

The 2025 context-engineering survey organizes the field around three broad phases — context retrieval and generation, context processing, and context management — then studies how those integrate into RAG, memory, tool-based reasoning, and multi-agent systems. A finer-grained, implementation-oriented decomposition, used as the organizing spine for the rest of this piece, is:

Framework 1 — The context-engineering pipeline

Source → Retrieve → Rank → Transform → Compress → Structure → Place → Consume → Evaluate → Persist/Forget

The last two stages are easy to skip and are where most production systems leave value on the table: context engineering should not end when the prompt is assembled.

2. The Seven Layers, and Where Context Comes From

A useful way to organize the field is as seven layers: sources, retrieval, transformation, compression, memory, tools/multi-agent acquisition, and lifecycle management. This piece works through each. The first — sources — is deceptively simple: potential context can come from the current user request, conversation history, system or developer instructions, documents, databases, knowledge graphs, vector indexes, web search, APIs, environment observations, multimodal input, execution traces, user profiles, episodic memory, semantic memory, procedural memory, other agents, previous model outputs, code repositories, and runtime state.

That list immediately differentiates context engineering from RAG. RAG is one context-acquisition mechanism among many, not context engineering itself — a distinction worth stating plainly because the two terms are frequently used interchangeably in practice. Lewis et al.'s original RAG formulation (NeurIPS 2020) explicitly combined a model's parametric knowledge with external non-parametric memory retrieved at inference time — a specific architectural pattern, not a synonym for "giving the model outside information" in general.

3. Context Retrieval: Beyond Top-k Similarity

Retrieval, in its narrowest form, asks:

R (q, K) {d1, …, dk}
(2)

But modern context engineering requires substantially more than top-k semantic similarity over a query q against a corpus K. A well-designed system determines whether retrieval is necessary at all, what should be searched, which data source to search, how many results are required, how diverse the evidence should be, whether results conflict with each other, and whether another retrieval iteration is warranted before generation proceeds.

The conventional architecture — query → embedding → similarity search → top-k chunks → LLM — works surprisingly well in practice, but it introduces a specific, well-documented failure mode: wrong query leads to wrong documents leads to an apparently plausible answer built on the wrong evidence. The retriever can become the hidden bottleneck of an otherwise capable model, silently, because a fluent wrong answer gives no external signal that retrieval — not generation — is where the system failed.

3.1 HyDE, Self-RAG, and Corrective RAG

HyDE (Gao et al., ACL 2023) changes retrieval from embedding the question directly to first generating a hypothetical answer document, embedding that, and using it to find real documents — the generated hypothetical acts as a semantic bridge into the document space, often closer in embedding space to a real answer than the bare question is.

Self-RAG (Asai et al., ICLR 2024) introduced a concept that is now foundational to context engineering more broadly: retrieval should itself be conditional. The model is trained to decide whether retrieval is required at all, and to critique both the retrieved evidence and its own generation, rather than blindly retrieving a fixed number of passages for every query regardless of whether the query needs external evidence.

Corrective RAG (CRAG) (Yan et al., ICLR 2025) goes a step further by explicitly evaluating retrieval quality and activating corrective behavior when the retrieved evidence appears weak:

retrieve evaluate
accept rerank expand query rewrite query switch source web search abstain
(3)

The branching structure in Equation 3 is closer to what "context engineering" means in practice than simple RAG is: retrieval as a conditional, self-correcting subroutine rather than a fixed lookup step.

3.2 Hierarchical Context and RAPTOR

Flat chunk retrieval often destroys document structure. A 200-page report has a real hierarchy — document, chapter, section, subsection, paragraph, sentence — and flattening it into independent, uniformly-sized chunks loses much of that structure. RAPTOR (Sarthi et al., ICLR 2024) addresses part of this by recursively clustering and summarizing chunks into a tree that holds information at multiple levels of abstraction simultaneously, so a query can be answered from a document-level summary, a section-level synthesis, or an exact paragraph, depending on what it actually needs. Context, in other words, does not have to have a single resolution — context engineering increasingly resembles multiresolution information processing rather than flat lookup.

4. Context Transformation, Compression, and the Context-Window Fallacy

Retrieved information usually should not enter the model's context unmodified. The path is better thought of as raw information → relevant information → model context, with a transformation layer in between doing extraction, deduplication, summarization, canonicalization, entity resolution, contradiction detection, reranking, metadata augmentation, chunk merging, query-focused summarization, compression, structured conversion, and graph expansion as needed. Many production RAG systems still effectively do retriever → concatenate documents → model; a stronger architecture is retrieve → filter → rerank → reconcile → compress → structure → model. This middle layer is underappreciated relative to how much failure-mode reduction it buys.

Context also has a real inference cost: for Transformer models, longer input sequences affect memory, latency, and attention computation. That makes a central context-engineering question genuinely quantitative: what is the minimum representation that preserves task-relevant information? LLMLingua (Jiang et al., EMNLP 2023) frames prompt compression explicitly as a method for reducing inference cost while preserving useful semantic content, reporting substantial compression ratios under the conditions it tested. LongLLMLingua (Jiang et al., ACL 2024) extends the same approach to long-context settings, explicitly modeling both information density and positional effects together.

Working principle

More context ≠ better context.

It is sometimes true that Accuracy(Ccompressed) > Accuracy(Craw) — because compression, done well, increases the signal-to-noise ratio the model has to reason over.

4.1 The Context-Window Fallacy

One of the more important conclusions in this literature is that maximum context length and effective context length are different quantities. A model advertised as accepting hundreds of thousands or millions of tokens does not necessarily reason equally well over every one of them — this is exactly what Lost in the Middle demonstrated with positional sensitivity. That motivates a concept of context utilization (CU):

CU = task-relevant information successfully utilizedtask-relevant information supplied
(4)

A model with a 1M-token window and poor utilization can underperform a model with a 32K-token window supplied with carefully selected evidence — which is why the metric worth optimizing is not raw token count but effective information density:

Effective Information Density = useful informationtokens
(5)
A scope note on Equations 4 and 5: neither CU nor effective information density is a standardized, published metric with an agreed-upon measurement protocol. They are useful conceptual framings that follow directly from the Lost in the Middle result, not numbers you will find reported in a benchmark leaderboard. Treat them as a way to reason about context quality, not as a metric to cite as established.

5. Context Placement as a Layout Problem

Information ordering itself can influence performance, which means assembling context is partly a layout problem, not only a selection problem. A context ordered SYSTEM → TASK → GLOBAL CONSTRAINTS → RELEVANT MEMORY → RETRIEVED EVIDENCE → CURRENT OBSERVATION → IMMEDIATE QUESTION may behave differently from the same components concatenated in a different or random order. That creates a largely unexplored optimization problem:

π* = arg maxπ (C) Performance (M, π (C))
(6)

where π is a permutation or structured arrangement of the available context components. The long-context positional effects documented by Lost in the Middle make this more than a stylistic formatting preference — it is a genuine, currently under-studied lever on model performance.

6. Memory Is Context Across Time

Memory becomes relevant the moment context engineering spans more than one interaction. A useful decomposition distinguishes five kinds: working memory (information actively required by the current reasoning episode), episodic memory (specific historical experiences — "the user requested architecture A," "the experiment failed because of X," "we changed parameter Y and the result improved by Z"), semantic memory (generalized knowledge extracted from experience — "for this codebase, component X should never directly call Y"), procedural memory (how to perform recurring operations — "when deploying: run tests → build → validate → deploy"), and profile memory (stable facts about users, organizations, environments, or projects). Each of these plausibly needs a different retention policy, retrieval algorithm, storage representation, decay rate, and confidence estimate — treating "memory" as one undifferentiated store tends to blur those distinctions in practice.

6.1 MemGPT and Virtual Context

MemGPT (Packer et al., 2023 — a research preprint from the Letta/MemGPT project, not a peer-reviewed venue) introduced one of the clearest conceptual analogies available for this problem: LLM context can be managed somewhat like memory in an operating system. Instead of putting everything into a limited context window, the architecture manages memory tiers and decides, dynamically, what should occupy the active context at any given moment. The analogy extends further than the original paper spells out, as shown in Table 2.

Operating SystemContext-Engineered Agent
RAMContext window
DiskPersistent memory
CacheWorking memory
Page retrievalMemory retrieval
EvictionForgetting
SchedulerContext controller
Process stateAgent state
FilesystemKnowledge store
Table 2. The MemGPT operating-system analogy, extended column-by-column beyond the original paper's framing.
memgpt Persistent Memory retrieve Archival Memory Active Context Window LLM

This is a stronger conceptual foundation for long-running agents than the common alternative of simply appending every new conversation turn to a growing context indefinitely — a pattern that runs directly into the context-window fallacy discussed above.

6.2 Reflection as Context Generation

Not all useful context exists externally — a system can generate new context from context it already has. Generative Agents (Park et al., UIST 2023) demonstrated this through a loop of observation, memory retrieval, reflection, and planning. The distinction between external context (obtained from outside the model — a document, a tool result) and derived context (produced by processing existing information — a summary, a plan, a hypothesis, a lesson learned, a synthetic query) is a useful one to keep explicit in system design, since the two carry different provenance and reliability characteristics (see §10). Formally:

Ct+1 = f (Ct, Ot, At, Rt)
(7)

where Ot is an observation, At an action, and Rt an outcome at step t. Under this formulation, context is a dynamical state that evolves with the system, not a static prompt assembled once.

Memory, reflection, and the experience layer above them are large enough topics that this series treats them in dedicated depth separately — a forthcoming Memory Engineering piece and a Reflection Engineering piece go further into retention policy, consolidation, and self-critique than this survey-level treatment does. This section stands as the literature foundation those pieces will build on.

7. Tools, Coding Agents, and Knowledge Graphs as Acquisition Mechanisms

Tool use changes a model from a passive text generator into an information-seeking process. ReAct (Yao et al., ICLR 2023) established the now-standard reason → act → observe → reason pattern, where actions obtain information and observations become new context for the next reasoning step. Toolformer (Schick et al., NeurIPS 2023) approached the same problem from the training side: teaching a model which APIs to call, when to call them, how to parameterize them, and how to incorporate the returned result. Under this view, tools are context operators — a search call injects knowledge context, a calculator injects deterministic computation, a database query injects structured state, a vision call injects perceptual information, and a filesystem read injects implementation context. The agent does not merely "use tools"; it uses tools to alter its own epistemic state.

7.1 Context Engineering for Coding Agents

This becomes particularly concrete for large codebases, where the naïve strategy — read the repository, put a large amount of code into context — is inefficient and often infeasible. A more scalable strategy is progressive: repository map → module map → class signatures → function signatures → selected implementation → dependent tests, expanding only along task-relevant branches:

C0 ⊂ C1 ⊂ C2 ⊂ ⋯
(8)

The agent initially needs structural context — what exists and how it's organized — not full implementation detail, expanding deeper only when a specific task branch requires it. This minimizes token expenditure while preserving navigational awareness of the whole codebase.

7.2 Knowledge Graphs as a Context Channel

Knowledge graphs provide a complementary retrieval channel to vector search. Vector retrieval answers, approximately, "what text looks semantically similar?" Graph retrieval answers a different question: "what entities and relationships are structurally connected to this query?" A sophisticated context system can draw on both, plus keyword search, memory, and tool results, as a unified candidate pool before a selection stage narrows it down:

C = Cvector ∪ Cgraph ∪ Ckeyword ∪ Cmemory ∪ Ctool
(9)

This hybrid-retrieval framing is particularly attractive for complex, multi-hop reasoning tasks where a single retrieval channel systematically misses one type of relevant evidence. This series' Knowledge Fusion Engineering and Graph Engineering pieces cover this territory — vector versus graph retrieval, hop limits, provenance through the graph — in dedicated depth.

8. Multimodal and Multi-Agent Context Engineering

The field becomes considerably more interesting once models are multimodal. Context can consist of text combined with several other modalities and derived signal types at once:

C = {text, image, video, audio, OCR, ASR, objects, faces, events, spatial relations, temporal relations}
(10)

Simply converting everything into text can discard important structure. For video specifically, the context problem is not "can my LLM accept a long transcript?" It is: which moments, frames, speakers, visual entities, OCR strings, actions, temporal relationships, and retrieved prior events should actually be available for answering this particular question? That requires modality-aware context selection, not uniform flattening to text — a theme this series' Beyond OCR piece develops in more depth for the document-specific case.

8.1 Multi-Agent Context and Isolation

Multi-agent systems introduce a distinct challenge: the context universe becomes a union of private, shared, message, and environment context —

C = Cprivate ∪ Cshared ∪ Cmessages ∪ Cenvironment
(11)

— and each agent should probably not receive everything known by every other agent. Context can instead be partitioned by role, authority, relevance, privacy, task stage, and expertise, giving a per-agent selection function:

Cit = Select (It, rolei, taski, budgeti)
(12)

This creates an opportunity for context routing analogous to message routing in distributed systems. An underappreciated design principle follows directly from it: information one subsystem knows does not automatically need to become context for every other subsystem. Deliberate context isolation buys lower token usage, reduced distraction, lower information leakage, easier debugging, and a smaller prompt-injection surface. In a typical research-agent pattern, a writer agent does not need the researcher's complete reasoning trajectory — it may only need validated findings, with the researcher's raw exploration kept private. This series' Graph Engineering piece — on coordinating multiple loops — treats multi-agent context routing as one of its central topics.

9. Provenance, Freshness, Conflict, and Security

Four related problems become important once context is assembled from many sources rather than authored by a human: knowing where it came from, how long it stays valid, what to do when sources disagree, and what to do when a source is actively adversarial.

9.1 Context Provenance

Production context systems should store more than the bare content of a context unit. Storing only {"text": "X happened."} loses information a system needs to reason about reliability later. A richer record looks closer to:

json{ "content": "...", "source": "...", "created_at": "...", "retrieved_at": "...", "confidence": 0.93, "authority": "...", "entity_ids": [], "valid_from": "...", "valid_until": "...", "security_scope": "...", "transformation_history": [] }

Context without provenance is a real risk, not a theoretical one, because generated summaries and retrieved facts otherwise become indistinguishable in the model's input once concatenated as plain text — the model has no way to weight a fabricated-sounding but actually well-sourced claim differently from a fluent hallucination unless the system tracks that distinction upstream.

9.2 Context Freshness

Context relevance changes over time, and at very different rates depending on what the content is. For some information, value decays roughly exponentially:

Value (c, t) = Value (c, 0) · e−λt
(13)

with λ varying enormously by content type, as shown in Table 3.

ContextDecay Rate
Mathematical theoremAlmost zero
User preferenceSlow
Company org chartModerate
Stock priceExtremely fast
Current webpagePotentially fast
Task execution stateImmediate
Table 3. Context freshness decay varies by orders of magnitude across content types — a single global TTL is a poor fit for most real systems.

A strong context system needs explicit freshness semantics per content type rather than one global expiry policy.

9.3 Context Conflict

Contradictory context is a largely underdeveloped problem in production systems. If retrieved context contains both "Product price = $99" and "Product price = $129," naïve RAG may simply expose both claims to the model and let it guess. A context-engineered system should instead determine source reliability, temporal ordering, whether both claims can genuinely be true simultaneously, and whether one supersedes the other.

Working principle

Context Selection ≠ Similarity Search.

It includes epistemic reconciliation — deciding which of two conflicting, similarly-retrievable claims is actually true right now.

9.4 Context Poisoning and Security

As tool-using agents ingest arbitrary external content, context becomes an attack surface. Potential attacks include direct and indirect prompt injection, poisoned retrieved documents, malicious tool descriptions, corrupted memory, and instruction/data confusion. Every context unit should carry both a semantic role (what this information means) and an authority level (what this information is permitted to influence) — which argues for keeping instruction, observation, evidence, untrusted content, memory, and tool result as distinct, taggable categories rather than concatenating everything into undifferentiated text where a retrieved document can silently masquerade as a system instruction.

10. Context Caching and the Context Lifecycle

Many inference requests share context — system instructions, a repository map, API documentation, a company ontology often remain unchanged across many calls. That intersects context engineering directly with prefix caching, KV-cache reuse, semantic caching, retrieval caching, and prompt caching as system-level mechanisms. StreamingLLM (Xiao et al., ICLR 2024) illustrates the model-level counterpart of the same problem: managing attention state efficiently across very long streaming interactions. Infini-attention (Munkhdalai et al., 2024 — an arXiv preprint from Google Research, not yet peer-reviewed at time of writing) explores compressive memory built directly into the attention mechanism itself, rather than managed externally.

Pulling the pipeline stages introduced throughout this piece together, context is best understood as passing through a full lifecycle, not a one-shot assembly step:

lifecycle ACQUIRE VALIDATE SELECT RANK TRANSFORM COMPRESS STRUCTURE INJECT CONSUME EVALUATE PERSIST FORGET
Figure 1 The context lifecycle: acquisition and validation through injection, consumption, evaluation, and a final branch to either persistence (memory) or deletion (forgetting).

This lifecycle view — not "prompt building," a one-shot metaphor — is the more useful mental model for engineering a context system: every stage after INJECT still has work to do, and the branch at the bottom, back into memory or into deletion, is what makes the system capable of learning across calls rather than starting from zero each time.

11. Measuring Context Quality

Context quality is worth evaluating along at least ten independent dimensions, since a system can score well on one and poorly on another — high recall with poor signal-to-noise is a common and specific failure pattern, not a contradiction:

QC = f (R, C, A, F, D, P, S, T, U, K)
(14)

where R is Relevance, C is Completeness, A is Accuracy, F is Freshness, D is Diversity, P is Provenance, S is Signal-to-noise, T is Temporal consistency, U is Utility, and K is Cost. Several of these dimensions have narrower, more classically IR-flavored formalizations worth stating explicitly. Context Recall asks how much of the required information actually made it in:

CR = required information includedrequired information available
(15)

Context Precision asks the complementary question — how much of what made it in was actually useful:

CP = useful information includedall information included
(16)

Most naïve systems optimize recall alone, by adding more context — but a high-performing system needs both together. A third metric distinguishes a failure of retrieval from a failure of the model to use what it was given, which the earlier Lost in the Middle discussion (§4.1) shows is not the same failure at all:

CUR = supplied relevant evidence reflected correctly in outputsupplied relevant evidence
(17)

And two efficiency measures turn context construction into an explicit resource-allocation decision rather than an unconstrained "add more" default:

CE = Δ TaskPerformanceContext Tokens CE$ = Δ TaskPerformanceInference Cost
(18–19)
A caveat that applies to this whole section: Equations 14 through 19 are a proposed measurement framework synthesized from the literature's recurring concerns, not a single benchmark's standardized metric suite that every paper reports against. Different papers cited in this piece measure different subsets of these dimensions, using different exact definitions. Treat this as a checklist for what to ask about a context system's quality, not a scorecard with agreed-upon published baselines to compare against.

12. Context Rot, Compilation, and Model-Awareness

As conversations or agent traces grow, context accumulates outdated assumptions, duplicate facts, irrelevant intermediate reasoning, superseded plans, contradictory state, and tool noise. Performance can deteriorate even though the model technically still has capacity available — a system-level phenomenon worth naming explicitly as context entropy accumulation. Good context engineering performs active entropy reduction through summarization, consolidation, deletion, canonicalization, supersession, and deduplication, rather than letting context grow monotonically.

"Context entropy accumulation" and "context rot" are useful descriptive terms for a real, widely-observed pattern, but they are working vocabulary from this synthesis and adjacent practitioner writing (see Anthropic's engineering post on context engineering in the References below) rather than an established term with a formal, agreed-upon definition in the peer-reviewed literature.

A second useful analogy treats context engineering like compilation: raw information → parsing → intermediate representation → optimization → dead-context elimination → compression → context scheduling → model-specific serialization → inference. Under this framing, raw information plays the role source code plays in a compiler, structured knowledge plays the role of an intermediate representation, and the assembled prompt/context plays the role of an executable representation — and just as different compilation targets need different code generation, different target models plausibly should not receive identical context. That motivates a model-aware framing of the core optimization:

Context* = f (Model, Task, State) rather than Context* = f (Task)
(20)

since different models genuinely differ in effective context length, instruction-following behavior, positional biases, tool ability, domain knowledge, and reasoning strength — an optimal context compiler, in this framing, has to be model-aware rather than task-only.

13. Adaptive and Control-Theoretic Context Engineering

Static context pipelines typically fix parameters like chunk size, top-k, and summary length in advance. A better system determines these dynamically: a simple factual question might warrant retrieving one or two facts; a comparative question might warrant multiple independent sources; long-form synthesis might warrant a full retrieve → cluster → summarize → recursively retrieve loop; uncertain reasoning might warrant retrieve → reason → identify missing evidence → retrieve again. Context construction, under this view, becomes a policy:

πC (st) atC
(21)

where atC is a context-management action taken given the current state st. This creates a genuine connection to control theory: at every inference step, a system observes its current state, estimates an information deficit, acquires context, reasons, observes the outcome, and updates state —

st+1 = F (st, at, ot)
(22)

— which makes context engineering, at this level of maturity, a closed-loop controller rather than a one-shot template fill. That is considerably closer to agent architecture than to prompting.

13.1 Agentic Context Engineering

This direction became explicit with Agentic Context Engineering (ACE) (Zhang et al., ICLR 2026), which treats context as an evolving playbook that accumulates and refines strategies through generation, reflection, and curation, rather than remaining a static, hand-authored prompt. Microsoft Research's own summary of the same work (see References — company R&D page, not a substitute for the ICLR paper itself) and independent coverage in VentureBeat both frame it around a specific failure mode the paper names "context collapse." That work marks a useful staged framing for the field's own trajectory:

Staged framing

Context Engineering 1.0 — humans design the context pipeline.
Context Engineering 2.0 — agents participate in managing their own context.
Context Engineering 3.0 — the context-management policy itself learns from outcomes.

That last stage begins to look like self-improving inference infrastructure — the territory this series' later Evolution Engineering piece is aimed at.

14. What Context Engineering Is Not, and a Reference Architecture

Context engineering should not become a fashionable replacement name for every LLM technique that touches information flow. It is not synonymous with prompt engineering, RAG, vector databases, memory, agents, long-context models, or knowledge graphs individually — each of those is a component or mechanism within it. A more precise statement:

Context Engineering = Systematic Management (Prompt, Retrieval, Memory, Tools, State, History, Knowledge, Observations)
(23)

Pulling every mechanism discussed in this piece together into one architecture gives something close to what a genuine context engine looks like in production:

architecture User / Task Intent / Task Model Context Requirements Estimator Retrieval Memory Tools Context Candidate Pool Validation Ranking Deduplication Conflict Resolution Compression Structuring Context Compiler LLM Result / Tool Action Evaluation → Memory Update
Figure 2 A proposed reference architecture for a genuine context engine — from intent modeling through retrieval/memory/tools, a candidate-pool assembly pipeline, the context compiler, inference, and a closing evaluation step that feeds back into memory.
This reference architecture is a synthesis, not a published system diagram from a specific paper — it is one reasonable way to compose the individually-cited mechanisms above into a single pipeline. Production systems (LangChain, LlamaIndex, and similar frameworks) implement varying subsets of it under different names.

15. Open Research Gaps

Several problems in this literature remain genuinely open rather than merely under-optimized. Learned context selection would replace handcrafted top-k rules with a learned P (ci | task, state, model) based on eventual task success, rather than a fixed heuristic. Counterfactual context evaluation would ask, for each included item, whether removing it would change the answer — enabling causal attribution of context utility rather than correlational guesses. Context attribution would determine which specific context units actually influenced each generated claim, which matters directly for auditability in regulated domains. Multi-resolution memory would maintain raw episodes, compressed episodes, semantic abstractions, and procedural knowledge simultaneously rather than collapsing to one representation. Context uncertainty would store confidence not only in model outputs but in the context itself, propagating uncertainty rather than discarding it at retrieval time. A context graph would represent relationships between facts, evidence, agents, tasks, decisions, and outcomes explicitly, instead of keeping a flat memory store. Context garbage collection — agents need genuine, principled forgetting mechanisms, not just growing stores with occasional manual pruning. Security-aware context routing would ensure untrusted information does not inherit authority merely by having entered the context window (see §9.4). Cross-agent context optimization would determine, formally, what information one agent should transmit to another rather than relying on ad hoc message design. And multimodal context compression — compressing video or multimodal history without destroying cross-modal relationships — remains substantially harder than text summarization and is comparatively under-studied relative to text compression.

16. A Research Hierarchy, and Relationship to Adjacent Disciplines

The field organizes reasonably cleanly into three top-level concerns — acquisition, processing, and management — feeding a common assembly and evaluation loop:

taxonomy Context Engineering Acquisition Processing Management Retrieval Generation Tools Sensors Other agents Filtering Ranking Compression Structuring Resolution / Validation Memory Persistence Forgetting Refresh Routing / Budgeting Context Assembly Model Agent Outcome Evaluation Context Update
Figure 3 A research hierarchy for context engineering: acquisition, processing, and management feed a common assembly and inference loop that closes back through evaluation into a context update.

Context engineering also relates to, but is distinct from, several adjacent named disciplines: prompt engineering optimizes instructions; RAG engineering optimizes external knowledge retrieval specifically; memory engineering optimizes information persistence across time; tool engineering optimizes external actions and observations; agent engineering optimizes decision loops; harness engineering optimizes the execution infrastructure surrounding the model; and context engineering, in this framing, optimizes what information reaches inference, and when. These areas overlap substantially in practice but operate at genuinely different levels of abstraction — which is also why this series treats several of them as separate, dedicated pieces rather than folding everything into one.

17. The Core Principle, and Where the Field Is Going

Across this literature, the operating principle worth keeping is stated plainly at the top of this piece: the objective is not to give the model more information — it is to give the model the right information, in the right representation, at the right time, and to remove it once it stops being useful. Figure 4 shows how SelfMinds AI organizes this principle into a concrete stack.

SeMAI context engineering diagram
Figure 4 SelfMinds AI's context engineering stack: acquisition, management, and assembly layers feeding a governed context into inference.

The trajectory this body of work points toward is a progression from prompt engineering, through RAG, through memory, to context engineering, and — per the ACE line of work discussed in §13.1 — toward context policy learning. In that end state, a system would not have a developer hand-specifying top_k = 5, history = last_10_messages, include_summary = True; instead, a context controller would learn, from outcomes, what to acquire, retain, transform, retrieve, expose, hide, compress, or forget before the next inference step, given the current model, task, history, token budget, uncertainty, and available information. At that point, context engineering stops being infrastructure that merely surrounds an LLM and becomes an adaptive information-management layer that materially governs the model's effective cognition — which is also the reason two systems built on exactly the same underlying model can differ dramatically in reliability, cost, memory, and autonomy: the difference is their context architecture, not their model weights.

18. Limitations of This Synthesis

In the interest of the rigor this piece is trying to model, three limitations are worth stating explicitly rather than leaving implicit. First, several of the formalizations above — the ten context-quality dimensions in Equation 14, the efficiency metrics in Equations 18–19, "context entropy accumulation," and the reference architecture in §14 — are proposed synthesis frameworks, not results independently validated in a peer-reviewed benchmark. They are included because they are useful ways to reason about the problem, and each is flagged in-line where it appears; they should not be cited as established metrics with agreed measurement protocols. Second, evidentiary weight varies substantially across the works cited: peer-reviewed venues (NeurIPS, ICLR, ACL, EMNLP, and similar) carry more weight than arXiv preprints, which in turn carry more weight than a company blog post or news article — status is labeled explicitly for every entry in the References section below specifically so this distinction stays visible rather than getting flattened by uniform-looking citation links. Third, this is a synthesis of a fast-moving literature as of mid-2026; several of the papers cited here (ACE, several long-context and compression preprints) are recent enough that independent replication and follow-up critique are still ongoing, and some of the field's own terminology — "context engineering" as an umbrella term — is itself less than two years old at time of writing.

References

Peer-reviewed venue, company/institutional research publication, or unreviewed preprint — these carry meaningfully different evidentiary weight, so each entry is labeled. Links go directly to the primary source. No content beyond what was supplied for this piece's drafting was introduced during writing — this list matches the source bibliography in full.

Context Engineering & Surveys
RAG & Retrieval Foundations
In-Context Learning & Prompting
Long-Context Modeling
Context & Prompt Compression
Memory
Tools & Agents
Agent Planning & Self-Improvement
Industry & Practitioner Sources

SelfMinds AI publishes this kind of literature review because we build directly on top of it — our retrieval, memory, and multi-agent architectures are informed by the same research this piece surveys. If you're evaluating how a context-engineered system should be architected for your data, we're glad to talk through it.

Discuss Your AI Architecture →
Continue Exploring — Living Intelligence Series
Beyond RAG: Building Enterprise Memory Beyond the Vector Database: Knowledge Graphs as a Learning System Why RAG Pipelines Fail in Production Beyond OCR: Intelligent Document AI Part 2 — Harness Engineering (coming) Part 6 — Memory Engineering (coming) Part 8 — Reflection Engineering (coming)
Kourosh Meshgi
Kourosh Meshgi
Chief Scientist, SelfMinds AI
PhD, Machine Learning, Kyoto University. Publications at CVPR, ACL, Interspeech, ICIP. Research in continual learning, computer vision, NLP, and agentic systems. h-index 10.
0:00 0:00