Context engineering
Context engineering: curating what a model sees, and what to log
A definition of context engineering from primary sources, the techniques that follow from it, and the records worth keeping for each.
Context engineering is the work of curating and maintaining the best set of tokens a model sees at inference time, including everything beyond the prompt itself. Anthropic gave that definition in September 2025. It matters because a longer context is not automatically a better one: accuracy and recall can degrade as tokens accumulate. The techniques are careful system prompts, just-in-time retrieval, compaction, notes outside the window and sub-agents.
Last verified 2026-10-07. Sources are listed at the end of the page.
What is context engineering?
Anthropic’s post “Effective context engineering for AI agents”, published on 29 September 2025, defines it as “the set of strategies for curating and maintaining the optimal set of tokens (information) during LLM inference, including all the other information that may land there outside of the prompts.” Prompt engineering, in the same post, is the narrower job of writing effective prompts, particularly system prompts.
The shift matters for agents. An agent runs in a loop and every pass adds tool results, messages and files to what the model reads next. Somebody has to decide what stays, what is summarised and what is dropped. That decision is context engineering.
Why does a bigger context window not solve it?
Anthropic’s post says context must be treated as a finite resource with diminishing marginal returns, and that as the number of tokens in the window grows, the model’s ability to recall information from it decreases. The Claude API documentation says the same in its own words: more context is not automatically better, accuracy and recall degrade as token count grows, and this is known as context rot. That page lists context windows of up to 1M tokens for some current models, as published on 2026-10-07, and still says curating what is in context matters as much as how much space exists.
An academic result points the same way. Liu and colleagues found that performance is often highest when relevant information sits at the beginning or end of the input and degrades significantly when it sits in the middle of a long context. Those results were measured on the models available to them in 2023, so check the position effect on your own model before building around it.
Everything in the request counts toward the window: the system prompt, every message, tool results, images, documents and tool definitions, plus the model’s own output and thinking. Prompt caching changes what you pay for cached tokens, not whether they occupy the window.
What are the main context engineering techniques?
Each technique is named in Anthropic’s post unless the note says otherwise.
| Technique | What it does | Cost or risk |
|---|---|---|
| System prompt design | Sets instructions at the right altitude: specific enough to guide behaviour, flexible enough to give strong heuristics. Avoids brittle hard-coded logic and vague guidance. | Over-specified prompts break on cases you did not foresee. Vague ones assume shared understanding that is not there. |
| Tool design | Gives the agent self-contained tools with clear intended use and avoids bloated tool sets. | Tool definitions use tokens on every request. Overlapping tools make the model guess. |
| Just-in-time retrieval | Keeps lightweight identifiers such as file paths, stored queries and links, and loads the data into context at runtime when needed. | Adds latency and tool calls. A bad identifier or a wrong query loads the wrong thing. |
| Compaction | Summarises a conversation nearing the limit and starts a new window with the summary. | Detail lost in the summary cannot be recovered. The Claude API documentation also describes server-side compaction and tool-result clearing, in beta for some current models as of 2026-10-07. |
| Structured note-taking | The agent writes notes to storage outside the window and reads them back later. | Notes can be wrong, stale or injected. Decide who may write them and how long they live. |
| Sub-agents | Specialised agents work on focused tasks in clean windows and return condensed summaries to a main agent. | More model calls, and the summary is a lossy hand-off. Anthropic’s multi-agent research post reports roughly 15 times the tokens of chat. |
How do you engineer the context for an agent task?
List what the model needs
For one task, write down the facts, instructions and tools it needs to finish. Anything else is a candidate for removal.
Decide what is always in, and what is fetched
Keep the stable instructions in the system prompt. Fetch records, files and documents when the step needs them, by identifier.
Set a budget
Pick a token ceiling per run below the window. Trim tool outputs to the fields the next step uses.
Choose a plan for long runs
Decide in advance when to compact, what to write to notes and when to hand a sub-task to a fresh agent.
Test with a fixed task set
Re-run the same tasks after each change and compare outcomes. A context change that looks tidy can lose the one fact a task needed.
What should you log about context?
- The assembled context per call: system prompt version, retrieved items by identifier, tool results by identifier and token count per section.
- What was dropped, summarised or cleared, with the version of the summary.
- Where each item came from, and for retrieved text, the access check applied. Context from a tool or a retrieved document is untrusted input.
- Which notes were written, by whom and when, and which were read back.
- The model and its context window for that call, so a change of model explains a change of behaviour.
Where does the Model Context Protocol come in?
The Model Context Protocol (MCP) describes itself as an open-source standard for connecting AI applications to external systems: data sources such as files and databases, tools such as search engines, and workflows such as specialised prompts. Each connected server is a way for tokens to enter the window. That makes every server a source to name in your logs and a place where untrusted text can arrive. See MCP and tool security.
Where Swfte fits: governed sources of context
Status words: Built, In progress, Roadmap. No dates are given.
| Part | Status | Note |
|---|---|---|
| Company brain as a governed source of context: graph of people, groups and reporting lines with evidence statuses | Built | Readable through the local API. Each fact carries a status and an age. |
| Document content from the brain with per-object access lists | In progress | Libraries exist, endpoints are not finished. See the RAG pipeline. |
| Context packages and an MCP server for agents | Roadmap | Designed for. See agents on the brain. |
| Cortex knowledge bases | Built | On-device RAG over files, folders and URLs. Cortex does not read the company brain yet, which is Roadmap. |
| Optional prompt compression in Connect | Built | It calls an external MCP engine and fails open: if compression fails, the request goes through uncompressed. So you cannot assume a given request was compressed. This page states no saving. |
This page does not describe a memory feature beyond the above. If your agents run in one framework with one small corpus, you can apply every technique here yourself and you do not need Swfte for it. For the governance side see AI observability.
Sources and last verified
Every dated or technical fact on this page was read from the pages below on 2026-10-07. Anything that could not be confirmed is left out or marked as not verified.
- Anthropic, Effective context engineering for AI agents. The definition and the techniques: system prompts, tools, just-in-time retrieval, compaction, notes and sub-agents.
- Claude API documentation, Context windows. What counts toward the window, context rot, window sizes, compaction and context editing.
- Model Context Protocol, introduction. MCP as a standard for connecting AI applications to data sources, tools and workflows.
- Liu et al., Lost in the Middle (arXiv 2307.03172). How position of relevant information affects performance in long contexts.
- Anthropic, Building effective agents. The augmented LLM with retrieval, tools and memory.
- Anthropic, How we built our multi-agent research system. Token usage of multi-agent systems compared with chat.
Frequently asked questions
What is context engineering?
Context engineering is the practice of curating and maintaining the set of tokens a model sees during inference, including everything outside the prompt such as tool results, history and retrieved documents. Anthropic’s definition contrasts it with prompt engineering, which focuses on writing prompts, especially system prompts.
How is context engineering different from prompt engineering?
Prompt engineering is about wording the instructions. Context engineering decides everything that enters the window across many turns: instructions, retrieved data, tool outputs, history and notes. For a single question the two overlap. For an agent in a loop, context engineering is the larger job.
What is context rot?
Context rot is the decline in accuracy and recall as the number of tokens in the context window grows. The Claude API documentation uses the term and says it makes curating what is in context as important as how much room is available. Test it on your own model and tasks.
What is the difference between compaction and note-taking?
Compaction summarises a conversation near the limit and restarts with the summary. Note-taking has the agent write selected facts to storage outside the window and read them back when needed. Compaction is automatic and lossy. Notes are deliberate and can be kept for much longer.
Does RAG count as context engineering?
Yes. Retrieval is one way to decide what enters the window. Just-in-time retrieval, where an agent holds identifiers and loads data when needed, is a form of it. The choice of chunks, their order and their number are context engineering decisions, as is leaving something out.
Does Swfte compress prompts?
Optionally. Connect can call an external MCP compression engine, and it fails open, meaning a failed compression sends the request uncompressed. It is not a guarantee that a given request was compressed, and this page does not state a saving. Treat it as an option to test on your own traffic.