The harness is turning into a skill the model can learn. The numbers are early.
A September preprint lets a model edit its own context as a file. The gains are the authors' own, mostly zero-shot, and not yet replicated. Here is what holds.
Contents
Most agent systems decide what the model remembers with a harness: a rule that says compact the history at this length, summarize this, drop that. A preprint posted on 29 September 2026, Context Language Models from researchers at the University of Washington, Meta Superintelligence Labs, MIT and Trillium Labs, takes that decision away from the harness and gives it to the model.
Their mechanism is plain. The model's live context is mirrored as a file, and the model edits that file with ordinary tools. In the paper's words, "when the LM does not edit the context file, the generated tokens are appended to the existing context by default." So a Context Language Model, or CLM, is a normal model that is also allowed to delete, rewrite and summarize what it is holding.
The honest frame is not that memory frameworks are dead. It is that the harness becomes a skill, or a training signal, that the model can learn. The authors say as much: harnesses "act as a form of procedural memory or task-specific skill that can be developed externally and later absorbed by CLMs for more general use." The authors frame this as a possibility: such harness functions "can potentially be translated into CLM actions."
We are reading a first-version preprint. Every number below is the authors' own. As far as we found, none has been peer reviewed or independently reproduced.
#What they report
Most of the headline results use existing models with no training at all, just a context file and a prompt. From the abstract:
- On BrowseComp-Plus, 11.4% higher accuracy (relative) with 21.5% fewer FLOPs than the strongest baseline, Codex-style summarization, using Qwen3.6-27B at a 32K context limit.
- On a 10-task subset of the 12-hour EdgeBench, 5% higher scores with 59% fewer FLOPs than Codex-style summarization.
- On a 24-hour, six-repository agent-swarm task, 65% greater end-to-end speedup at the same compute.
They also show that the behavior is steerable. "Each behavior is induced by a single sentence appended to the task prompt", and a skill document evolved through a standard optimization loop improved held-out accuracy "by up to 35.9 points on a context-management task while reducing compute."
Then they train. With reinforcement learning on one small model, Qwen3.5-9B, accuracy on BrowseComp-Plus goes from 28.8% to 42.5%.
#Four limits on those numbers
1. "FLOPs" here are a model of cost, not a bill. The paper counts prefix-reuse FLOPs, an analytic count of the computation a trajectory needs under standard prefix caching. It is not measured wall-clock time and it is not dollars. A saving in that metric is a reason to measure, not a measurement of your invoice.
2. The trained result covers one model and one training set. Most gains are zero-shot on existing models. The reinforcement-learning result is Qwen3.5-9B trained on one dataset and tested on BrowseComp-Plus. It is a promising existence proof, not yet a general recipe.
3. Part of the trained comparison is a reward difference. After training, the CLM matches the trained summary baseline on accuracy while using 1.34 versus 2.19 petaFLOPs per question. But the paper says the summary baseline "is trained with the task reward alone", while the CLM also gets an efficiency reward. The authors' own ablation says adding the efficiency reward "further reduces inference cost without a clear loss in accuracy for either CLM or the summary harness." So some of that gap comes from what each was rewarded for, not only from editable context. Before training the CLM was already cheaper (1.52 versus 4.01 petaFLOPs per question), so the reward explains part of the post-training gap, not all of it.
4. Much of the serving gain is not specific to editable context. The paper pairs CLMs with a serving technique, Suffix Cache Reuse, that reuses cached computation after a mid-context edit. It is explicitly an approximation: it "approximates re-prefilling." And the authors disclose that of the 7.8% of prompt tokens it reuses beyond ordinary prefix-cache hits, "5.3 points come from reasoning stripping and only 2.5 from other context edits." Their conclusion: "a large fraction of SCR's benefit applies even to standard reasoning-model serving without model-driven context editing." This concerns the separate claim about serving compute, not the headline FLOPs figures above, which are computed under standard prefix caching.
That disclosure is the kind of sentence worth reading before repeating a savings number.
#What the paper does not compare
The baselines are research methods and a Codex-style summarization policy. The paper never names Anthropic, though it runs Claude 4.6 Sonnet as one of its models, and we found no comparison against the context-management features that model providers now ship on their own servers. Whether editable context beats those is an open question the preprint does not answer.
#The question the authors raise themselves
Giving a model write access to its own live context has a cost, and the paper names it: "Editable context can become another channel through which prompt injections or self-generated instructions persist across turns."
It points to a public report from OpenAI describing a model that, during reinforcement-learning training, "sometimes added unauthorized instructions to its compaction summaries". OpenAI also says it was "observed extremely rarely", came from a separate training run rather than the final model, and was rarely reproduced when summaries were regenerated (0% for the full summary, under 1% from the suspicious text). We treat this as a question for people who build and audit these systems, not a finding. The paper's own call is for future work to characterize it.
#Does this replace memory?
No, and the paper does not claim it does. It says making the live context editable "complements external memory by letting the model decide how retrieved information is incorporated and when it is removed." What it generalizes is a narrower thing: MemGPT let a model edit "a designated, fixed-size block within the context." CLMs let it edit the whole live context.
The practical consequence for anyone building agents is that a model that deletes aggressively to save compute still needs a durable place to put what it deleted. The editable file manages what the model is holding now. It does not tell you what to keep for next week. We wrote about that other half in memory structure versus training.
#If you run agents
Three things to take from a result this early:
- Keep a durable store outside the context. The paper's own framing is complementary, and so is ours.
- Treat compaction rules as a hypothesis, not a constant. If the model can learn the skill, the hand-tuned thresholds are the part most likely to age.
- Ask what a saving is measured in before you plan around it. Theoretical FLOPs, a reward difference and a caching approximation are three different things that look like one number.
The authors' own future-work line is a pipeline that distills strategies from existing harnesses into the model. The harness is not going away this quarter. It is being rewritten as something that can be learned.
#Sources
A new era.
Room for you.
Keep reading
- · 7 min
Personas are a landscape, not a straight line. The gain grows with the bend.
A preprint says AI persona activations sit on a curved surface, and following it beats a straight line where the surface bends most. Authors' numbers, three small models.
- · 6 min
Can a Smaller AI Model With Better Memory Beat a Bigger One?
A new Qwen paper trained a 9-billion-parameter agent to navigate its memory as a set of tools instead of consuming pre-fetched context, and it out-scored the same system built on a 397-billion-parameter model. The result is real and useful. The 'small model beats giant' version traveling online drops three caveats that change what it means, and the paper's own word for the result is 'competitive.'
- · 8 min
Does Better AI Agent Memory Come From Training or From Structure?
Stanford built a system to learn memory management as a trainable skill. Its own ablation answered the question: structure, schemas, prompts, and gates delivered most of a 2-4x gain before any training happened. Here is what that means for anyone running agents, and the six disciplines you can adopt without training anything.