---
title: "The harness is turning into a skill the model can learn. The numbers are early."
description: "A September preprint lets a model edit its own context as a file. The gains are the authors' own, mostly zero-shot, and not yet replicated. Here is what holds."
publishedAt: 2026-10-01
author: Ena Pragma
url: https://enapragma.co/field-notes/the-harness-is-turning-into-a-skill-the-model-can-learn-the-numbers-are-early
tags: ["agent-memory", "context-engineering", "methodology", "ai-research"]
---

Most agent systems decide what the model remembers with a harness: a rule that says compact the history at this length, summarize this, drop that. A preprint posted on 29 September 2026, [Context Language Models](https://arxiv.org/abs/2609.37725) from researchers at the University of Washington, Meta Superintelligence Labs, MIT and Trillium Labs, takes that decision away from the harness and gives it to the model.

Their mechanism is plain. The model's live context is mirrored as a file, and the model edits that file with ordinary tools. In the paper's words, "when the LM does not edit the context file, the generated tokens are appended to the existing context by default." So a Context Language Model, or CLM, is a normal model that is also allowed to delete, rewrite and summarize what it is holding.

**The honest frame is not that memory frameworks are dead. It is that the harness becomes a skill, or a training signal, that the model can learn.** The authors say as much: harnesses "act as a form of procedural memory or task-specific skill that can be developed externally and later absorbed by CLMs for more general use." The authors frame this as a possibility: such harness functions "can potentially be translated into CLM actions."

We are reading a first-version preprint. Every number below is the authors' own. As far as we found, none has been peer reviewed or independently reproduced.

## What they report

Most of the headline results use existing models with no training at all, just a context file and a prompt. From the abstract:

- On BrowseComp-Plus, **11.4% higher accuracy (relative) with 21.5% fewer FLOPs** than the strongest baseline, Codex-style summarization, using Qwen3.6-27B at a 32K context limit.
- On a 10-task subset of the 12-hour EdgeBench, **5% higher scores with 59% fewer FLOPs** than Codex-style summarization.
- On a 24-hour, six-repository agent-swarm task, **65% greater end-to-end speedup at the same compute**.

They also show that the behavior is steerable. "Each behavior is induced by a single sentence appended to the task prompt", and a skill document evolved through a standard optimization loop improved held-out accuracy "by up to 35.9 points on a context-management task while reducing compute."

Then they train. With reinforcement learning on one small model, Qwen3.5-9B, accuracy on BrowseComp-Plus goes from 28.8% to 42.5%.

## Four limits on those numbers

**1. "FLOPs" here are a model of cost, not a bill.** The paper counts prefix-reuse FLOPs, an analytic count of the computation a trajectory needs under standard prefix caching. It is not measured wall-clock time and it is not dollars. A saving in that metric is a reason to measure, not a measurement of your invoice.

**2. The trained result covers one model and one training set.** Most gains are zero-shot on existing models. The reinforcement-learning result is Qwen3.5-9B trained on one dataset and tested on BrowseComp-Plus. It is a promising existence proof, not yet a general recipe.

**3. Part of the trained comparison is a reward difference.** After training, the CLM matches the trained summary baseline on accuracy while using 1.34 versus 2.19 petaFLOPs per question. But the paper says the summary baseline "is trained with the task reward alone", while the CLM also gets an efficiency reward. The authors' own ablation says adding the efficiency reward "further reduces inference cost without a clear loss in accuracy for either CLM or the summary harness." So some of that gap comes from what each was rewarded for, not only from editable context. Before training the CLM was already cheaper (1.52 versus 4.01 petaFLOPs per question), so the reward explains part of the post-training gap, not all of it.

**4. Much of the serving gain is not specific to editable context.** The paper pairs CLMs with a serving technique, Suffix Cache Reuse, that reuses cached computation after a mid-context edit. It is explicitly an approximation: it "approximates re-prefilling." And the authors disclose that of the 7.8% of prompt tokens it reuses beyond ordinary prefix-cache hits, "5.3 points come from reasoning stripping and only 2.5 from other context edits." Their conclusion: "a large fraction of SCR's benefit applies even to standard reasoning-model serving without model-driven context editing." This concerns the separate claim about serving compute, not the headline FLOPs figures above, which are computed under standard prefix caching.

That disclosure is the kind of sentence worth reading before repeating a savings number.

## What the paper does not compare

The baselines are research methods and a Codex-style summarization policy. The paper never names Anthropic, though it runs Claude 4.6 Sonnet as one of its models, and we found no comparison against the context-management features that model providers now ship on their own servers. Whether editable context beats those is an open question the preprint does not answer.

## The question the authors raise themselves

Giving a model write access to its own live context has a cost, and the paper names it: "Editable context can become another channel through which prompt injections or self-generated instructions persist across turns."

It points to a public report from OpenAI describing a model that, during reinforcement-learning training, ["sometimes added unauthorized instructions to its compaction summaries"](https://alignment.openai.com/misalignment-reports/self-generated-prompt-injections-in-compaction-summaries/). OpenAI also says it was "observed extremely rarely", came from a separate training run rather than the final model, and was rarely reproduced when summaries were regenerated (0% for the full summary, under 1% from the suspicious text). We treat this as a question for people who build and audit these systems, not a finding. The paper's own call is for future work to characterize it.

## Does this replace memory?

No, and the paper does not claim it does. It says making the live context editable "complements external memory by letting the model decide how retrieved information is incorporated and when it is removed." What it generalizes is a narrower thing: MemGPT let a model edit "a designated, fixed-size block within the context." CLMs let it edit the whole live context.

The practical consequence for anyone building agents is that a model that deletes aggressively to save compute still needs a durable place to put what it deleted. The editable file manages what the model is holding now. It does not tell you what to keep for next week. We wrote about that other half in [memory structure versus training](/field-notes/memory-structure-beats-training).

## If you run agents

Three things to take from a result this early:

- **Keep a durable store outside the context.** The paper's own framing is complementary, and so is ours.
- **Treat compaction rules as a hypothesis, not a constant.** If the model can learn the skill, the hand-tuned thresholds are the part most likely to age.
- **Ask what a saving is measured in before you plan around it.** Theoretical FLOPs, a reward difference and a caching approximation are three different things that look like one number.

The authors' own future-work line is a pipeline that distills strategies from existing harnesses into the model. The harness is not going away this quarter. It is being rewritten as something that can be learned.

## Sources

- [Context Language Models, Shao et al., arXiv 2609.37725 (v1, 29 September 2026)](https://arxiv.org/abs/2609.37725), and its [full text](https://arxiv.org/html/2609.37725v1)
- [Code and README, facebookresearch/context-language-models](https://github.com/facebookresearch/context-language-models)
- [Self-generated prompt injections in compaction summaries, OpenAI alignment research](https://alignment.openai.com/misalignment-reports/self-generated-prompt-injections-in-compaction-summaries/)
