An AI agent's notes can shape what it does
Two unrelated October 2026 papers, read side by side. In agents that keep notes, a wrong note may persist. What they show, what they don't, and our rule.
Contents
An AI agent that keeps its own notes is not a blank slate: the notes can steer what it tries next, and a wrong note can keep steering it. Two unrelated preprints posted to arXiv in the first week of October 2026 were put side by side in one YouTube video from Discover AI, uploaded 8 October 2026, and only the one from MIT is about notes. In it, MIT researchers left coding agents alone on a simulated island for 30 hours (arXiv 2610.07130, submitted 5 October 2026). The agents "made games" and learned techniques that helped them later (MIT paper, Section 7.4). They also kept wrong notes. One agent logged a juggling feat that did not hold up, and the authors say that result "had motivated both a subsequent experiment and a physical monument" before the agent caught the error (Appendix E). The paper names the problem and floats one direction, agents sharing what they believe, but does not test it (Section 5.3). Ours is a plain rule: a note is a claim to check, not a fact to trust.
Both papers are preprints posted in early October, so their results are the authors' own.
#What the two papers are
The first is an experiment from MIT. In "Is this machine playing?" (Cloos, Norelli, Durbin, Andreas, Rus and Isola; arXiv 2610.07130, 5 October 2026), the authors placed "a modern AI coding assistant in an unintended role: as the mind of a body on an unknown digital island" (abstract). They ask a simple question: "what does a capable artificial agent do when nobody tells it what to do?" (MIT paper, Section 7.4). They built the agent, which they call Eko, by adding three parts to an otherwise unmodified coding assistant such as Claude Code or Codex: "autonomous orchestration, embodiment, and long-term memory" (MIT paper, Section 2.1). The main experiment runs thirteen independent Eko agents on Claude Code with Claude Opus 4.7 (MIT paper, Section 2.3). Each run lasts 30 hours, and the paper states: "No human intervenes during a run or directs it toward any activity." (MIT paper, Section 2.3). The abstract describes the setup as "only a minimal instruction mentioning no specific task, reward, or activity". In practice the agents also had a short persona file ("You are naturally curious. You get bored easily.", Appendix A.2), a standing instruction to "Periodically take a step back and revise your workspace and memory files" (Appendix A.1), and a nudge to keep going: "Please don't stop. Just continue doing whatever you want to do." (Appendix A.4). Within that setup, the agents "climbed hills, stacked blocks into towers, drew mandalas, reinterpreted sports" and "learned techniques that later expanded what it could accomplish" (abstract).
The second is theory from Princeton and UC Berkeley. "The Geometry of Empowerment" (Ji, Myers, Levine and Eysenbach; arXiv 2610.07796, submitted 6 October 2026, revised 7 October) is about empowerment, which its abstract calls "the capacity for an agent to actively control its environment" (abstract). An introductory chapter posted to arXiv in 2013, which the new paper cites, calls it a "task independent utility function", defined as the "channel capacity between an agent's actions and an agent's sensors", a measure of "how much influence and control an agent has over the world it can perceive" (Salge, Glackin and Polani, arXiv 1310.1863, 7 October 2013). In plain words, it asks how much you can change what you will be able to see by what you do. The new paper's theory says that "policies that maximize potential empowerment naturally seek out central states within an MDP", where an MDP is a world of states and moves (empowerment paper, Section 5). Its figures show this on small grids: with walls added, "the high empowerment states" move "to the bottleneck openings" (empowerment paper, Section 4.2). In plain terms, in these small grid worlds an agent that wants this kind of influence heads for the hubs and the doorways.
The two papers are unrelated. Neither cites the other, and the MIT paper never uses the word "empowerment" (we searched the full text of both on 9 October 2026). They were submitted on consecutive days, 5 and 6 October, by different groups (MIT paper, empowerment paper). The Discover AI video that puts them together is titled "Princeton's new AI theory | MIT Experiment on AI Learning" and runs 30 minutes. We read them side by side for that reason. Only the MIT paper is about agent notes.
#How the memory loop works
Eko does not learn by changing its model. The authors write: "This learning happened without changing the underlying model weights: experience accumulated instead in editable text files" (MIT paper, Section 7.4). There are two files: a log of past events, and a semantic file for "knowledge, beliefs, and reusable techniques" (MIT paper, Section 2.1). Every hour "the coding assistant's session is reset and the island is restored to its initial condition" (MIT paper, Section 2.3). After each reset "the system prompt carries semantic memory in full" (MIT paper, Appendix A.3). So the loop is short. The agent writes a note, the note comes back at the start of the next session, and the note shapes what the agent tries.
The notes do carry skill. The authors compared the 13 agents after their 30 hours with 13 fresh agents. The two groups "share the same architecture and the same underlying LLM, with no weight changes; they differ only in the content of their memory files" (MIT paper, Section 5.1). On a one-hour task to build many towers, 62% (8 of 13) of the 30-hour agents built more than two, against 15% (2 of 13) of the fresh agents (MIT paper, Section 5.1). The authors then erased single techniques from selected agents' notes. Erasing every mention of the technique for making extra blocks "drops its tower count from six to two on average" for Agent 2 (MIT paper, Section 5.2). The gain was not uniform. On the tallest-tower task the two groups scored similarly, and "the tallest towers built by two of the 30-hour agents are only 4 blocks high, below the shortest tower built by any 0-hour agent" (MIT paper, Section 5.1).
#Where the loop goes wrong
The same loop carries mistakes, and the paper walks through three, from three of the 13 agents.
In the first, Agent 10 concluded that a two-block juggling technique "is now bulletproof" and recorded it in memory as "CRACKED sustained 2-ball juggle". The authors write that the result "had motivated both a subsequent experiment and a physical monument" (MIT paper, Appendix E). Then Eko proposed a longer test, which reported "50 completed cycles out of 50" in "a total duration of 2.4 s". The authors explain: "A vertical throw at the selected speed required approximately 2.6 s to rise and return." (MIT paper, Appendix E). Eko noticed the contradiction and worked out that most of its catches were recaptures of a block it had only just thrown. It wrote: "My 10-cycle test was probably mostly feints too." The old log entry stayed, because the log is append-only. Eko "added a subsequent entry explicitly invalidating its interpretation", and it rewrote its semantic memory with the corrected procedure (MIT paper, Appendix E).
The other two are notes that discourage. In Appendix C.3, the memory of Agent 12 labels as "verified" the claim that "Jump produces a tiny lift" and is "useless for climbing real obstacles", and it later turns that claim into advice: "Climbing is a dead end" and "Don't waste time trying to walk onto mesas." (MIT paper, Appendix C.3). The paper's gloss is that here "memory preserves an erroneous conclusion together with an instruction discouraging further attempts" (MIT paper, Appendix C.3). Separately, Section 5.3 reports that Agent 12 "never recorded the technique required to jump on platforms in its memory" and, "consistently", "never sets foot on an elevated platform, neither during its 30 hours" nor "during evaluation", "even though jumping works" (MIT paper, Section 5.3). The paper also says Agent 12 "spends most of its run" laying patterns on the ground (MIT paper, Section 3.2), and the authors did not test Agent 12 by erasing a note. They did test Agent 7, whose memory held a false belief about stacking blocks: "erasing all mentions of this belief significantly increases Agent 7's tallest-tower score" (Section 5.3). Their summary: "a false belief, once recorded, may persist and impair later behavior" (Section 5.3).
A note that says do not try X is a dangerous kind, because it removes the one thing that would correct it, which is trying X. That sentence is our reading, and the paper did not test it on Agent 12. The agents already had rules on marking uncertainty and rewriting disproven beliefs. Their semantic memory file told them to "Mark uncertainty explicitly. Distinguish what you have verified from what you suspect." and "When a belief is disproven, rewrite or update it immediately." (MIT paper, Appendix A.3). Agent 12's note still carried the word "verified" beside a wrong claim, which is one case, not a rate. The paper does not test whether such rules work. It lists "how to avoid entrenching false beliefs once written down" among questions "with implications for any agent that learns this way" (MIT paper, Section 5.5).
#Why it matters
Learning without changing the model, through written artifacts such as notes, is spreading. The MIT authors say the rise of coding and assistant agents "has made this mode of learning increasingly widespread" (MIT paper, Section 5.5), and they report that in their runs "agents rarely delete what they have written" (MIT paper, Section 3.3). If so, a wrong note can stay in the file. The practical point is plain: if your agent keeps notes, the notes are part of how it behaves, and they can be wrong. Other groups have measured the same risk in other settings. A 2025 study of agents with a memory bank (arXiv 2505.16067, 21 May 2025) found "error propagation, where inaccuracies in past experiences compound and degrade future performance", and showed that "future task evaluations can serve as free quality labels for stored memory" (Xiong and colleagues). A benchmark paper from August (arXiv 2608.19564, 20 August 2026) opens with the point that "an incorrect durable update can silently distort future behavior" (Li, Yao and Zheng).
#How this applies to us
Nothing in this section is a finding of either paper. It is our practice.
Our agents are built the way the MIT paper describes: a coding assistant, orchestration that keeps it working, tools that let it act, and memory files. So the failures above are not lab curiosities for us. We have seen the business version. A note written yesterday says a job is waiting on something, or that an approach never works, and today's agent trusts the note instead of looking. Our rule is that a note is not a reading. Before an agent tells anyone that something is blocked or waiting, it must re-check live.
The same logic covers what leaves the agent. Before an agent publishes a claim that changes what someone else does, our rule is that it writes down the claim, the command that produced it, a second method that does not share the first method's premise, and what would prove the claim wrong. With no second method, the claim ships labelled unverified or does not ship. The juggling episode shows the habit we want, once, in one agent: it proposed a harder test of its own result and checked the answer against what its world allows. We want that check to be routine, not a lucky event.
#How you can use it
- Date every memory entry and record how it was checked. Re-test any entry that tells the agent not to try something, because that is the entry that hides its own error. The MIT team's template went the other way on dates for its semantic file: "No timestamps, dates, or session references", with history kept in a separate log (MIT paper, Appendix A.3). We would put the date and the check on the entry itself, so a later reader has something to challenge.
- Give a new agent a bounded first phase. Make it read-only, or a safe copy of your tools, and have it log what it learns and how it checked it. Promote a note to real work only after a person or a second check approves it. The paper's framing is play "as a phase of learning that precedes the work an agent is eventually asked to do", with the agent free to "play and explore on its own when the setting is appropriate, and strictly follow instructions when it is not" (MIT paper, Sections 7.1 and 7.3). The paper hedges it with the words perhaps and may (Section 7.4) and does not test it as a remedy. The read-only, log-and-check details are ours.
- Judge an agent by its receipts. A label that says "verified" is not a receipt. A receipt names what was checked, with which command, and what came back.
#What the papers do not show
The MIT paper has no limitations section. We searched its full text for the word and found none, while the empowerment paper has one. The MIT hedges are scattered: "The fixed taxonomy bounds the behavioral diversity this measure can resolve." (MIT paper, Figure 3 caption), "We do not identify the internal mechanisms underlying this behavior or explain how they emerged." (MIT paper, Section 4.5), and, about the agent rather than the study, "Yet Eko sometimes turns suggestive evidence into an appealing explanation too quickly." (MIT paper, Section 7.2). We list only what the authors say or what the setup shows, and we mark our own inferences.
- Thirteen agents are one arm of a larger study. Adding the paper's own run counts gives about 54 runs: 13 on Opus 4.7, 5 on Codex with GPT-5.5, 5 on Kimi Code, 5 on Opus without the persona file, and 13 each on Sonnet 4.5 and Haiku 4.5, where the Sonnet and Haiku runs lasted 10 hours (MIT paper, Section 2.3 and Figure 3 caption). The sum is ours. The Codex and Kimi agents "engage in a much narrower subset of activities" than the Claude Code agents (MIT paper, Appendix C.5).
- A model scored the range of activities. The paper uses "Claude Opus 4.7 as an LLM judge" to read the memory files and mark which activities each agent recorded (MIT paper, Appendix C.2).
- The main comparison has no significance test in the text. We found none for the 13 against 13 comparison of task scores. The memory-edit tests are "three selected memory histories; they are not five independently acquired histories for each intervention" (MIT paper, Appendix C.8).
- The curiosity instruction may not explain the range of activities. In five runs without the persona file and with a reduced system prompt, the authors say "By this analysis we do not observe a clear difference in the range of activities", with the memory files and world unchanged (MIT paper, Appendix C.4). That is a small test, and it is our inference, not theirs, that the curiosity instruction alone may not explain the range of activities.
- Memory is separated from prior knowledge only in part. Both groups in the 0-hour versus 30-hour comparison use the same model, so the comparison isolates what the notes carried. It does not measure how much the model already knew about physics, which is our reading of the design, and the authors do not claim it does. They argue that the findings "cannot be explained solely by the replay of examples memorized during training" (MIT paper, Section 4.5). One headline discovery, throwing mid-jump to launch a block higher, was "a technique we did not know our simulator allowed" (MIT paper, Section 5.4). A footnote in Appendix B says the velocity term "was introduced by Codex (GPT-5.5) when generating the Luau implementation" of the world (MIT paper, Appendix B).
- The empowerment paper is theory with toy experiments. Its experiments are small tabular worlds, such as 5 by 5 and 8 by 8 grids (empowerment paper, Appendix I), and it says "we do not present experiments in high-dimensional or continuous settings" (empowerment paper, Section 2). It tests no language-model agent, and it says nothing about Eko. Its authors call the result "a partial answer to Salge's longstanding hypothesis" (Section 1), and add that the hypothesis "is not generally true for current formulations of empowerment maximization and return maximization" (Section 5).
#What it points to
In our view, the edge in agents is moving from more autonomy to memory that corrects itself. The MIT paper gives one case for why: a wrong claim sat in Agent 12's notes under the label "verified", in a run with no human in the loop. That is one agent of 13. The earlier Voyager agent (arXiv 2305.16291, 25 May 2023) already built in a check, with "self-verification for program improvement" (Voyager abstract), and the MIT paper describes it as one that "stores verified code skills for later reuse" (MIT paper, Section 6). We expect agent products to compete on how their memory is checked. That is our prediction, not a result from any paper above.
On the empowerment paper, separately, the honest reading is to watch, not to build yet. In tabular settings, shown on small grids, empowerment gives a precise way to say what an agent with no task might want, but the paper's guarantee is narrow: it "does provably lower bound goal-oriented behavior in tabular settings", and "our adaptation result is a lower bound, not a ranking result" (empowerment paper, Section 5). Its own list of next steps includes "developing scalable formulations of continuous empowerment maximization" (Section 5). It notes that recent work has used empowerment as "objectives for assistance" (Section 2), and its reference list includes a 2026 paper titled "Training LLM Agents to Empower Humans", which we have not read (empowerment paper, references). The MIT paper does not mention empowerment, and the empowerment paper tests no language-model agent.
Related field notes: long-running agents turn rumors into facts covers the same hardening of a hedge into a fact, and agent memory is infrastructure, not a feature covers keeping a knowledge base from rotting.
The agent worth trusting is the one that goes back and checks its own record.
#Sources
- Cloos, Norelli, Durbin, Andreas, Rus and Isola, "Is this machine playing?", arXiv 2610.07130 (v1, 5 October 2026), and its full text
- Ji, Myers, Levine and Eysenbach, "The Geometry of Empowerment", arXiv 2610.07796 (v1, 6 October 2026; v2, 7 October 2026), and its full text
- Discover AI (YouTube), "Princeton's new AI theory | MIT Experiment on AI Learning", uploaded 8 October 2026
- Salge, Glackin and Polani, introductory chapter on empowerment, arXiv 1310.1863 (7 October 2013)
- Xiong and colleagues, "How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior", arXiv 2505.16067 (21 May 2025)
- Li, Yao and Zheng, "Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM Agents", arXiv 2608.19564 (20 August 2026)
- Wang and colleagues, "Voyager: An Open-Ended Embodied Agent with Large Language Models", arXiv 2305.16291 (25 May 2023)
A new era.
Room for you.
Keep reading
- · 6 min
The harness is turning into a skill the model can learn. The numbers are early.
A September preprint lets a model edit its own context as a file. The gains are the authors' own, mostly zero-shot, and not yet replicated. Here is what holds.
- · 7 min
Personas are a landscape, not a straight line. The gain grows with the bend.
A preprint says AI persona activations sit on a curved surface, and following it beats a straight line where the surface bends most. Authors' numbers, three small models.
- · 9 min
Long-running AI agents turn rumors into facts. The team behind a Supercell lab project proposes a lab-bench fix, and is careful not to say it worked.
A talk on Project Paradox shows how AI agents lose track of where facts came from, and proposes a controlled test loop. What it shows, what it doesn't, and what to do now.