Same task, twice: how to test whether giving your AI assistant more context saves you money
A demo halved one AI bill with extra context. Here is a twenty-task test an owner can run, and why that demo's own tokens and dollars fell by different amounts.
Contents
At the AI Engineer World's Fair in San Francisco on July 2, 2026, Peter Werry, a founding engineer at Unblocked, put two sessions of an AI coding assistant on screen. He had asked it to plan an optimization of one component of his own company's software, once without extra company context and once with it. The usage panels read $2.59 without and $1.29 with. He called that "a 50% reduction in cost which is pretty incredible." The video was posted to YouTube on October 10, 2026.
Those numbers are on screen in the video, and they are one pair of sessions, shown by the vendor. This note covers what the talk was about, what its own numbers show, what the vendor's other savings figures rest on, what independent research says, and the test we would run before paying for any tool that promises to make an assistant cheaper by giving it more to read: twenty of your own tasks, each done twice.
#What the talk was about
The title promised a relational context engine that cuts token burn (a token is a fragment of a word, the unit AI models are billed in). Strip the jargon and the idea is this. Context is the extra information you hand an AI assistant along with a question: the code, the old tickets, the chat thread where a decision was argued. A context engine, as the vendor describes it, is a layer that connects a company's systems (the vendor lists GitHub, GitLab, Jira, Linear, Confluence, Notion, Google Drive, Slack, Teams, Sentry and Datadog among them), relates what is in them, and hands the assistant a short bundle with citations, limited to what the person asking may see. In the talk's words the result is context that is "interrelated," "intent specific," "personalized" and "permissions aware."
By the video's own chapter list, about 12 of its nearly 20 minutes, from "Let's build" at 6:17 to "Context engine simulator" at 18:30, go to building a structured-query engine, and the thesis behind it is narrow. Werry says retrieval "has two halves essentially": semantic retrieval, which finds passages that read like your question, and "the structured part." His examples of what the first half cannot handle are which pull requests (proposed code changes) he merged last week and who reviewed the payments code most. His verdict: "these are queries. They're not semantic searches."
That is the useful idea, and it has little to do with tokens. Which jobs closed last week and who has the most open orders are lookups over records, not searches over text. An assistant that only finds similar-sounding passages is the wrong tool for them. (That is our reading, not the speaker's.) The title says token burn, but in the auto-generated captions the speaker never says "token." The spoken claim is about dollars. Token counts appear on the usage panels.
To show the structured half, Werry builds a small version live over GitHub data. The vendor's repository calls it a "Workshop project." It is a teaching scaffold: the talk calls it "a component of a context engine." The vendor's own post says some questions "are answered through relational queries against structured data," so the idea is in the product, but we found nothing saying the whole engine is built like the workshop project. His own framing: "generating the queries is actually the easy part. The hard part is the scaffolding that sits around it." The scaffolding is simple to describe:
- Learn the layout of the records by sampling them, mostly with what he calls "traditional procedural methods" rather than an AI model.
- Match a name someone typed to the same person across systems by fuzzy matching, with an AI model optionally brought in to "act as a tiebreaker." Werry says that is "more or less how we do it" in the vendor's product.
- Check the generated query before it runs. Dangerous operations are refused, and sensitive fields, such as which customer's data a query may touch, are kept away from the model and added underneath by the system. In the demo, a request for pull requests with "authentication" in the title was bounced: "regular expressions aren't allowed you're getting bounced."
- If the check fails, feed the error back and try again, with a cap. "It's just a loop with some sort of maximum retry." In the demo it "took a couple attempts" and landed on an exact-title match that found nothing, so he adds a text search: "So we need one more thing which is this full text search."
The product is sold per user. The vendor's comparison page lists $29 per user per month with a 21-day trial, and custom pricing for Enterprise.
#What the demo showed, and what it did not
The two usage panels are the best evidence in the talk, because they are on screen and not in a slide.
| Without extra context | With extra context | |
|---|---|---|
| Total cost | $2.59 | $1.29 |
| Model time (the panel's "Total duration (API)") | 3m 6s | 1m 33s |
| Time the session was open (the panel's "Total duration (wall)") | 1h 9m 23s | 1h 11m 50s |
Cost fell 50.2% and model time fell 50.0%. The speaker's "about a minute and 30 seconds" against "3 minutes" is model time. The two sessions were open for about the same stretch of the clock.
The panels also list tokens, and they split them four ways: fresh input (what you send), output (what the model writes), and two kinds of cache, which is text the model has already seen in the session, kept on hand so it can be reused. A recent independent study of coding agents found that "prompt-cache creation and reads dominate the measured input-side cost". So what you count as a token changes the answer. We recomputed the drop four ways from the numbers on screen:
| What is counted | With | Without | Fewer |
|---|---|---|---|
| Fresh input and output only | 31.2k | 38.3k | 18.6% |
| Plus cache writes | 159.3k | 225.7k | 29.4% |
| Plus cache reads, no cache writes | 460.7k | about 2.04M | about 77% |
| All four at par | 588.8k | about 2.23M | about 73 to 74% |
| Dollars | $1.29 | $2.59 | 50.2% |
Both fell. But tokens fell by less than dollars under the two narrow counts and by more under the two that include cache reads, so the figure in the title and the figure on the bill are different quantities. The without-context session averaged about $1.16 per million tokens against about $2.19 for the with-context session, and cache reads were a bigger share of the without-context session's tokens (about 90% against 73%). That is consistent with cache reads costing less per token than fresh text, which is our inference from those two averages. We did not look up a price list. The without-context totals are approximate, which is why the cache-read rows say "about": they include a small side call to another model, and the screen rounds that session's cache reads to "2.0m." On that rounding the two cache-read rows could move by about a point.
What this pair cannot tell you matters as much. It is one session on each side, with no repeats shown. The without-context session produced about twice the output tokens (12.0k against 6.1k), so the two plans were probably not identical work; that is our inference, and nothing on screen compares quality. The task was a plan, not finished work, on the speaker's own software, and the bundle it was handed included Notion architecture documents, pull requests and Slack conversations "that were had about optimizing this particular component." The demo shows that context can cut cost on a task like this. It does not show how often.
#The vendor's other numbers, and what each rests on
The talk's dollar figure is one of several the vendor publishes. Read the denominators.
| Figure | What it rests on |
|---|---|
| 48% fewer tokens, 60% lower cost, 83% faster | One Kotlin SDK task in the vendor's own codebase, described by April 28 at the latest. The July 29 post gives the percentages and says an AI judge scored it. The outcomes page describes it as the same task run twice, once each way. |
| 41.6% fewer tokens, 26.6% faster | One frontend task, from the same July 29 post. The comparison run "ran out of turns and never produced working code," so, we infer, its token count stops where it gave up. |
| A session that "cost about 10x more" | One question the vendor ran internally once each way, in the same July 29 post. It says the session without the context engine spent 81 turns exploring the codebase, "roughly 84% of a session that cost about 10x more," and "still came back with the wrong answer," while the session with it answered correctly in 40 seconds. No dollar amounts and no repeat count are given. |
| 66% fewer tokens per agent task | Shown on the outcomes page under the line "Across nearly 1,000 repositories." The vendor's comparison page traces it to one engineering manager benchmarking "the same question with and without," 5,000 tokens against 15,000. |
| 40% lower median token spend, 30% faster median time | Twenty tasks, the best-sized figure of the set, run in a customer's own test setup and quoted on the vendor's site. The page shows no date, but its metadata says it was published September 18, 2026. In the same story the first single run finished three times faster while using a few more tokens. |
The vendor describes the evaluation it runs during a customer's proof of concept as "an A/B comparison: the same tasks in the same harness, one session with the context engine and one without," with a blind AI judge scoring quality. That is the right shape. The post that describes it mentions that evaluation in one paragraph, separate from the figures above, and gives no sample size or repeat count.
Credit where it is due: the vendor does not ask for trust. Its July 29 post says "You shouldn't take our word for any of that" and tells readers to "Point it at your own repo, use your own tooling and your own model key, and see what the delta looks like on your codebase instead of ours." The talk closes with an open-source simulator, a test harness that runs the same task both ways, which Werry says you can run "entirely locally." (It is a way to repeat the experiment, not a record of how the April Kotlin run was made: the repository was created on May 29, 2026, and the April 28 post points to an earlier harness, unblocked-compare.) And the vendor's own comparison with plain retrieval (RAG, the standard approach of searching documents for passages similar to the question) concedes the boundary: "Use RAG when you have a single authoritative corpus," and "If your agents consistently ship working code, your RAG is sufficient." A vendor that publishes its harness and says when you do not need its product is inviting exactly the test below.
#What independent research says
We found no independent test of this vendor's token or cost claims. The two customer figures above, the 20-task test and the 5,000 against 15,000 tokens, were run by the customers themselves, not by an outside evaluator, and as far as we found they appear only on the vendor's own pages. We searched arXiv titles and abstracts, with a known paper as a positive control, and about 25 web queries on October 10, 2026. That describes our search, not the world. What exists is work on neighbouring techniques, and it is mixed.
- Extra context files cost more and did not generally help. A study of repository-level context files for coding agents found that providing them "does not generally improve task success rates, while increasing inference cost by over 20% on average". These are static notes files, not a live engine. A second study, 288 runs on 17 tasks from three repositories, found that context strategy "does not measurably move correctness", with a limit: it could only rule out differences larger than 10 to 15 percentage points.
- Fewer tokens did not reliably mean a smaller bill. In a study of three token-reduction approaches for coding agents, the largest setup cut tool-output tokens by 38.4% "but increased billed cost by 6.8%". Across tasks, token reduction was only weakly related to cost reduction (a correlation of 0.15, where 1.0 would mean the two move in lockstep). The authors say to judge "cost per successful task."
- More elaborate search was not automatically better. On a code question-answering benchmark, plain semantic search answered 65.2% of questions correctly against 46.2% for deep agentic search, at less than half the cost per correct answer.
- Graph retrieval and plain retrieval each win some tasks. A systematic comparison found "distinct strengths" for each, and that combining them led to "consistent performance improvements." It was first posted in February 2025 and its current version is dated March 4, 2026.
None of these tests the vendor's product. All are arXiv preprints, and we did not check whether any has been peer reviewed. Put together they say what the demo's own numbers already said: more context is not automatically cheaper, and tokens are not dollars.
#The test: twenty of your own tasks, each done twice
This is the method we would use. It has the same shape as the vendor's own harness, with the steps spelled out for an owner.
- Pick twenty real tasks from the last month, written down before you start. Not demos. Include the dull ones. If you cannot find twenty, run what you have and say how many it was.
- Run each task twice with only the context changed. Same assistant, same model, same instructions, same number of steps, a fresh session each time, and a coin flip for which version goes first. The vendor's simulator lists three steps for the baseline run and six for the context run, because the context run adds steps that gather context and extract patterns. Those steps may be part of what you are buying, but they are also a difference besides context, so decide before you start whether you are testing the context or the whole workflow.
- Count what you pay in. Where you pay per use, count dollars per finished task from the bill, not tokens. Where you pay a flat rate, count minutes to an acceptable result and the number of corrections you had to make. Either way, put the price of the context tool itself on the "with" side, and count a run that never finishes as a failure, not a cheap run.
- Score quality blind. Have someone who does not know which run had the extra context grade both. The vendor uses an AI judge; for work that matters, use a person.
- Read a spread, not a headline. Twenty runs per side is a rough check, not a statistical guarantee. Look at the median (the middle result), the worst case, and how many tasks went the wrong way. The vendor's own customer story shows why: a single run finished three times faster and used a few more tokens.
- Do not print a savings percentage in a deck, a board update or a renewal until you have measured it on your own work.
#How this applies if you run a small business
The product in the talk is aimed at engineering teams: code, pull requests, tickets. We have not tested it, and nothing here says it will or will not pay off for you. What transfers is the method, and three habits.
- Sort your questions by shape. Find me the passage about our refund policy is a search. Which invoices are past due and who has the most open orders are lookups over records. Match the tool to the shape before you compare prices.
- Keep a live check for anything that changes daily. Open orders, today's stock, overdue invoices: a summary the assistant was given last week is wrong about today. For those, have it look the record up when you ask. The vendor says this is how its engine behaves: "every answer is computed against the current state of your code, conversations, and documents at the moment you ask."
- Ask which records each person may see. A layer that reads every system also reads what people in those systems were never meant to share. The vendor says it maps each system's permissions and "enforces them at query time," filtering to what the asking person is authorized to see, and its comparison page lists SOC 2 Type II and CASA Tier II (we did not check those reports). We found no independent test of how well it enforces each source system's permissions, so ask for a demonstration on your own records: have someone without access to a document ask for its contents.
We run AI agents for small businesses, so this is the question we want answered before paying for any extra context layer. A vendor's number tells you what happened on its task. Twenty of your own tell you what to expect on yours.
Related field notes: Two talks, two missing layers on what an agent knows once it starts, a reasoning paper's 62.6% token cut reported without a variance, and Claude Opus 5.5 sells the same work for less on cost per finished task.
How this was researched: the talk's figures were read from the video itself, its auto-generated captions and its on-screen usage panels, on October 10, 2026. Every figure from a written source was read from the linked page the same day. The vendor's figures are the vendor's own and were not independently verified. Of the vendor pages cited, the outcomes page is undated. The customer story and the simulator's README show no date on the page; the story's metadata says published September 18, 2026, and the simulator repository was created May 29, 2026. The unblocked-compare repository was created April 16, 2026, and the document-query-engine repository was created June 26, 2026, with its files dated June 29. We did not test the product.
#Sources
- AI Engineer, Beyond RAG: A Relational Context Engine That Cuts Token Burn
- arXiv, Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?
- arXiv, Do Context Files Help Coding Agents? A Two-Agent Ablation Study on Real Repositories
- arXiv, Token Reduction Is Not Cost Reduction
- arXiv, Deep Agentic Search for Repository-Level Code Question Answering: An Empirical Study
- arXiv, RAG vs. GraphRAG: A Systematic Evaluation and Key Insights
- Unblocked, A Pile of MCP Connectors Is Not a Context Engine
- Unblocked, Inside the Unblocked Context Engine: How Relevance Is Decided
- Unblocked, The next era of Unblocked: context for agents
- Unblocked, Context Engine vs RAG: Retrieval Isn't Enough
- Unblocked, Unblocked vs Sourcegraph Cody: Which Gives Coding Agents Better Context? (2026)
- Unblocked, Cut token costs by half (outcomes page, undated)
- Unblocked, customer story with a 20-task agent test (no date on the page; metadata says September 18, 2026)
- Unblocked, context-engine-simulator (GitHub; README shows no date, repository created May 29, 2026)
- Unblocked, unblocked-compare (GitHub)
- Unblocked, document-query-engine (GitHub)
A new era.
Room for you.
Keep reading
- · 10 min
Which AI model is really answering? A new test is a drift alarm, not a lie detector
Tsinghua researchers can often tell which AI model sits behind an agent from how it works. What the paper shows, what it doesn't, and how to use the idea at any level.
- · 9 min
AI agents still struggle to read a PDF. The benchmark that shows it was built by the company selling the fix.
Jerry Liu of LlamaIndex argues documents are the missing context layer for AI agents. His own benchmark supports much of it, and it is still a vendor's benchmark.
- · 9 min
Claude Opus 5.5 sells the same work for less. Its benchmark table runs at a setting you may not use.
Anthropic's new flagship is pitched on cost per task, not peak score. The first independent measurement landed the same day, and it complicates that pitch.