Which AI model is really answering? A new test is a drift alarm, not a lie detector
Tsinghua researchers can often tell which AI model sits behind an agent from how it works. What the paper shows, what it doesn't, and how to use the idea at any level.
Contents
If you use AI through someone else's product, you usually can't check which model actually did the work. You pay for one model, and the service tells you it answered. A new paper from Tsinghua University asks whether you can check that claim from the outside, by watching how an AI agent works rather than asking it what it is. The answer is a qualified yes, and the most useful way to read it is as an early warning that something changed, not as proof that anyone lied.
#What the researchers did
The paper, "Who Is Behind the Harness? Fingerprinting LLMs through Agentic Behavior", is by Chuyi Wang, Xiaohui Xie, Tongze Wang, Fangchen Luo and Yong Cui of Tsinghua University. It was posted on September 23, 2026, as a preprint, which means it has not been peer reviewed, and the authors have not released code or data.
The problem it starts from is real. A provider, a relay or a compromised router, in the paper's words, "can silently serve a different model while continuing to present the advertised identity." Asking the model who it is doesn't settle it, because the service controls what the model says. And more and more AI work runs through an agent "harness": the software around the model that runs its tools, files and steps.
#How it works
The method is called LIDAR. It gives an AI coding agent three small repair jobs, each in a normal and a slightly altered version:
- Verify: does the agent check its own change without being told to?
- Recover: the first test run fails once for a temporary reason. Does it retry, investigate, or flail?
- Resolve: a test contradicts the written spec, and the agent is told to explain the conflict and follow the spec. Does it?
The idea is that a model's habits also show in how its behaviour changes when one condition flips while the task stays the same. In the paper's tests, using both versions of each task did better than either alone. Six runs make one sample, which is turned into 958 measurements of what the agent did and in what order. Each model gets a reference profile, and a new sample is matched to the closest one.
#What it found
Across 36 models from seven families, run inside two open-source harnesses, LIDAR named the right model first most of the time:
| Harness | Right model named first | Right model in top three |
|---|---|---|
| OpenCode | 88.4% | 98.4% |
| mini-swe-agent | 95.1% | 99.6% |
That is roughly one wrong name in nine tests in the first harness and one in twenty in the second, where each test is six runs. It scored higher than earlier fingerprinting methods that look at the model's text, though in OpenCode the best text method was within its margin of error, and those methods were built for direct question-and-answer use, not for agents.
The most striking result is how it handles simple disguises. The authors tested two instructions a provider could slip in. Under an instruction to hide or change what the model says about its identity, LIDAR's first-guess accuracy fell by up to 11 points ("its largest Top-1 drop is 11.11 percentage points"), and the text-based methods fell by up to 33. In one harness, though, the two strongest text methods held steady under that instruction while LIDAR dropped about 8 points. Under a second instruction, to start every reply with a fixed notice, LIDAR did not change at all, while one text method lost 83 points. The paper's own summary is that LIDAR "degrades only modestly." That test used six closely related models the authors picked as a hard case.
#What it does not show
A video summary of the paper described the result as near 100% accurate and virtually impossible to get around. The paper supports a narrower version of the first and does not test the second.
- 100% appears only in the six-model robustness test. On the full 36 models, first-guess accuracy is 88% to 95%.
- "Impossible to bypass" was never tested. The only disguises tried were those system instructions. The authors write plainly: "These signals do not make the fingerprint immune to imitation." They note that other research has shown "a provider can tune a weaker model to imitate a fingerprint." They also describe their result as "not remote attestation," meaning it is evidence, not a cryptographic guarantee.
- It only chooses from a known list. The true model is assumed to be one of the 36 already profiled. Whether it can tell you "this is none of them" has not been tested.
Two questions stay open, and the paper does not answer them:
- Is it telling models apart, or the companies serving them? We found nothing in the paper saying which provider served each model, and other research shows the same model can behave measurably differently depending on who serves it.
- Would a settings change look like a different model? For example, the same model run with less "thinking time," or a compressed version of it. The paper does not test this, and we found no study that has.
#Why it is a drift alarm, not a lie detector
The authors are careful about what their fingerprint captures: "The resulting trace is therefore a joint product of the model and its execution environment." Change the environment and the behaviour can change even when the model doesn't.
That matters because the best-documented case of "my AI got worse" that we found was exactly that kind of change. In April 2026, Anthropic published a postmortem on a drop in quality in Claude Code and two related products. It found three causes: a lower default reasoning setting, a bug that kept discarding the model's earlier reasoning, and a new system-prompt instruction. It says its API was not affected. In its words: "We never intentionally degrade our models, and we were able to immediately confirm that our API and inference layer were unaffected." That is the company's own account, not an independent audit, but it shows how much can change around an unchanged model.
The serving company matters too. OpenRouter, which routes requests to many providers, reported, from real usage, that tool use and its accuracy "vary between providers significantly more than the standard benchmarking would suggest." The sample it published comes from providers of Moonshot's Kimi K2 model. Moonshot itself publishes a vendor verifier to monitor the quality of K2 across the providers that serve it. And a separate study found that "changing the harness changed what unchanged model weights could accomplish."
So the honest reading is this: if a check like LIDAR says your agent no longer behaves like the one you measured, something in the setup changed. That is worth knowing. It is not, on its own, proof that anyone swapped the model. The paper's own framing fits: replay the check over time, and "a sustained deviation is evidence of identity drift."
#What this means at your level
This is our reading of the paper and the sources above, not a tested method.
- Using an AI assistant or coding tool: you can't easily run this yourself. The code isn't released, and it needs reference runs for every model you'd want to recognise. If a tool "feels worse", the documented case we found was settings and harness changes, not a swapped model. "Is it worse?" and "is it the same model?" are different questions, and each needs its own check.
- Building on an API or an aggregator: this is where it bites. The same model name can behave differently from provider to provider. Pin your provider where you can, and keep a small set of fixed tasks you rerun to spot changes.
- Buying an agent platform for a business: ask what the vendor can show about which model ran, beyond the name the provider returns. A behavioural check like this is a cheap tripwire, not a contract-grade guarantee. For contractual assurance, one line of research argues that "software-only methods are fundamentally unreliable" and that the gap "can be more effectively closed with hardware-level security."
The same idea also cuts the other way. Another paper shows a website can identify which model is driving a visiting browser agent from its actions and timing, "with up to 96% F1," which lets an attacker tailor an attack to that model. It is safest to assume your agent's model can be identified from how it behaves.
#The throughline
What. A way to identify the AI model behind an agent by how it works, not what it says: 88% to 95% right on the first guess among 36 known models.
Who. Five researchers at Tsinghua University, in a preprint with no code released yet.
How. Small, paired coding tasks that expose a model's working habits, compared against saved reference profiles.
When. Now, for anyone paying for a specific model through a service they don't run.
Why. The name a service reports is a claim, not a measurement. Watching behaviour gives you a second, independent signal, as long as you read a mismatch as "something changed" rather than "someone lied."
#What Ena Pragma takes from this
We are the buyer in this paper's story. Our own software build workers run models from providers we don't operate, inside OpenCode, one of the two harnesses the paper tested. Our records of which model did the work come from the model name in the provider's own response. That catches a provider that says it switched. It cannot catch one that switches and keeps the label, which is exactly the case the paper describes. We have no evidence that has happened to us, and these records alone could not show it hadn't.
One of our core values is "Receipts over optimism," and a model name in a response is a claim, not a receipt. Our setup suits a LIDAR-style check, because we pin the harness version, each job requests a fixed model name, and we can already replay a job through the same setup. A small set of our own tasks, replayed on a schedule and compared with its own history, would give us a drift alarm on a model we pay for but don't host. It would not prove identity against a provider determined to fool it, and we won't claim it does.
How this was researched: Branden, one of our founders, shared a summary of a video about the paper. Our AI research agent found the paper and the surrounding research and wrote a graded evidence sheet. We then re-read the sources quoted here directly on September 27, 2026: the paper's full HTML version and its abstract page, Anthropic's postmortem, OpenRouter's post, Moonshot's verifier repository, and the abstracts of the three other papers cited. We did not run LIDAR or any of the code, and none has been released.
#Sources
- Wang, Xie, Wang, Luo and Cui, "Who Is Behind the Harness? Fingerprinting LLMs through Agentic Behavior" (arXiv 2609.28559)
- Anthropic, "An update on recent Claude Code quality reports" (April 23, 2026)
- OpenRouter, "Provider variance: introducing Exacto"
- Moonshot AI, K2 Vendor Verifier
- "Same Model, Different Harness: Different Coding-Agent Results" (arXiv 2608.26218)
- Cai et al., "Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs" (arXiv 2504.04715)
- "Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces" (arXiv 2605.14786)
- Earlier Field Notes: An 8B model beat 32B baselines by changing the harness, and Union Alpha was never a mystery model. It is a procurement question.
A new era.
Room for you.
Keep reading
- · 10 min
Self-improving AI agents overfit their own tests. Two September papers each show a brake that works. Neither tests them together.
Google and Oxford researchers each published a way to keep self-improving AI agents from overfitting. What the papers do, what a popular video got wrong, and what EP takes.
- · 12 min
Shopify's AI coding gates are built so the agent can't grade its own work. A YouTube rebuild runs looser.
Shopify built Helix to rebuild its apps with AI agents that must pass four gates. A video recreated it in Claude Code. What Shopify said, what changed, what the evidence shows.
- · 14 min
Spec-driven development has early support. The course's biggest promises are barely tested.
A DeepLearning.AI and JetBrains course teaches writing the spec before an AI agent writes code. Benchmarks back clearer specs. Its own workflow is barely tested.