How we actually run this
Working methodology and copy-ready methods. The thinking behind how we run agents and the systems around them in production, plus the tests, checks, and disciplines from our posts, each on its own page with a copy button. Take them; they work anywhere.
Methodology
How we run the systems, end to end.The agent knowledge architecture
The full picture behind the memory, vault, and wiki concepts in our posts: the layers, the file structure, the loops, and the gates, as one diagram and one annotated tree.
How we run agent memory in production
Agent memory treated as infrastructure: what the system watches, the principles it runs on, and why it compounds instead of decaying.
Methods
Copy-ready tests, checks, and disciplines. Take them; they work anywhere.The Two-Hop Gap Test
A ten-minute test of whether your AI can combine facts it already knows, run on your own business data. Direct versus stepwise, scored in two columns.
Cold Review
An independent, tools-denied reviewer you run over a finished deliverable before it ships. It may read and open sources; it may only trace, never recompute.
Session Audit
A forensic self-audit over your AI's work records that the producing system cannot grade itself on: pinned inputs, mechanical extraction, blind adversarial review.
Mechanize the Lesson
The checklist to run the second time your AI repeats a corrected mistake: convert the written rule into a mechanism that sits in the path.
Consistency Review
The inward-facing check your source-checking cannot do: stated counts, arithmetic, repeated facts, and references verified against the text itself.
The Agent-Readiness Check
Where your site stands with AI agents, in ten minutes: a neutral yardstick, a five-point form self-check, and a quarterly trigger watchlist.
From the post Should Your Website Be Ready for AI Agents Yet?Six Agent-Memory Disciplines
Six no-training memory disciplines distilled from Stanford's AutoMem ablations: consult before write, upsert over append, and the rest of the improvement loop.
fable5-delegate: escalate to Fable 5, fall back to Opus 4.8
A small escalation boundary for Claude Fable 5. Sends the hardest slice of a job to the most capable model, handles the four API changes that break older code, and falls back to Opus 4.8 on a refusal.
The Intervention Surface Framework
What to change when an agent fails: the failure-to-surface mapping (memory, harness, tools, guardrails, model), the two gates every surviving self-improving system shares, and the ledger line that makes changes auditable.
The Founder's Attack-Surface Checklist
Run an opponent on your own startup idea: the seven ways a company idea dies, each with its base rate, the disconfirming question to ask, and the flip-condition that would make you walk away.