Applied Intelligenceby Arseny Gorokh · @agorokh
applied AI research · independent
in this issue ·ai-13 ·applied research note ·20 aug 2026 ·13 min read

Governance is a metabolism, not a manifest.a deliberate stress test of agent governance: what failed the measurement, the self-hosted reviewer that survived it, and the dated audit of my own last claim.

Nobody hands you evidence about which agent-governance mechanisms actually change behaviour, so I ran the experiment deliberately on my own fleet: ship the candidate gates, measure each against behaviour, keep what proves itself, delete what does not, publish the ledger either way. I expected most of it to survive; it did not: 1,908 lines of advisory hooks retired by a live A/B, a retired synchronizer's zombie that ran CI red for six straight weeks, a hook deprecated within seventy-two hours of promotion. What survived obeys one selection rule: bind to a named enforcer with a planted-violation proof, or expire. The note is the results, the rule, the self-hosted code-review daemon earning its life on a real yield ledger, and, because every claim in this series ships with a re-read date, the audit of the previous note's claim: ratified by four more unattended cycles, twice qualified, one break found while writing.

Read the note  ›
1,908
lines of advisory governance deleted in one measured pass, after a live test showed the prose layer changed no behaviour an enforcing gate did not already control.
924 of 2,993
review threads whose findings were acted on: the yield that keeps the one surviving reviewer organ alive, next to its honest thirty-hour outage.
7, and 4
invariants bound to a specific enforcer with a planted-violation proof, and four on an expiring allowlist that only shrinks. Prose rules are debt with a fuse.
0
escalations adjudicated since the cycle the last note published as proof; nine queued behind a broken doorbell, and the newest cycle wrote no row at all. The audit is dated.

the archive

every note carries one folio. the magnitudes are corpus-specific; the structural claims are what transfer.
ai-1320 aug 2026governance stress test Governance is a metabolism, not a manifest.A deliberate stress test of agent governance on one production fleet: three deletions with the measurement behind each; the selection rule that decides what survives, bind to a named enforcer with a planted-violation proof, or expire; the self-hosted review daemon earning its life; and a dated audit of the previous note's claim, ratified and twice qualified. 13 min ai-1214 aug 2026agent memory An agent learns to remember in the harness, not the weights.Every agent program makes the memory bet; almost none measures it at the moment of use. I measured mine: the same convention went eight of eight delivered as versioned harness context and one of eight through the live retrieval layer, which added zero over a disciplined harness. The channel is the finding. The harness specified, and the two-week decision rule for your own stack. 12 min ai-1125 jul 2026self-improvement Self-improvement is an actuation problem.A weekly loop mines the operator's corrections out of fleet transcripts into versioned behavioural policy. The audit that found the loop was not closing without a human, the held-out acceptance gate that replaced the per-change reviewer, and three unattended weekly cycles honestly reported, including the one that ran green and blind. 13 min ai-094 jul 2026agent verification "It's done." The most expensive tokens an agent produces.Agent completion as a validated workspace state, independently replicated with a different agent family on the authors' released harness, unmodified; four refusals on the way, every one a real method error, and the science explicitly separated from the five practices an enterprise agent program can adopt now. 12 min ai-082 jul 2026engineering memory The meeting forgets. The pull request remembers.Conventional engineering culture discards the reasoning behind its decisions at a measured, well-documented rate; an agent-first workflow preserves the chain as queryable artifacts and enforces the reading. The measured half, the hypothesis half, and the falsifiers. 11 min ai-0730 jun 2026model routing Three open models drove Claude Code through its governance gatesGLM 5.2, Qwen 3.7 Plus, and Kimi K2.7 Code solved an identical fix perfectly, so they separate only under governance friction; a scan of 627 sessions found zero faked-compliance and exonerated the model the first read accused. 14 min ai-0613 jun 2026agent data MCP harvesting: trustworthy data when your agent has a connector, not an APIAn agent that can only reach a system through a flaky MCP connector treats its output as evidence to check, not a value to trust; an entailment gate caught 98 percent of unsupported claims where a similarity gate caught 27. 18 min ai-055 jun 2026agent memory A memory-and-policy layer above the model: the build-versus-buy caseA gateway routes a prompt but does not know the repository, policy, or budget; own the in-path layer that governs memory, policy, and cost while the model stays swappable, and prove it ports across providers. 12 min ai-0419 may 2026memory substrate Choosing the memory substrate for enterprise agents: LightRAG, Graphiti, and the weighting that decidesTwo purpose-built substrates and a filesystem baseline against a five-criterion adoption bar on an SDLC corpus; production weighting separates them by 36 points where uniform weighting calls it a tie, and the entity graph fabricated four times where the chunk-text retriever never did. 14 min ai-0315 may 2026agent telemetry Claude Code through DIAL: eight models, 192 runs, and metering every requestA POC connecting Claude Code to DIAL across eight models; the cheapest that passes everything is open-source, and the routing adapter meters every request for a per-project view of AI-coding cost. 17 min ai-0214 may 2026open models Which models can run Claude Code through DIAL? Five upstreams, and the costThree Anthropic tiers and two open Qwen coders through one gateway; an open 480B passed every task at $0.101 each, the cheapest of the field, and capability tracks scale, not vendor. 16 min ai-0128 apr 2026ingest model Picking the extraction LLM for knowledge-graph ingestion: breadth beats densityEleven models measured as the entity-extraction LLM for a knowledge-graph RAG system; Gemini 2.5 Flash cleared every retrieval cell at roughly three percent of the strongest commercial model's ingest cost, because breadth of extraction beats density of relations. 16 min

An applied-AI reading style, for work that has to survive diligence.

Applied Intelligence is an independent practitioner publication. Each note takes one piece of real engineering, an agent, an adapter, a gateway, a gate, and reports what held and what did not, with the magnitudes kept honest and the structural claims stated so another team can re-derive them.

The register is deliberate: calm on the surface, hard engineering underneath. A finding leads; the evidence follows; the method and the limits sit where a reader can check them. The companion repositories are artefacts, one click away, never decoration: sdlc-dial-adapter and agentic-memory-mcp.

Written by Arseny Gorokh. The notes draw on real systems and name the public platforms they study; all client-specific material is excluded, and the notes are not official publications of any employer.

the publication
notes12
latestai-13
registerapplied AI research
set inNewsreader
elsewhere
An Applied Intelligence publication · independent · 2026 by Arseny Gorokh Set in Newsreader & Spline Sans Mono · the Applied Intelligence reading style