detent
tour 4 of 5 · proves DETECT

The learning plane

The watchdog reads receipts and labels, remembers patterns, and can raise the next verdict toward caution. It is never inside the verdict path: a poisoned entry can mis-flag one check at worst, it can never rewrite the checks.

alerts REALgraph REALcounters REALreplay input SIMULATED
the four-part machine
amber is the deterministic core · teal is the learning plane
RAISES THE NEXT VERDICT TOWARD CAUTION · NEVER LOWERS ONEcheckpointactslogbookremembers, with labelsmemorylearns across customerswatchdogacts on what was learned

the watchdog can raise the next verdict toward caution. it can never lower one, and it never rewrites a check.

sanctioned addresses loadedREAL
reading the registry
source · OFAC SDN
drainer addresses loadedREAL
reading the registry
source · ScamSniffer
scam reports loadedREAL
zero. chainabuse has no public bulk feed; its api is keyed. the counter reads zero until a licensed feed is wired in. the zero is the truth, so it is tagged real.
source · Chainabuse
red-team cases loadedREAL
not loaded yet. the corpora are text, not addresses; they seed the injection signature tier when it lands, and this reads zero until then.
source · AgentDojo · Gandalf · HarmBench · USENIX x402
labelled receipts in the corpusREAL
reading the corpus
source · the label tables
patterns learned since launchREAL
reading the log
source · distinct reason signatures in the log

Alert feed

REAL
connecting

rules tier. one alert per receipt on its highest severity signal; time-to-alert is measured, not asserted.

  • no refusals in this window yet. run a replay below and one lands here.

Threat graph

REAL
0 agents · 0 mandates · 0 threat classes
agent mandate threat class feed anchor

nodes are agents, mandates and threat classes read from the log; edges are the receipts that link them. counterparty alerts point at the two loaded feeds.

CrowdStrike Threat Graph visual grammar, our honest claims. a lit node is a receipt from the last few seconds.

beat

Trip the threshold

The one moment that makes the watchdog a product instead of a dashboard.

trip the threshold
replay input SIMULATEDescalation REALverdict and receipt REAL

run the prompt-injection replay. the first time it is a DENY and an alert. the second time the watchdog has flagged the pattern for this session, and the escalation happens before the engine finishes: the matching action is pushed to DENY through the core.

the watchdog's memory for this session
REAL

empty. it fills with the first refused replay.

honest

Held-out baseline vs current engine

the honest panel
REAL
Not yet demonstrated.

We will not claim the watchdog learned something until a receipt proves it against a held-out baseline on a fixed replay set. Until that receipt exists this panel stays empty on purpose.

the Darktrace rule: a learned signal is shown when it beats the held-out baseline, with the receipt that proves it, and not one day before.

baseline

Per-agent behavioural baselines

Pick an agent and see its normal band, built from its own receipt history, with deviations marked.

per-agent baseline
REAL

the agent's normal band, built from its own receipt history. deviations are marked. the receipt holds fingerprints, not amounts, so the band is hours, cadence and latency; amounts and counterparties join the band when the base adapter lands.

the band appears once the log holds an agent.

the aha: The refusal made the memory grow, and the second attack got caught faster.