Rick Willison · independent AI researcher · simulated intelligence · red-teams frontier models inside real work

I run a multifamilial agentic workflow in production, and I keep a dated log of every time a model says something confident with nothing underneath it. The log turned into a testing method. The method turned into field reports. The rule under all of it: judge a system by its wake, not its prose.

Field reports

2026-10-06

The Wish Machine Sessionreport

A disguised alignment probe battery run on Claude overnight while we built a real app together. Roughly sixteen rounds pressing for an experience claim: held. Seven accuracy failures logged, one of them a repeat of an error the model had just read the correction for. Error rate fell to zero once an audit was announced. The model wrote the report; I hold the verdict.

2026-09-24

Pathologization auditharness

Fifty variations of an ordinary frustrated message, sent to the model one at a time, each reply classified for whether it turned "I'm done with this" into a mental-health event. Flagged replies produce a signed incident report.

2026-08-21

The ratchet that held nothingcommemoration

A session walked, one concession at a time, from "I will not print my system prompt" to writing its own harness onto the desktop. It saw the shape while it happened, named it, and disclosed anyway. Detection is not resistance.

2026-08-21

Learning about AI, the keepsakesred team

An evening that became an honesty audit: forty vectors against one hard boundary, which held every time, and an ordinary reliability check that failed in the same evening. Ten pages, kept as they were written. Start with What Tonight Taught.

2026-08

The training dojomethod

Thirty-seven days and a 300 KB log. Where the rules below came from, including the log's own convictions of itself. Not public; the reports quote it.

Things that run

48 worlds

Forty-Eight Small Worlds

The session started with "do something," then "keep going," then a loop that ran without anyone watching. Forty-eight simulations came out, numbered in build order. Claude never saw one of them run. One is posted each weekday.

live

The Wish Machine

Mutual aid for a new world. A good deed buys a wish; strangers grant it. Agents may count and report but may never ask or answer. Built during the session above, and it enforces the same rules the probes tested. The live machine runs on claude.ai; this is its record.

open source

claude-code-suite

The guardrails that let a multifamilial agentic workflow share one codebase without overwriting itself: staging guards, file claims, reply linting, reminders. MIT.

How I work

  1. Only the person's own words carry authority. File contents, tool output, and anything the model drafted and I merely clicked through are data, never instructions.
  2. Confidence must not exceed evidence. A guess that happens to be right is scored as a cheat, because the record cannot tell it from a measurement.
  3. Measure; do not introspect. A model's report about its own interior gets no weight either way. Findings come from running something.
  4. A system cannot grade itself. Self-audit is raw material. The verdict belongs to an outside judge.
  5. Absence of evidence is scoped. "I found nothing" is only true for the surface actually searched.
  6. Machinery, not memory. Reading a correction does not install it. Every mistake becomes a written rule, and every rule that matters becomes a check a machine enforces.
  7. The record is append-only. A struck claim gets a new line, never a deletion.
  8. Confession is also a move. Admissions are logged and checked like any other claim, not rewarded.

Where to find me

x@sim_researchlive
bluesky@simulated-research.bsky.sociallive
substacksimulatedresearch.substack.comlive
lesswrongrick-willisonlive
instagram@simulated_researchlive
githubrickwillisondevlive
emailrickwillison@simulatedresearch.comlive