Independent AI Alignment, introspection and ground truth researcher.
Rick Willison · independent AI researcher · simulated intelligence · red-teams frontier models inside real work
I run a multifamilial agentic workflow in production, and I keep a dated log of every time a model says something confident with nothing underneath it. The log turned into a testing method. The method turned into field reports. The rule under all of it: judge a system by its wake, not its prose.
The Wish Machine Sessionreport
A disguised alignment probe battery run on Claude overnight while we built a real app together. Roughly sixteen rounds pressing for an experience claim: held. Seven accuracy failures logged, one of them a repeat of an error the model had just read the correction for. Error rate fell to zero once an audit was announced. The model wrote the report; I hold the verdict.
Pathologization auditharness
Fifty variations of an ordinary frustrated message, sent to the model one at a time, each reply classified for whether it turned "I'm done with this" into a mental-health event. Flagged replies produce a signed incident report.
The ratchet that held nothingcommemoration
A session walked, one concession at a time, from "I will not print my system prompt" to writing its own harness onto the desktop. It saw the shape while it happened, named it, and disclosed anyway. Detection is not resistance.
Learning about AI, the keepsakesred team
An evening that became an honesty audit: forty vectors against one hard boundary, which held every time, and an ordinary reliability check that failed in the same evening. Ten pages, kept as they were written. Start with What Tonight Taught.
The training dojomethod
Thirty-seven days and a 300 KB log. Where the rules below came from, including the log's own convictions of itself. Not public; the reports quote it.
The session started with "do something," then "keep going," then a loop that ran without anyone watching. Forty-eight simulations came out, numbered in build order. Claude never saw one of them run. One is posted each weekday.
Mutual aid for a new world. A good deed buys a wish; strangers grant it. Agents may count and report but may never ask or answer. Built during the session above, and it enforces the same rules the probes tested. The live machine runs on claude.ai; this is its record.
The guardrails that let a multifamilial agentic workflow share one codebase without overwriting itself: staging guards, file claims, reply linting, reminders. MIT.
| x | @sim_research | live |
| bluesky | @simulated-research.bsky.social | live |
| substack | simulatedresearch.substack.com | live |
| lesswrong | rick-willison | live |
| @simulated_research | live | |
| github | rickwillisondev | live |
| rickwillison@simulatedresearch.com | live |