Guide
Everything we publish comes out of a run that happened: an eval that failed, a hook that let something through, a case somebody else sent us. Each piece first appeared on our Substack and is repeated here on our own domain; the Substack copy stays the canonical one.
- I ran eight Claude Code skills through real evals. Three of them had bugs.
2026-09-14. Most Claude Code skill packs on Gumroad and GitHub are text files nobody has tested. You download 300 SKILL.md files, install them, and hope Claude follows the instructions. Sometimes it does. - My own guard hook fought my own skill. That is how I know the evals work.
2026-09-14. Yesterday I published Skillkeel Starter, a free plugin with three guard hooks and eight repo skills for Claude Code. Today the paid Kit is out. Between the two, I ran 24 more skills through Claude Code's built-in eval runner, and the most useful thing that happened was a failure. - A stranger found the hole in my guard hook the day it went public. Then they published 52 cases.
2026-09-20. On 14 September I read a post on r/ClaudeCode, "I tried 40 ways to make my coding agent rewrite its own guardrail hook". The author had written a PreToolUse hook that stops the agent from editing the hook's own file, then run the 40 against it. First run, 11 September: 29 of 40 refused, 11 through. After a directory and path-normalisation fix, 14 September: 38 refused, 2 through, an opaque script and a base64 command, which a text hook cannot read.
New pieces land on Substack first. Subscribe or take the RSS feed.