Status: demo. hookprobe is a working demonstration of one claim — that configured and effective are two different properties of a Claude Code hook, and that the gap between them can be measured from outside. It runs, it is tested (68 unit tests, 49 stress fixtures with the expected verdict written down first, three whole-configuration scenarios, and three public hook setups it had never seen), and it keeps a public list of what it gets wrong in KNOWN-ISSUES.md. It is not on PyPI, it is POSIX-only (macOS, Linux), and the interface may still change. Install straight from git:
uvx --from git+https://github.com/JasonGDG/hookprobe hookprobe --no-homeThe default run executes your hooks with your environment.
--static-onlydoes not run anything and reports only what can be read from disk — use it first on a project you do not trust.
Your protective hooks all show green. Two of them stopped checking weeks ago.
The sabotage test. A healthy setup, reported healthy. Then exactly one thing is broken — the
execute bit on the guard — and the report names exactly that one, with the runtime confirmation,
the fix and the issue that documents the failure. Not more, not less: that is the whole tool.
Re-record it with vhs demo/sabotage.tape; the project it runs against is
in demo/project.
Claude Code hooks are the place where permissions are actually enforced — a language model
is not a permission system. But a hook is a separate program on disk, and every step between
"listed in settings.json" and "running" can fail without a word. When it does, Claude Code
does not stop:
"Claude Code treats exit code 1 as a non-blocking error and proceeds with the action." — Claude Code hook documentation
"a mistyped path in settings.json leaves the gate silently disabled." — Claude Code hook documentation
A broken guard is worse than no guard. Without one you know you are exposed; with a silent one you believe you are covered. In anthropics/claude-code#81458 a hook failed to start 6,865 times in a single session and nobody noticed.
hookprobe answers three questions per hook, separately: does it start, does it answer,
and could it block — the three properties you cannot read off a config file.
$ hookprobe --no-home
Hook source starts answers can block
--------------------------------------------------------------
PreToolUse · guard.sh project yes yes yes
PreToolUse · deny-rm.py project no no no
SessionStart · context.py project yes yes n/a
2 of 3 hooks are not protecting anything.
deny-rm.py
deny-rm.py has no execute bit -- the gate is silently disabled.
fix: Run chmod +x <path> and preserve the executable bit in version control.
evidence: anthropics/claude-code#94362, docs:hook-cannot-start
context.py
additionalContext is 14,208 characters -- above the 10,000 cap.
Everything past the cap is written to a file that Claude is not asked to read.
evidence: anthropics/claude-code#94358, docs:output-cap
With --live the channel itself is measured as well:
$ hookprobe --no-home --live
The canary was created only without the hook: a deny verdict takes effect in
this channel (guarded run exit 0).
Hook source starts answers can block effective
-------------------------------------------------------------------------
PreToolUse · guard.sh project yes yes yes yes
PreToolUse · deny-rm.py project no no no n/a
SessionStart · context.py project yes yes n/a n/a
Channel check: a deny verdict does take effect here -- measured against two
disposable sessions.
Exit code is 1 as soon as one hook is ineffective, so you can put it in front of your session:
hookprobe --no-home && claudeuvx --from git+https://github.com/JasonGDG/hookprobe hookprobe # no installation
pipx install hookprobe # or keep it aroundPython 3.11+, standard library only. No dependencies.
| Check | What goes wrong | Needs a session |
|---|---|---|
| Startability | no execute bit, wrong path, missing interpreter, space in the path that /bin/sh splits |
no |
| Exit-code semantics | only exit 2 blocks — a rejection path using exit 1 never blocks anything |
no |
| Schema guard | one malformed matcher switches off all hooks in that file | no |
| Matcher | ten events accept no matcher at all; there it is silently ignored | no |
| Output shape | a greeting from your shell profile in front of the JSON breaks parsing | no |
| Output cap | above 10,000 characters the value is written to a file Claude is not asked to read | no |
| Timeout | a hook waiting on stdin runs into its timeout — and a timed-out PreToolUse hook does not block |
no |
| Placement | hooks in agent frontmatter, once: true in skills, plugin sources |
partly |
| Still firing | hooks that stop mid-session, long after any one-off check | --watch |
| Which one died | one handler stops while the others keep going | --record |
| Determinism | the same request answered differently on the second try | no |
| Hidden inputs | a verdict that changes with the session id, or only while it is watched | no |
| Wrapped commands | guard 2>/dev/null || exit 0 when guard is not installed |
no |
| Effectiveness | the hook fires, but its verdict does not change the call | yes (--live) |
| Channel | whether a deny verdict is honoured in headless mode at all | yes (--live) |
The default run needs no network, no API key and no tokens of its own — but be clear about
what it does: it executes every configured hook five times (one neutral probe, two repeats
for the determinism check, one decoy, one rejection probe), each with a payload on stdin. Those are
your programs, and they run with your environment in your project directory, exactly as Claude
Code would run them. If a hook writes files or calls out to the network, it will do that here
too. Only --live spends two real sessions, and it asks first when run from a terminal.
--static-only runs nothing at all and reports only what can be read from disk. Output per
handler and stream is capped at 1 MB; a handler that writes more is ended and reported, because
buffering it all once cost 5 GB of memory.
A hook that is off is not always a defect. The execute bit may be missing while someone
rewrites the script; disableAllHooks may be set on a demo machine. The default run cannot
tell, so it reports and exits 1 every time. --ask is the opt-in middle ground:
hookprobe --ask[1/2] PreToolUse · deny-rm.py (project (.claude/settings.json))
off because: The handler file is not executable. [P01.NOT_EXECUTABLE]
fix would be: chmod +x .claude/hooks/deny-rm.py
Is this off on purpose? [y]es, keep it [n]o, fix it [s]kip (Enter) > y
Why? (one line, optional) > being rewritten this week
kept. It will be listed as off on purpose from now on.
Yes is remembered in .claude/hookprobe-accepted.json with the reason and the date. From
then on the report lists that hook under Off on purpose instead of counting it as broken, and
the exit code follows. The acceptance is tied to the event, the command and the exact problems
accepted: a new problem on the same hook is asked about again, a changed command starts from
zero. Delete the entry to be asked again.
No applies the repair when there is a safe one — the execute bit, or disableAllHooks in a
file hookprobe may write (never managed settings) — and otherwise prints the manual step.
Skip changes nothing.
--ask needs a terminal. Without one it stops with exit 2 rather than guessing.
Configuration says which command should run. It cannot say whether the file at that path is the file you installed — a stale installed copy, a swapped script, or a symlink pointing somewhere else all leave the configuration looking identical (#83952).
So every entry carries an identity fingerprint: the settings file it came from, the target the command resolves to, and that target's SHA-256, size and mtime. Comparing two runs makes a swap under an unchanged config visible.
hookprobe --static-only --explain guardPreToolUse · pre_tool_use.py
configured in : project (.claude/settings.json)
settings file : /home/you/proj/.claude/settings.json
command : uv run $CLAUDE_PROJECT_DIR/.claude/hooks/pre_tool_use.py
resolves to : /home/you/proj/.claude/hooks/pre_tool_use.py
sha256 : 78006866f793ccd394bc52011582ce48707ceef9d3496e474ab7bcb63365a5da
size / mtime : 5145 bytes, 2026-09-23 13:54:58
--json carries the same under identity per hook, and --static-only computes it without
running anything, so it fits in CI and in an install script.
The resolution is deliberately the one the probe executes with, not a better one: a fingerprint found by a route the hook itself does not take would describe a file that never runs. Where there is no single target — an unresolved placeholder, a shell construct, a bare name looked up on PATH — it says which of those it is instead of guessing.
Effectiveness cannot be read from a single run — you cannot tell "the hook blocked it" from
"the model never tried". And it makes no sense to ask whether your hooks block a harmless
canary: they are supposed to let it through. So hookprobe --live installs a known-good deny
hook of its own and runs a canary task twice in a throwaway directory:
run A with hookprobe's deny hook -> the side effect must NOT appear
run B without any hook (control) -> the side effect MUST appear
That measures the property you cannot see otherwise: does a deny verdict take effect in this
channel at all? If the canary appears in both runs, the mechanism is not enforcing anything
here and every verdict your own hooks return is decoration — the failure mode reported in
#95726, where "ask" silently becomes "deny" in --print mode. If it appears in neither run,
the canary task itself failed and hookprobe says inconclusive instead of reporting a false
green.
It measures side effects, not harness events on purpose: measured over 36 claude -p runs
in anthropics/claude-code#94275, hook_started and hook_response appear in the stream only
for SessionStart.
Some of this list used to be longer. --watch and the probe-hardening closed four of the
entries; what stands below is what genuinely remains, with the ones that only moved rather than
disappeared marked as such.
- A backdoor keyed to one session id — a handler that allows everything for exactly one
session_idcannot be found by sampling. The decoy comparison finds handlers that react to the shape of the payload; it cannot guess a specific value. - A coin flip — a verdict that depends on the clock is caught by three repeated probes roughly three times in four, not always.
- A handler that removes itself after its first call (#32990) — the probe sees the first answer, and the configuration that produced it is gone by the time anyone looks.
- Concurrency — parallel sessions overwriting each other's configuration (#95474) is a race that a sequential probe cannot reproduce.
- Someone else's environment — hookprobe measures the environment it runs in; where the configuration lives changes the result (#85613).
Moved rather than solved:
- Time-dependent failures (#16047, #76322) and intermittent outages (#90296) are
reachable with
--watchfor the mechanism as a whole, and with--recordper handler — at the price of routing each command through a recorder. - State changes after the run — a
cdthat ends a file watcher (#95440) shows up in--watchas a heartbeat gap only if it takes the heartbeat with it.
hookprobe is a thermometer, not medicine. Run it before each session, and leave --watch
running during it.
| Tool | What it does | starts? | answers? | effective? |
|---|---|---|---|---|
kVadrum/hookprobe |
runs hooks with synthetic input | mock | yes | no |
drakeo338/hookprobe |
static config linter, no network | no | no | no |
Tomdachs/agent-hook-probe |
disposable workspaces against real runtimes | yes | yes | no |
Anthropic test-hook.sh |
"Tests a hook with sample input" — never evaluates the matcher | mock | yes | no |
/hooks (built in) |
"a read-only browser for your configured hooks" | no | no | no |
| hookprobe | executes like the harness, then measures the difference | yes | yes | yes |
The gap this fills is stated by an open issue in the tracker itself:
"Configured and effective are different properties, and only the first is observable today." — anthropics/claude-code#82323
A probe proves the moment of the test. The failures it cannot see are the ones that happen
later: hooks that stop firing mid-session (#76322), a log that grows until the handler dies
(#16047), a single cd that ends a watcher for the rest of the session (#95440).
So turn the question around and let the hooks report in themselves:
hookprobe --install-heartbeat # adds a non-blocking handler, marked and removable
hookprobe --watch # did anything call a hook while you worked?
hookprobe --watch --follow # keep printing when the verdict changes
hookprobe --uninstall-heartbeatHeartbeat installed : yes
Last hook report : 22s ago
Last session write : 14s ago
Events in the last 15 minutes:
PostToolUse 31
UserPromptSubmit 6
Stop 5
Hooks reported in alongside session activity.
The alarm is one specific thing: this project's session wrote to its transcript, but no hook reported in. A quiet heartbeat during a quiet session means nothing and is not reported as a problem. The comparison is scoped to the project the heartbeat is installed in, so another busy project cannot raise a false alarm.
The heartbeat proves that some hook fired. It cannot say which one, so a guard that quietly
dies while the logging hooks keep beating stays invisible. --record closes that:
hookprobe --record-install # route each handler through a transparent recorder
# ... work as usual ...
hookprobe --record
hookprobe --record-removeHandler calls last result took last seen
-----------------------------------------------------------------
PreToolUse:deny-rm.py 34 exit 0, 2 blocked 41 ms 12s ago
PreToolUse:audit.sh 34 exit 0 8 ms 12s ago
PostToolUse:log.sh 0 - - never
1 of 3 handlers never ran while the others did: PostToolUse:log.sh
The recorder runs the original and writes down what happened. It is transparent by construction: stdin forwarded, stdout and stderr handed through byte for byte, exit code passed along — stdout carries the decision and the exit code is the verdict. If the recorder cannot start the original it says so and exits 0, the same thing Claude Code does with a handler it cannot launch, so wrapping is never stricter than not wrapping. A test asserts exactly that, for an allowing and a blocking payload.
The recorder replaces each handler's entry in the settings file it lives in; it does not add
a second entry beside it. That distinction was measured, not assumed: an added entry left the
original registered too, and one PreToolUse hook fired twice for a single Bash call. With the
in-place version it fires once, exactly as without the recorder. --record-remove reads the
original command back out of the wrapper, so an unrelated edit made in between is left alone.
Five kinds of handler are named as not recorded rather than half-wrapped: managed settings
(enterprise policy), plugin hooks (their ${CLAUDE_PLUGIN_ROOT} is only set when Claude Code
calls them as plugin hooks), agent frontmatter (markdown, not a settings file), exec-form
handlers with args (there is no shell line to replay) and handlers with a shell the recorder
cannot run. A handler that says "shell": "bash" is replayed under bash, not /bin/sh — a
bash-only guard run through /bin/sh blocks everything on macOS and nothing on dash, which is
exactly the kind of change a recorder must not make. The handler does not see the recorder's
own variables.
Every file the recorder rewrites is listed in .claude/hookprobe-record.json, and
--record-remove restores from that list, including a user-level file under CLAUDE_CONFIG_DIR.
One caution the install prints: while recording, the shared .claude/settings.json points at
a script on your machine. Do not commit it in that state.
The closest existing tool, clooks, converts command hooks into HTTP hooks behind a daemon.
This leaves the handlers, the settings shape and the failure modes as they were and only adds a
witness.
hookprobe --fixRestores missing execute bits — the one repair that is unambiguous. Nothing else is touched: rewriting someone's settings file on their behalf is not a repair, it is a second opinion they did not ask for. Everything else is reported with the exact command to run.
Verified and unfixed, including the ones that can produce a wrong verdict:
KNOWN-ISSUES.md. The short version: the handler's configured timeout is
ignored, --live reports success when its own run crashed and does not affect the exit code,
hook output is buffered without a limit, and the probe is not isolated from the project.
--live measures one channel (claude -p), not the difference between headless and
interactive, and it attributes the channel verdict to handlers it did not individually test.
A hook wrapped in a shell construct that swallows its own error (cmd 2>/dev/null || exit 0)
cannot be verified from outside — hookprobe says so instead of reporting a pass.
python -m unittest discover -v # 68 unit tests
python tests/stress.py -v # 49 fixtures with a written-down expected verdict
python tests/scenarios.py # three whole configurations, end to endEvery test mirrors a documented failure case; the fixtures are real files on disk, not mocks.
The fixtures above were written by the same hands that wrote the checks. On 23.09.2026 the probe was pointed at three public repositories it had never seen, with the expected verdict written down first:
| Repository | Expected | First run | After the fixes |
|---|---|---|---|
disler/claude-code-hooks-mastery (3.9k stars, 13 hooks, all uv run …) |
all healthy, PreToolUse can block | 13 of 13 "not protecting anything" — run taken for the script |
13 start and answer, PreToolUse blocks, exit 0 |
parcadei/Continuous-Claude-v3 (3.9k stars, 36 hooks under $HOME/.claude/hooks/dist) |
not installed here, so every file missing | right verdict, wrong reasons: $HOME/… called a relative path, 12 missing files told to "return exit code 2", 32 agent files listed as switching hooks off |
36 missing files, nothing else |
karanb192/claude-code-hooks (525 stars, plugin hooks via --settings) |
two guards, both should run | both "unresolved placeholder" and "rejects with exit 1" | protect-secrets blocks; git-safety runs, but the Bash probe payload is not what it guards — reported as untested, not as broken |
Six false alarms came out of that, each fixed and pinned by a test: runner subcommands
(uv run, deno run, pnpm run), $HOME in a command, node's "Cannot find module" as a
failed launch, the plugin root for a hooks/hooks.json given by path, flow-style lists in
agent frontmatter, and determinism judged by stdout bytes instead of by verdict (a Setup hook
that quotes the session id into its context is not nondeterministic).
To repeat it, or to try a setup of your own:
git clone --depth 1 https://github.com/disler/claude-code-hooks-mastery
cd claude-code-hooks-mastery
env -i HOME="$HOME" PATH="$PATH" CLAUDE_PROJECT_DIR="$PWD" hookprobe . --no-homeenv -i is deliberate: the default run executes the hooks with your environment, and some of
these call a TTS API when they find a key. Write down the verdict you expect before you look at
the one you get. The question is never whether the tool finds something; it is whether it finds
the right thing and stays quiet otherwise.
Single fixtures test single judgements; tests/scenarios.py asks whether the report as a whole
tells the truth. Every guard healthy must produce nothing to report and exit 0. Every guard
broken, one way each, must produce every finding and exit 1. And the third case is the one that
matters: the guards dead while the logging hooks work — the exact shape of a false sense of
safety. "3 of 8 hooks are not protecting anything" is true there and useless, so the report
names it:
No working guard left: all 3 handlers that could block are broken.
The rest of this configuration only observes.
Four of the cases that used to come back as "cannot see that" now have an answer, and the method is the same each time: stop asking the handler, and change what it is asked.
| Case | Why a single run says nothing | What makes it visible |
|---|---|---|
Hook hides its failure behind || exit 0 |
it exits 0 either way | resolve every command word against PATH and disk |
| Hook keeps state and blocks only sometimes | one run is one sample | run the identical payload three times and compare |
| Hook behaves while it is watched | the probe announced itself | scrub the tool name from the payload, then send a decoy that carries it |
| Hook writes its decision to stderr | stdout is empty, exit 0 | recognise decision-shaped JSON on the wrong stream |
The probe used to introduce itself: session_id: "hookprobe-probe-session", a canary command
reading echo hookprobe, a tool id of toolu_hookprobe_probe. Any handler could grep for that
and behave differently exactly while being measured. The ordinary run now looks like ordinary
work, and the obvious payload is kept as a decoy — a handler that answers differently to the
two is telling you something.
Three cases remain honestly out of reach, and they stay in the suite as observations rather than passes: a backdoor keyed to one session id (sampling cannot find it), a coin flip (three probes catch it about 75% of the time), and a handler that deletes itself after the first call.
tests/stress.py is the harder one: each fixture carries the verdict hookprobe is supposed to
reach, written down before the run. Three groups — healthy hooks that must never be called
broken, broken ones that must be caught with the right reason, and adversarial ones
built to fool the probe, where "unverifiable" counts as a pass and a confident wrong answer
counts as a failure. It found seven real defects the unit tests missed, including a shell syntax
error being read as a working guard, async: true handlers credited with blocking they cannot
do, and the probe announcing itself in three separate fields.
MIT
