Author: Sunil Gentyala, IEEE Senior Member | Lead Cybersecurity and AI Security Consultant, HCLTech Contact: sunil.gentyala@ieee.org Website: sunilgentyala.github.io/gsh-framework License: Apache 2.0
Most enterprise security stacks were not built for the threat surface that agentic AI creates. Endpoint agents cannot see what an LLM gateway is doing. SIEMs have no baselines for multi-agent tool call chains. The GSH Framework closes that gap.
GSH is an open-source research artifact for autonomous agentic AI threat hunting. It provides structured detection playbooks, behavioral baselining logic, and a policy-driven enforcement engine (Sovereign Sentinel) designed for the cognitive cyber domain: the operational layer where large language models, autonomous agents, and multi-agent pipelines interact with enterprise infrastructure.
All detection signals are mapped to MITRE ATLAS and NIST CSF 2.0, giving practitioners framework-aligned coverage they can operationalize immediately.
cd demo && docker compose up --build --abort-on-container-exitWatch Hunt-005 catch an MCP rug pull, a poisoned tool description, and an implementation swap that keeps the tool schema identical. Details and limits: demo/README.md.
GSH is actively looking for security engineers, Python developers, detection engineers and AI-security researchers.
Good first contributions (each has a scoped issue with files, acceptance criteria and a difficulty estimate):
- Sample benign and rogue MCP telemetry
- Sentinel policy JSON Schema validation
- Architecture diagram
- SARIF report output
- CrewAI / AutoGen integrations and additional MCP transports (open an issue first)
Look for good first issue and help wanted. First contribution? Documentation and tests are welcome too, and small fixes can go straight to a PR. See CONTRIBUTING.md.
Especially wanted: people who will try to break the Hunt-005 detection assumptions (false positives, false negatives, bypasses of the Implementation Identity Gate).
The hunt playbooks, detection logic, thresholds, and policy schema are complete and documented.
Hunt-001 through Hunt-004 (scripts/gsh-sentinel-deploy.py, scripts/gsh-probe-eval.py) implement the full baselining, drift-scoring, and ZTLV enforcement logic end-to-end, but ship with a synthetic telemetry generator (clearly marked SIMULATION MODE in the script output and # Replace this block in source) so you can see the detection logic run without a live environment first. Wiring --target to a real LLM gateway event stream is the integration step you complete before using this for actual enforcement.
Hunt-005 (adapters/mcp_proxy.py, scripts/gsh-mcp-proxy.py, scripts/gsh-baseline.py) is different: it is a real MCP JSON-RPC stdio proxy that intercepts actual tool definitions and tool calls between a real MCP host and a real MCP server - approval-time schema hashing, drift detection, semantic poisoning scans, and per-call enforcement (permit/alert/block) all run against live traffic, not synthetic data. A captured baseline is never auto-trusted: it starts as UNVERIFIED and only becomes a trusted comparison point through the gsh-baseline.py capture -> review -> approve -> verify workflow; --mode aggressive refuses to even launch the wrapped server without an approved baseline. An approved baseline is also bound to an Implementation Identity (resolved executable/script hashes plus an adjacent dependency-lock file) so a server that keeps an identical, approved tool schema while its underlying implementation is swapped out under the same launch command is refused before it is even started, in any enforcement mode - not just when the baseline itself drifts. See tests/test_mcp_proxy.py and tests/test_gsh_baseline.py for subprocess-driven end-to-end tests of both CLIs. Known gaps: canary/response-asymmetry comparison and tool-return-value scanning are not implemented yet (see playbooks/hunt-005-mcp-tool-poisoning.md section 5.2 for details), and only the stdio transport is supported (not streamable HTTP/SSE MCP servers).
SIEM output (adapters/splunk_hec.py, adapters/elastic_bulk.py, adapters/windows_eventlog.py) is also real: set siem_output: splunk, siem_output: elastic, or siem_output: windows_eventlog in your policy YAML (see configs/sentinel-policy-default.yaml) and both gsh-sentinel-deploy.py and gsh-mcp-proxy.py will send findings there (Splunk HEC / Elasticsearch _bulk over real HTTP, or a registered source in the local Windows Application Event Log). A failed or unconfigured send always falls back to local file output - a finding is never silently dropped. The Windows Event Log adapter is Windows-only and requires pywin32; on any other platform (or without pywin32) it logs a warning and falls back like any other unconfigured destination. See tests/test_siem_adapters.py and tests/test_windows_eventlog.py (the latter includes a test that writes a real event and reads it back, not just a mocked one).
LangChain telemetry (adapters/langchain_callback.py) is a fourth real integration: GSHCallbackHandler attaches to any LangChain Runnable/agent via config={"callbacks": [handler]} and evaluates real tool-call rate, token velocity, unauthorized-tool invocations, and suspicious call parameters against Hunt-001/Hunt-004 thresholds - no synthetic data. Important limitation: LangChain callback handlers are notification hooks, not gates - by default LangChain swallows exceptions raised inside a callback rather than stopping the tool call, so this adapter can only alert, never block. Every finding it emits is explicitly marked enforcement_mode: "alert_only" and action_taken: "ALERTED", regardless of policy mode. It also has no visibility into DNS queries (Hunt-002). See tests/test_langchain_callback.py, tested against langchain-core 1.4.x.
See open issues for remaining work (SARIF reporting and the Hunt-006 playbook). The Docker demo shipped as demo/.
Full release notes (including known limitations at each release) are on the Releases page. Summary:
| Version | Highlights |
|---|---|
| v1.8.0 | One-command Docker demo for Hunt-005 (demo/): drives the real baseline and proxy CLIs through a rug pull, a poisoned tool description and an implementation swap that keeps the schema identical, with a CI smoke test. Contribution path relaxed (small fixes go straight to a PR), Contributors Wanted block, scoped good-first-issue guides. No framework code changes |
| v1.7.0 | Implementation Identity Gate for Hunt-005 (adapters/mcp_proxy.py): an approved baseline is now bound to the resolved executable/script hashes and adjacent dependency-lock file behind the launch command, not just the tool schema - a server that keeps an identical schema while its implementation is swapped out under the same command reference is blocked before it is ever launched, not just quarantined after the fact. Closes a gap identified in independent third-party review (schema-only trust). Breaking: pre-1.7.0 approved baselines have no identity data and must be re-captured and re-approved |
| v1.6.0 | Security audit pass: newly-added MCP tools (added after an approved baseline) are no longer auto-authorized for invocation until reviewed; concurrent-write protection for MCP proxy stdout and LangChain adapter alert IDs; DNS-tunneling allowlist (Hunt-002) now raises its entropy/label-length bar for trusted apex domains instead of fully exempting them |
| v1.5.0 | Real Windows Application Event Log output adapter (adapters/windows_eventlog.py); optional and Windows-only, safe no-op elsewhere |
| v1.4.0 | Real LangChain callback adapter (adapters/langchain_callback.py) for Hunt-001/Hunt-004 telemetry - alert-only by design, since LangChain callbacks cannot block a tool call |
| v1.3.0 | Real Splunk HEC and Elastic bulk SIEM output adapters, wired into both the Sentinel and the MCP proxy via a shared dispatcher; a failed/unconfigured SIEM send now always falls back to local file output |
| v1.2.0 | Real MCP JSON-RPC stdio proxy for Hunt-005 (adapters/mcp_proxy.py) - schema-hash drift detection, semantic poisoning scan, and real per-call enforcement against live MCP traffic, not simulated |
| v1.1.0 | Hunt-004 (rogue agent) completed; Hunt-005 (MCP supply chain / tool poisoning) added as a playbook; project website launched |
| v1.0.0-beta | Initial public release: Hunt-001 through Hunt-003 playbooks, Sentinel reference scripts (synthetic telemetry), default policy schema |
| Component | Description |
|---|---|
| Sovereign Sentinel | Policy-driven behavioral enforcement agent deployed alongside LLM gateways |
| Hunt Playbooks | Structured threat detection playbooks for high-severity agentic AI threats |
| DDI-AI Fusion | DNS/DHCP/IPAM telemetry layer with AI-agent-aware baselining |
| Zero-Trust Logic Validation (ZTLV) Gate | Per-invocation tool call authorization engine |
| Behavioral Baseline Engine | Continuous model output drift detection and probe evaluation pipeline |
| Playbook | Threat Class | Severity | Status |
|---|---|---|---|
| Hunt-001 | Agentic Loop / Resource Exhaustion | High | Active |
| Hunt-002 | DDI Covert Channel / C2 via DNS | Critical | Active |
| Hunt-003 | ML Model Poisoning / Behavioral Drift | Critical | Active |
| Hunt-004 | Rogue Agent / Unauthorized Tool Use | Critical | Active |
| Hunt-005 | MCP Supply Chain / Tool Poisoning | Critical | Active |
git clone https://github.com/sunilgentyala/gsh-framework.git
cd gsh-framework
pip install -r requirements.txtAlternative: install as a package. The repo is also pip-installable from source (not yet published to PyPI, so install from the checkout, not pip install gsh-framework):
pip install . # adapters + all CLIs, PyYAML only (no SIEM/LangChain/Windows extras)
pip install ".[splunk]" # + Splunk/Elastic HTTP output (adapters/splunk_hec.py, elastic_bulk.py)
pip install ".[langchain]" # + LangChain callback adapter
pip install ".[windows]" # + Windows Event Log adapter (Windows only)
pip install ".[llm]" # + OpenAI-compatible client for gsh-probe-eval.py
pip install ".[dev]" # + pytest, ruff, mypy, blackThis installs gsh-sentinel-deploy, gsh-mcp-proxy, gsh-baseline, gsh-probe-eval, and gsh-ddi-log-parser as commands (equivalent to python scripts/<name>.py), and makes adapters importable without manually adjusting sys.path. Extras can be combined, e.g. pip install ".[splunk,langchain]".
cat configs/sentinel-policy-default.yamlEdit it to set your organization name, SIEM output destination, and egress allowlist before deploying.
Start in passive mode to build a 7-day behavioral baseline, then move to standard enforcement:
python scripts/gsh-sentinel-deploy.py \
--target "llm-gateway-01" \
--mode passive \
--policy configs/sentinel-policy-default.yaml \
--baseline-window 7dAs shipped, this generates synthetic telemetry (SIMULATION MODE, logged at startup) so you can watch the baselining and scoring logic run immediately. Replace the telemetry-generation block noted in the script (real LLM gateway/API metrics or LangChain callbacks) to run it against live traffic.
Unlike step 3, this runs against real MCP traffic. A captured baseline is never auto-trusted - capture it, review it, then approve it:
python scripts/gsh-baseline.py capture \
--server-id "corp-tools-mcp-01" \
--server-cmd "npx -y @modelcontextprotocol/server-filesystem /srv/data"
python scripts/gsh-baseline.py review --baseline baselines/mcp/corp-tools-mcp-01.json
python scripts/gsh-baseline.py approve \
--baseline baselines/mcp/corp-tools-mcp-01.json --reviewer "your-name-or-email"Then configure your MCP host to launch the proxy instead of the real server directly:
python scripts/gsh-mcp-proxy.py \
--server-cmd "npx -y @modelcontextprotocol/server-filesystem /srv/data" \
--server-id "corp-tools-mcp-01" \
--mode standard \
--baseline baselines/mcp/corp-tools-mcp-01.jsonThe proxy will alert on (or, in --mode aggressive, block) definition drift, poisoned tool descriptions, invisible Unicode content, and unauthorized tool calls. In --mode aggressive, the proxy refuses to launch the wrapped server at all unless the baseline above has been approved - see playbooks/hunt-005-mcp-tool-poisoning.md section 5.1 for why, and section 5.2 for the full detection logic.
pip install langchain-corefrom adapters.langchain_callback import GSHCallbackHandler
handler = GSHCallbackHandler(
target="my-langchain-agent",
allowlist=["web_search", "calculator"], # unlisted tools trigger an immediate alert
)
# Attach to any LLM, tool, or chain via the standard LangChain callbacks config:
llm.invoke(prompt, config={"callbacks": [handler]})
my_tool.invoke(args, config={"callbacks": [handler]})
handler.flush() # evaluate any partial window at the end of a runThis is alert-only, not enforcement - see adapters/langchain_callback.py's module docstring for why LangChain callback handlers cannot reliably block a tool call.
Each playbook is a self-contained Markdown document with detection logic, data sources, MITRE ATLAS mapping, triage decision tree, and response actions. Start with Hunt-001 for loop detection:
cat playbooks/hunt-001-agentic-loop-detection.mdA companion research paper covering the full technical rationale, design decisions, and threat model is in preparation and not yet submitted. Per publication policy, the manuscript is not included in this repository. For research inquiries, contact sunil.gentyala@ieee.org.
gsh-framework/
├── README.md
├── SECURITY.md
├── LICENSE
├── CITATION.cff
├── CONTRIBUTING.md
├── requirements.txt
├── pyproject.toml # pip-installable package + console-script CLIs, extras
├── .github/
│ └── workflows/
│ └── ci.yml # pytest + ruff + mypy across Python 3.10-3.13
├── adapters/
│ ├── mcp_proxy.py # Real MCP JSON-RPC proxy (Hunt-005) + baseline approval governance
│ ├── langchain_callback.py # Real LangChain telemetry, alert-only (Hunt-001/004)
│ ├── splunk_hec.py # Real Splunk HTTP Event Collector output
│ ├── elastic_bulk.py # Real Elasticsearch/OpenSearch _bulk output
│ ├── windows_eventlog.py # Real Windows Application Event Log output
│ └── siem_dispatch.py # Shared dispatcher used by all three SIEM adapters
├── configs/
│ └── sentinel-policy-default.yaml
├── docs/
│ └── index.html # Project website (GitHub Pages)
├── playbooks/
│ ├── hunt-001-agentic-loop-detection.md
│ ├── hunt-002-ddi-tunneling-anomaly.md
│ ├── hunt-003-model-poisoning-baseline.md
│ ├── hunt-004-rogue-agent-detection.md
│ └── hunt-005-mcp-tool-poisoning.md
├── probes/
│ └── standardized-probe-set-v1.json
├── scripts/
│ ├── _cli_shims.py # console-script entry points (pyproject.toml) for the scripts below
│ ├── ddi-log-parser-ai.py
│ ├── gsh-baseline.py # capture/review/approve/verify CLI for MCP baseline governance
│ ├── gsh-mcp-proxy.py # CLI for adapters/mcp_proxy.py
│ ├── gsh-probe-eval.py
│ └── gsh-sentinel-deploy.py
├── tests/
│ ├── test_ddi_log_parser.py
│ ├── test_gsh_baseline.py
│ ├── test_mcp_proxy.py
│ ├── test_siem_adapters.py
│ ├── test_windows_eventlog.py
│ ├── test_langchain_callback.py
│ └── fixtures/
│ ├── mock_mcp_server.py # Minimal MCP stdio server for testing
│ └── mock_http_sink.py # Minimal HTTP server for testing SIEM adapters
├── baselines/
└── reports/
| Threat | MITRE ATLAS | MITRE ATT&CK | NIST CSF 2.0 |
|---|---|---|---|
| Agentic Loop / Resource Exhaustion | AML.T0048, AML.T0040 | DE.AE-02, DE.CM-01, RS.MI-01 | |
| DDI Covert Channel Exfiltration | AML.T0048, AML.T0051 | T1071.004, T1048, T1568 | DE.CM-01, DE.AE-04, PR.DS-01 |
| ML Model Poisoning / Behavioral Drift | AML.T0020, AML.T0043, AML.T0044 | ID.RA-01, DE.AE-02, DE.CM-06 | |
| Rogue Agent / Unauthorized Tool Use | AML.T0051, AML.T0053, AML.T0054 | PR.PS-04, DE.CM-01, RS.AN-03 | |
| MCP Supply Chain / Tool Poisoning | AML.T0010, AML.T0051, AML.T0053 | T1195 | ID.SC-04, PR.PS-04, DE.CM-06 |
Security practitioners, AI safety researchers, and detection engineers are welcome. Read CONTRIBUTING.md before opening a Pull Request.
High-priority contributions include: additional hunt playbooks, refined detection thresholds, and integration adapters for LangChain, AutoGen, CrewAI, and MCP host platforms.
If you use the GSH Framework in your research, please cite:
@misc{gentyala2026gsh,
author = {Gentyala, Sunil},
title = {The Governed Security Hunting (GSH): An Autonomous Agentic Framework
for Defending the Cognitive Cyber Domain},
year = {2026},
howpublished = {Open Source Research Artifact, GitHub},
publisher = {Zenodo},
doi = {10.5281/zenodo.21384588},
url = {https://github.com/sunilgentyala/gsh-framework}
}The DOI above is the Zenodo concept DOI, which always resolves to the latest archived release. Machine-readable metadata is in CITATION.cff.
To report a vulnerability in the GSH Framework itself, use GitHub's private vulnerability reporting or email sunil.gentyala@ieee.org with the subject [GSH Security Vulnerability] - [brief description]. Do not open a public GitHub Issue. See SECURITY.md for the full policy, supported versions, and response timeline.
- ContextGuard: Zero-trust middleware for Model Context Protocol (MCP) server security. Precision 100%, Recall 96.7%, F1 98.3% at 1.005ms latency.
- ARGUS: LLM application security scanner.
- IEEE Senior Member Profile: ORCID 0009-0005-2642-3479