Security for agents that act on real systems.
GuardClaw is runtime security for developers building tool-using agents. Policy checks on every tool call, seven layers of defense, and a tamper-evident receipt for every decision, running locally on your hardware.
01 / Where GuardClaw sits
We watch the agent act.
AI security gets checked at three moments: the code before it ships, the text passing through the model, and the runtime. GuardClaw works at the runtime, where an agent actually invokes tools, reads files, and calls out.
Code & dependency scanning
Catches vulnerabilities in your source. It runs before the agent does.
Model input / output filtering
Checks prompts in and responses out, at the model's edge.
Runtime agent security
Shell commands, file reads, edits, and web fetches the agent makes are checked against policy in the fast path, before they run.
02 / Try it
Send an attack. Watch it get denied.
Pick a hostile action an agent might attempt. GuardClaw runs it through all seven layers and stops it at the first that catches, then writes a signed receipt.
03 / Defense in depth
Seven layers. Each assumes the others can fail.
One filter that catches 95% misses 5 in 100. Seven layers, each built to catch what the ones before it missed, cut that a long way down. We publish the numbers we actually measure rather than a compound figure we cannot test.
Threat Intelligence
Known-threat detection from compiled patterns and live feeds. A vector one team reports protects everyone.
Input Validation
Catches prompt injection, data leakage, and malicious payloads before they reach your agent.
Policy Enforcement
Deny-by-default. Every action is checked against explicit rules: allow, deny, or escalate.
Capability Tokens
Short-lived, signed, single-use tokens scoped to one approved action. Replays fail automatically.
Sandboxed Execution
On macOS, supervised runs execute inside an OS sandbox profile that limits which files and network destinations the agent can reach. Linux runs get process isolation today, and filesystem confinement is in progress. Windows is not sandboxed yet.
Human-in-the-Loop
High-risk operations pause for approval, cryptographically bound to the exact request.
Receipt Chain
A tamper-evident, cryptographically linked audit trail. Every decision, reconstructable.
04 / What defines it
Three decisions, made on purpose.
Local-first
The detection engine runs on your hardware. Your code and your agent's actions are not sent to us or to any third party for scoring. Receipt metadata syncs to the cloud only if you turn that on.
Deterministic
Compiled pattern matching and rule-based policy in place of a model deciding. Ask why an action was blocked and you get the rule, the pattern, and the score that did it, recorded in a receipt. Given the same configuration and pattern set, the same input gets the same answer.
Defense in depth
Seven independent layers, each built assuming the others can be bypassed. Compound coverage is the only thing that holds against creative attackers.
05 / Drops into your stack
Works with the agents you already run.
06 / FAQ
Questions, answered.
What exactly does GuardClaw secure?+
The runtime. After the model produces a plan, GuardClaw checks each concrete action the agent takes, including shell commands, file reads, network calls, and tool invocations, against your policy before any side effect happens. Turn on deny-by-default and anything outside your allowlist is refused.
At what point does GuardClaw check an agent?+
At runtime, while the agent is acting. Code scanning happens earlier, on code at rest. Model filtering happens at the boundary of the model, on text going in and coming back. GuardClaw watches what the agent does once it starts calling tools, which is a separate moment and sits alongside the other two rather than replacing either.
Does my code or data leave my machine?+
Your code and your prompts stay on your machine, and security decisions are made locally with no network round-trip. Anonymous telemetry is on by default and carries decision counts, threat scores, timing, platform and version. It does not carry prompts, commands, file paths or personal data, and you can turn it off with guardclaw telemetry disable. Pattern updates check on a schedule you set. If you choose to sync receipts, a receipt carries the decision, the score, the tool name and action, and the file path or URL the tool touched. Your file contents and prompt text travel as hashes rather than as content. Free and paid tiers are treated differently for training, and the GuardClaw terms and the privacy notice are the authoritative statement of what is collected and what it is used for.
Is detection AI-based?+
No. GuardClaw uses compiled pattern matching (Bloom filters, Aho-Corasick, RE2) and rule-based policy. Given the same configuration and pattern set, the same input gets the same decision, and every decision is written to a receipt you can audit.
What are capability tokens and receipts?+
Every approved action gets a short-lived, signed, single-use token scoped to just that action; replays fail. Every decision is written into a cryptographically linked receipt chain, a tamper-evident record built for compliance and fast incident review.
Will it slow my agents down?+
Checks run in the fast path, locally, typically in well under a millisecond. Security that adds heavy latency gets disabled under pressure, so low-latency enforcement was a design requirement from day one.
Does it guarantee 100% detection?+
No tool does. GuardClaw publishes its rates (97%+ on our built-in corpus) and layers seven controls so a miss at one has another chance to be caught. You keep your security team and threat model. GuardClaw is infrastructure underneath them.
07 / Talk to us
Securing agents in production? Let's talk.
Questions
Integration, threat model, or deployment questions? We answer fast.
hello@takeinterest.ai ↗Press & research
Writing about agent security or runtime guardrails? We're glad to talk.
press@takeinterest.ai ↗Build with us
Want guardrails tailored to your stack? Tell us what you're running.
Book a call ↗Give your agents guardrails.
Deterministic, local, and layered from the first commit. Start free, check for pattern updates on a schedule you set, upgrade when you need the live feed.