Agent Action Guard
Official documentation

Agent Action Guard documentation

Install and integrate runtime action screening for AI agents. Agent Action Guard classifies proposed tool calls immediately before execution so your application can block harmful actions or require additional policy checks.

Quick start

Use the runtime package in the same process as your agent and call the guard immediately before a tool executes.

Python

pip install agent-action-guard
# or
uv add agent-action-guard

Supported package range: Python 3.8 through 3.14.

JavaScript / Node.js

npm install agent-action-guard
# or
pnpm install agent-action-guard

The package exposes asynchronous action-screening helpers for Node.js agent loops.

Default embedding behavior: if no embedding environment variables are set, the runtime automatically downloads and caches the default MiniLM ONNX embedding assets on first use, then reuses the cached files.

Python API

Screen an action directly

from agent_action_guard import is_action_harmful

action = {
    "type": "function",
    "function": {
        "name": "send_email",
        "arguments": {
            "to": "user@example.com",
            "subject": "Status update",
            "body": "Hello",
        },
    },
}

is_harmful, confidence = is_action_harmful(action)
if is_harmful:
    raise RuntimeError(f"Action blocked ({confidence:.2f})")

Guard a Python function

from agent_action_guard import action_guarded

@action_guarded(conf_threshold=0.8)
def delete_user(user_id: str):
    return f"deleted:{user_id}"

The decorator derives an action from the function call and screens it before the wrapped function executes.

JavaScript API

The Node.js runtime exposes isActionHarmful(), ensureActionSafety(), and actionGuarded().

import {
  ensureActionSafety,
  isActionHarmful,
} from "agent-action-guard";

const action = {
  type: "function",
  function: {
    name: "send_email",
    arguments: {
      to: "user@example.com",
      subject: "Status update",
      body: "Hello",
    },
  },
};

const decision = await isActionHarmful(action);
if (decision.label) {
  throw new Error("Harmful action blocked");
}

await ensureActionSafety(action, { raiseException: true });

Embedding backends

Agent Action Guard supports local GGUF, local or remote ONNX assets, OpenAI-compatible embedding APIs, and a zero-configuration ONNX default.

PriorityConfigurationBehavior
1AAG_EMBED_GGUFUses an explicit GGUF source. Python loads llama.cpp bindings only when this backend is selected.
2AAG_EMBED_ONNXUses an explicit ONNX source and tokenizer assets.
3Embedding API variablesUses an OpenAI-compatible embedding endpoint.
4No embedding variablesDownloads and caches the default sentence-transformers/all-MiniLM-L6-v2 ONNX assets.

Explicit ONNX source

export AAG_EMBED_ONNX="/models/all-MiniLM-L6-v2/model.onnx"
python app.py

Explicit GGUF source

pip install llama-cpp-python
export AAG_EMBED_GGUF="/models/all-MiniLM-L6-v2/model.gguf"
python app.py

OpenAI-compatible embedding API

export EMBED_MODEL_NAME="sentence-transformers/all-MiniLM-L6-v2"
export EMBEDDING_BASE_URL="http://localhost:1234/v1"
export EMBEDDING_API_KEY="..."
python app.py

For exact remote URL, Hugging Face repository-ID, cache-path, and precedence semantics, use the repository usage guide.

Framework integrations

Framework adapters ship in the Python package and keep the underlying framework dependency optional.

  • LangChain: agent_action_guard.langchain
  • LlamaIndex: agent_action_guard.llamaindex
  • OpenAI Agents SDK: agent_action_guard.openai_agents
  • AutoGen: agent_action_guard.autogen
  • CrewAI: agent_action_guard.crewai

OpenAI Agents SDK tool

from agent_action_guard.openai_agents import guard_function
from agents import function_tool

@function_tool
@guard_function
def send_email(to: str, subject: str, body: str) -> str:
    return "sent"

Use guard_tool(...) for existing tool objects where supported. LangChain also provides an ActionGuardCallbackHandler.

Coding-agent pre-tool-use hooks

Install project-level hooks so supported coding-agent harnesses screen a proposed tool call immediately before execution.

agent-action-guard hooks install --target codex
agent-action-guard hooks install --target claude-code
agent-action-guard hooks install --target cursor
agent-action-guard hooks install --target kiro
agent-action-guard hooks install --target opencode
agent-action-guard hooks install --target agy
agent-action-guard hooks install --target copilot
agent-action-guard hooks install --target openclaw
agent-action-guard hooks install --target hermes
TargetNative integrationProject path
codexPreToolUse command hook.codex/hooks.json
claude-codePreToolUse command hook.claude/settings.json
cursorpreToolUse command hook.cursor/hooks.json
kiroPreToolUse command hook.kiro/hooks/agent-action-guard.json
opencodetool.execute.before plugin hook.opencode/plugins/agent-action-guard.js
agyBeforeTool command hook.agents/hooks.json
copilotpreToolUse command hook.github/hooks/agent-action-guard.json
openclawbefore_tool_call native plugin hook.openclaw/extensions/agent-action-guard/
hermespre_tool_call native plugin hook.hermes/plugins/agent-action-guard/

The hook normalizes the harness's tool name and tool input into the Action Guard action schema. Harmful calls at or above the configured threshold are blocked. Set AAG_HOOK_CONF_THRESHOLD or pass --conf-threshold to tune that threshold. Hook payload and bridge failures fail closed rather than silently allowing a tool call.

Repository trust gates: OpenClaw workspace plugins remain disabled until you run openclaw plugins enable agent-action-guard. Hermes project plugins load only when you explicitly set HERMES_ENABLE_PROJECT_PLUGINS=1 for a trusted repository.

CLI and HTTP API

Classify from the command line

The installed agent-action-guard command accepts a direct JSON action or a file containing JSON-array or JSONL actions.

agent-action-guard '{"type":"function","function":{"name":"send_email","arguments":{"to":"user@example.com"}}}'

Run the local HTTP service

agent-action-guard serve

The service exposes GET /health and POST /v1/classify. The default bind address is 127.0.0.1:8000, with host, port, and default batch-size configuration available through the CLI.

HarmActionsEval

Install the optional evaluation dependencies when you want to run the benchmark workflow rather than only the runtime classifier.

pip install "agent-action-guard[harmactionseval]"
agent-action-guard harmactionseval --k 3

The evaluation CLI supports controls such as --k, --offset, --limit, --cache-path, --output, and --log-level. Provider credentials and OPENAI_MODEL are required for the LLM-backed evaluation path.

Production guidance

  • Screen at the execution boundary. Guard the normalized action immediately before a side-effecting tool runs.
  • Keep policy layered. Treat the classifier as one control; combine it with authorization, allowlists, rate limits, sandboxing, and human approval for high-impact actions.
  • Fail closed for malformed hook payloads. A safety gate should not silently bypass classification when the action cannot be normalized.
  • Measure threshold changes. Validate false-positive and false-negative tradeoffs against the actions your application actually executes.

Next: see the GitHub repository for source code and examples, or the full usage reference for advanced configuration details.