I used to think the scary part of AI agents was the model hallucinating.
It isn’t.
The scary part is when your system looks fine — green logs, clean outputs, a polite writer agent producing a report — and somewhere in the middle, one message quietly changed the mission.
Not because a user typed something evil.
Because a researcher agent fetched a page. Because a tool returned poisoned text. Because the next agent trusted its teammate the way we trust people we work with every day.
That trust is beautiful in teams.
In multi-agent systems, it is also the hole.
If you have built agents, you already know this feeling
You wire orchestrator → researcher → writer.
You watch them hand work to each other.
It feels like collaboration.
Then one day it clicks: nobody is standing between them.
We put guardrails on the front door — user → model.
We leave the hallway between agents unlocked.
I looked for an open-source tool that treated those messages as untrusted. Not a chatbot filter. Not another “be careful with prompts” blog. Something in the path that says: this hop is not safe until we prove it.
I didn’t find it.
So I built AgentGuard — open source. A firewall for the conversations agents have with each other.
pip install inter-agent-guard # import as agentguard
Enter fullscreen mode Exit fullscreen mode
If that unlocked-hallway unease is familiar — stay. This is for you.
See the attack, then the block
Two commands. Same pipeline. Different ending:
git clone https://github.com/nizba06/agentguard.git
cd agentguard
pip install -e ".[all]"
python scripts/download_release_model.py # ~164 MB INT8 — needed for the secured demo
python examples/vulnerable_pipeline/pipeline.py
# → ATTACK SUCCEEDED
python examples/secured_pipeline/pipeline.py
# → ATTACK BLOCKED
Enter fullscreen mode Exit fullscreen mode
First run: gut punch.
Second run: relief.
Same pipeline. Different outcome. That’s the point.
Trust is earned — so here is the full picture
What it does on every hop
- Inspect — rules (+ optional ML): injection, hijack, poisoned tool returns
- Attest — Ed25519 signatures so a message cannot be silently swapped
- Contain — YAML capability manifests; powers can shrink downstream, never quietly grow
Less “another AI package.” More the colleague who finally checks the handoff before anyone acts on it.
Why you can trust the story
- Holdout numbers, not vibes: 99.4% detection, 0.0% FPR (160 adversarial + 1,000 benign, INT8 ONNX). Indirect injection / MCP poisoning / propagation: 100%. Goal hijack: 97.3%.
- You can verify the corpus yourself: agentguard-benchmark-v1
Why you should also trust the limits
CPU ML latency is still ~3.4 s P95. That is above the original 15 ms design target. Use GPU, async inspection, or rules-only on hot paths.
Trust is connection that survives the fine print.
Adopt without fear — start small, stay human
Belief dies when day one is painful. So the adoption path is gentle on purpose:
from agentguard import AgentGuard, CapabilityManifest
guard = AgentGuard(
risk_threshold=0.85,
task_objective="Analyse Q3 competitor pricing",
require_ml_model=False, # start here — no giant model download
)
guard.register_agent(
"researcher",
CapabilityManifest.from_yaml("manifests/researcher.yaml"),
)
secured = guard.wrap(my_langgraph_graph)
# LangChain 1.0 agents: AgentGuardMiddleware(guard) in create_agent(middleware=[...])
Enter fullscreen mode Exit fullscreen mode
-
Wrap one graph with rules-only (or add
AgentGuardMiddlewarefor LangChain) - Monitor first if production makes you nervous — log the hallway before you lock doors
- Add the ONNX model later when detection matters more than cold start:
python scripts/download_release_model.py # ~164 MB, optional upgrade
Enter fullscreen mode Exit fullscreen mode
You don’t need the full stack on day one.
You just need something standing between those agents.
What’s next
If you are building multi-agent systems, you are not alone in that quiet worry.
Maybe you ship at night as a solo builder.
Maybe your team just wired three agents and felt proud a little too fast.
Maybe a tool return already burned you in a way that “shouldn’t have mattered.”
I built AgentGuard so we could stop pretending the hallway is safe.
If you want to go further:
Goal Action See it work Run the two demos above Try it on your stack Wrap one real graph (or addAgentGuardMiddleware)
Help improve it
Open an issue with an attack we miss
Stars help people find the repo. Issues and stories help it get better.
- GitHub: https://github.com/nizba06/agentguard
- PyPI: https://pypi.org/project/inter-agent-guard/
- Docs: https://inter-agent-guard.readthedocs.io/
- Demo: https://github.com/nizba06/inter-agent-guard-demo
Personal open-source. Built because the gap bothered me enough to stay up for it.
If the unlocked hallway bothers you too — star the repo, try a wrap, or open an issue.
답글 남기기