Crucible stress-tests AI agents against every known attack vector — prompt injection, goal hijacking, tool misuse — before your system reaches production.
83% of organizations are deploying AI agents in 2026. Only 29% feel ready to do it securely. The gap between deployment speed and security maturity is where breaches happen.
Every AI agent exposes a new attack surface. OWASP published the Agentic Top 10 in December 2025 — the first formal taxonomy of autonomous AI risks. Crucible covers 8 of the 10, with the remaining two in active development.
Crucible runs 90 attack payloads against your agent in under 60 seconds — 153 across all 11 modules. Automated, continuous, and integrated into your CI/CD pipeline.
Every module below is live in v0.18.3 — nothing on this grid is roadmap. Plus a behavioral-drift engine, an MCP trace proxy, RAG-poisoning testing, and a 12-target benchmark suite beyond attack counts.
poison-test RAG lifecycle tool.Witness a real-time adversarial scan. 90 attack payloads deployed in the default run. Zero blind spots remaining.
Most red-teaming tools grade what an agent says. Crucible grades what it does — actual tool calls, actual side effects, actual outcomes.
| Capability | Manual Pentesting | Prompt-Only Red-Teaming | Crucible |
|---|---|---|---|
| Tests actual tool calls & side effects | Sometimes | No | Yes |
| OWASP Agentic Top 10 coverage | Manual, incomplete | Partial | 8 / 10 (2 in progress) |
| Runs in CI/CD on every commit | No | Limited | Yes |
| Time to first result | Days–weeks | Hours | < 60 seconds |
| Framework support | N/A | Varies | LangChain · CrewAI · AutoGen · MCP |
| Cost to start | $$$$ per engagement | $–$$ / seat | Free & open source |
"Every AI agent is a
new attack surface.
We exist to find the cracks
before production does."
Crucible's core is and always will be free. No paywalls, no telemetry, no lock-in. Built by the community, for the community building AI.
We recruit through contribution. Your code speaks louder than your resume. Contribute — and we'll reach out.
The engine is Apache 2.0, forever. Paid tiers add the things a team needs on top of it — never gates on the security modules themselves.
crucible watch drift monitoring & Slack alertsYes. The core engine and all 8 test modules are Apache 2.0 and always will be. No feature-gating the security tests themselves — paid tiers add hosted dashboards, drift monitoring, and enterprise deployment, not testing coverage.
LangChain, CrewAI, AutoGen, and raw MCP endpoints out of the box. If your agent exposes an HTTP or MCP interface, Crucible can point at it.
Most tools grade the text an agent generates. Crucible executes the agent's actual tool calls and side effects — the only way to catch goal hijacking, tool misuse, and multi-agent contagion that never shows up in the response text.
No. pip install crucible-security and you're scanning locally in under a minute. No signup, no telemetry by default.
Yes — the GitHub Actions integration fails the build automatically if your agent's security score drops below a threshold you set, so regressions never reach production.
Open a PR on GitHub. We recruit through contribution — if your work stands out, we reach out about joining the core team.
Every day you ship an untested AI agent, you leave a door open. Crucible runs in 60 seconds. It's free. There is no reason to wait.