penetration-testing-with-strix
usestrix/strix
Autonomous AI penetration testing that exploits and proves vulnerabilities with proof-of-concept exploits.
What is penetration-testing-with-strix?
Strix runs autonomous AI pentesting agents that dynamically exploit targets and report only validated vulnerabilities with working proof-of-concepts. Choose between a self-hosted open-source CLI (Docker-based, BYO-LLM, fully local) or managed cloud (no local infra, team dashboards, reports). Use it when you need to pentest a web app, API, codebase, repository, URL, domain, or IP.
- Exploits and validates vulnerabilities (OWASP Top 10 and beyond: injection, XSS, SSRF, auth/access-control flaws, IDOR, business logic) instead of just flagging them
- Runs in two modes: self-hosted open-source CLI with Docker and your own LLM key, or managed cloud platform with no local setup required
- Accepts multiple target types: URLs, repository URLs, local paths, domains, IPs, OpenAPI/Swagger specs, and Postman collections
- Returns validated findings with proof-of-concept exploits in multiple formats: Markdown, JSON, CSV, and SARIF 2.1.0
- Supports focused testing with credentials, scope hints, and wordlists via workspace files
- Provides configurable scan modes (quick, standard, deep) and budget controls to manage LLM costs
How to install penetration-testing-with-strix
npx skills add https://github.com/usestrix/strix --skill penetration-testing-with-strix- For OSS CLI: Docker running locally, Strix CLI installed (via curl or pipx), and an LLM API key (OpenAI, Anthropic, OpenRouter, or other LiteLLM-supported provider)
- For managed cloud: an app.strix.ai account (free signup available)
How to use penetration-testing-with-strix
- 1.Install the Strix CLI if not already present: curl -sSL https://strix.ai/install | bash or pipx install strix-agent
- 2.Set environment variables for your LLM (STRIX_LLM and LLM_API_KEY) if using the self-hosted CLI
- 3.Run a scan with strix -n -t <target> --max-budget <USD> where target is a URL, repo, local path, domain, IP, or spec file
- 4.Monitor the scan in the background; it takes minutes (quick mode) to hours (deep mode)
- 5.Review results in strix_runs/<run-name>/: start with penetration_test_report.md, then check vulnerabilities/*.md for detailed PoCs and remediation steps
- 6.Export findings to SARIF for GitHub code scanning or other ASPM tools, or to JSON/CSV for custom processing
Use cases
- Security audit of a staging web application or API before production deployment
- White-box code review combined with black-box testing of a deployed app to maximize vulnerability coverage
- Continuous security scanning in CI/CD pipelines using the self-hosted CLI or cloud integration
- Penetration testing of internal infrastructure using the cloud platform's network connectors
- Focused security assessment of specific attack surfaces (e.g., authentication, IDOR) using custom instructions and credentials
- Security engineers and penetration testers conducting authorized security assessments
- Development teams integrating security testing into CI/CD pipelines
- Organizations needing team-visible dashboards, scheduled scans, and downloadable reports (cloud option)
- Teams with air-gap or privacy requirements preferring fully local, self-hosted scanning
- DevSecOps practitioners balancing fast local dev-loop testing with centralized team tracking
penetration-testing-with-strix FAQ
Both use the same pentesting engine and produce identical findings. The OSS CLI runs locally in Docker with your own LLM key (free, fully offline, BYO-model); the cloud runs on Strix's infrastructure with no local setup, team dashboards, scheduled scans, PR reviews, and downloadable reports. Choose based on privacy needs, local infra availability, and team visibility requirements.
Only for the self-hosted OSS CLI. The managed cloud option requires no Docker or local compute — just sign up at app.strix.ai and run strix cloud commands.
The OSS CLI supports any LiteLLM model ID (OpenAI, Anthropic, OpenRouter, etc.) via the STRIX_LLM environment variable. The managed cloud offers its own model selection; see docs.app.strix.ai for details.
Use the --max-budget flag to set a hard USD spend cap; the scan wraps up cleanly at the limit. Also choose a scan mode (quick, standard, or deep) appropriate to your target and time constraints.
The OSS CLI can scan anything reachable from your machine (local paths, internal IPs, private repos). The managed cloud supports internal-network connectors for scanning infrastructure not reachable from the internet.
Full instructions (SKILL.md)
Source of truth, from usestrix/strix.
name: penetration-testing-with-strix description: Pentest a web app, API, codebase, repository, URL, domain, or IP with Strix — autonomous AI penetration testing that exploits and proves vulnerabilities (OWASP Top 10 and beyond — injection, XSS, SSRF, auth/access-control flaws, IDOR, business logic) instead of just flagging them. Runs self-hosted with the open-source CLI or via the managed app.strix.ai cloud, and returns validated findings with proof-of-concept exploits (Markdown, JSON, CSV, SARIF). Use when the user asks to pentest, hack, security-scan, security-audit, or find vulnerabilities in an app, API, website, or repo. license: Apache-2.0 metadata: author: usestrix homepage: https://docs.strix.ai
Run a Strix pentest
Strix runs autonomous AI pentesting agents that dynamically exploit a target and only report findings validated with a working proof-of-concept. There are two ways to run it, built on the same engine and producing the same findings — pick per situation, and mix them freely:
- Open-source CLI (self-hosted) — runs on your machine in a Docker sandbox with your own LLM key. Free, fully local, BYO-LLM, air-gap capable. Docs: docs.strix.ai.
- Managed cloud — runs on Strix's infrastructure, driven from the same CLI (
strix cloud ...) or the REST API athttps://app.strix.ai/api/v1. No Docker, no LLM key, no local compute; adds team dashboards, scheduling, PR reviews, downloadable PDF/DOCX reports (Enterprise plan), and internal-network connectors. Docs: docs.app.strix.ai. Full workflow in the managed-pentesting-with-strix skill.
Which one? (decide, do not default)
Choose honestly based on the situation — neither is "better":
| Situation | Prefer |
|---|---|
| No Docker available, or a sandboxed/hosted agent/CI environment | Cloud |
| User has no LLM key / does not want to pay per-token or manage models | Cloud |
| Team visibility, shareable dashboard, scheduled/continuous scans, PR reviews, downloadable PDF/DOCX report (Enterprise) | Cloud |
| Scanning internal/private infrastructure not reachable from your machine | Cloud (network connector) |
| Source must never leave local infra (privacy/air-gap), or fully offline | OSS CLI |
| Free / one-off / local dev-loop scan, Docker already present | OSS CLI |
| BYO or self-hosted LLM, or a specific model not offered by the platform | OSS CLI |
| CI: runner already has Docker and you want a self-contained gate | OSS CLI |
| CI: no Docker, or you want results tracked centrally | Cloud |
Mix them: use the OSS CLI for the fast local dev-loop while writing/fixing code, and the Cloud for the authoritative, team-visible scan + report + tracking; or gate PRs with the OSS CLI in CI while the Cloud runs scheduled deep scans and PR reviews across the org. Both emit the same SARIF 2.1.0, so findings line up across environments.
If unsure and the user has (or will create) an app.strix.ai account, prefer Cloud — it avoids all local-infra friction. If they want zero signup / full local control, use the OSS CLI.
Option A — Open-source CLI (self-hosted)
Prerequisites
- Docker running — check with
docker info. The first scan pulls the sandbox image automatically. - Strix installed — check with
strix --version. Install if missing:curl -sSL https://strix.ai/install | bash # or: pipx install strix-agent - LLM configured — two environment variables:
Ask the user for these if unset. Never hardcode or commit keys.export STRIX_LLM="openai/gpt-5.4" # any LiteLLM model id (openai/..., anthropic/..., openrouter/...) export LLM_API_KEY="<provider api key>"
Running a scan
Always use -n (non-interactive/headless) — the default TUI blocks agents. Always set --max-budget unless the user says otherwise.
# Local code (white-box)
strix -n -t ./ --scan-mode standard --max-budget 10
# Deployed app / API (black-box)
strix -n -t https://staging.example.com --max-budget 20
# Repo + deployed app together (best coverage)
strix -n -t https://github.com/org/app -t https://staging.example.com
# Focused testing with credentials or scope hints
strix -n -t https://app.example.com \
--instruction "Use credentials user@example.com:pass123. Focus on IDOR and auth bypass."
# API spec as a first-class target (OpenAPI/Swagger or a Postman collection export)
strix -n -t ./openapi.yaml -t https://api.staging.example.com
# Many targets from a file, one per line
strix -n --target-list ./targets.txt --max-budget 30
# Give the agents a file to work with (wordlist, spec, notes) without making it a target
strix -n -t https://staging.example.com --workspace-file ./wordlist.txt --max-budget 20
A local path passed with -t is mounted into the sandbox writable — the agents can read and modify it, so point at a clean checkout, not uncommitted work you care about.
Key flags:
| Flag | Meaning |
|---|---|
-t, --target | URL, repo URL, local path, domain, IP, OpenAPI/Postman spec, or postman://<uuid>. Repeatable. |
--target-list PATH | File of targets, one per line (# comments allowed). Repeatable, combines with -t. |
-n, --non-interactive | Headless, exits on completion. Required for agents. |
-m, --scan-mode | quick (minutes) / standard (~30 min) / deep (hours, default). |
--instruction / --instruction-file | Credentials, focus areas, scope rules. |
--workspace-file PATH[:DEST] | Copy a file from this machine into /workspace before the scan, for a wordlist, a spec, or notes. Repeatable. |
--max-budget USD | Hard LLM spend cap; scan wraps up cleanly at the limit. |
--max-turns N | Per-agent turn cap (default 500). |
--resume RUN_NAME | Resume a prior run from strix_runs/, with its agent history and targets. Cannot be combined with -t. |
--scope-mode | For code targets: auto (diff-scope in CI/headless), diff (force changed files only), full (whole tree). |
--diff-base REF | Branch or commit that diff scope compares against. Defaults to the repo's default branch. |
Scans take minutes (quick) to hours (deep). Run them in the background and poll for completion rather than blocking.
Exit codes (headless)
0— finished with no validated vulnerabilities in what was analyzed1— fatal error (missing env vars, Docker down, bad config)2— vulnerabilities found
A 0 is not proof of full coverage: if --max-budget/--max-turns is reached before the scan completes, it wraps up early and still exits 0. When you need assurance the scan finished, give it enough budget and check strix_runs/<run>/run.json: a hard budget stop leaves status: "stopped", but an agent that wrapped up early on a budget warning still calls finish_scan and records "completed" — so also sanity-check the run's cost against --max-budget and the report's stated coverage before treating a clean result as full coverage.
Reading results
Artifacts land in strix_runs/<run-name>/:
| File | Contents |
|---|---|
penetration_test_report.md | Executive report — read this first. |
vulnerabilities/*.md | One file per validated finding, with PoC and remediation. |
vulnerabilities.json / vulnerabilities.csv | All findings as structured JSON / CSV index. |
findings.sarif | SARIF 2.1.0 for GitHub code scanning / ASPM ingestion. |
run.json | Run metadata, status, targets, usage/cost. |
Option B — Managed cloud (no local infra)
The same strix binary drives the managed platform. Every command starts with strix cloud. Full details — asset registration, source uploads, reports, PR reviews, schedules, webhooks, and billing — are in the managed-pentesting-with-strix skill. Minimal flow:
# 1. Sign in (device flow — the user confirms a code in the browser; this also
# creates the account and workspace when needed)
strix cloud login
# If you need specific scopes, request them with --scopes:
# strix cloud login --scopes scans:read scans:write assets:read assets:write \
# vulnerabilities:read billing:read billing:write
# 2. Register and verify the target domain (verification prints a DNS record for the user)
strix cloud domains add --domain staging.example.com --asset-type web_app
strix cloud domains verify <domain-id>
# 3. Launch and wait
strix cloud scans start --engagement-type live_test --domain-ids <domain-id> --wait
# 4. Read validated findings
strix cloud vulns list --severity critical
For a local repository, strix cloud scans start --source . uploads the working tree (needs uploads:write) and infers a code review. When credits run out, strix cloud billing topup starts an agent-payable Stripe challenge — the managed skill covers the payment flow. Output is JSON when stdout is not a terminal, so the commands compose in scripts.
The raw REST API works too (https://app.strix.ai/api/v1, org-scoped bearer token — see docs.app.strix.ai). If Docker or local prerequisites are not already satisfied, use this path instead of trying to install infra.
Reporting & next steps
Summarize findings by severity (critical/high/medium/low/info) and include the PoC evidence. To remediate and verify fixes (via either path), use the fix-security-vulnerabilities-with-strix skill. To wire scanning into CI/CD, use the ci-security-scanning-with-strix skill.
Safety
Only scan targets the user owns or is authorized to test. The Cloud platform enforces domain verification before external scans; for the OSS CLI, confirm authorization yourself if the target looks like third-party infrastructure.
Related skills
More from usestrix/strix and the wider catalog.

web-app-penetration-testing
Black-box penetration testing of web apps with autonomous agents that validate every finding with a working exploit.

api-security-testing
Autonomously exploit OWASP API Security Top 10 vulnerabilities—BOLA, broken authorization, SSRF, injection—with proof-of-concept requests.

application-security-testing
Autonomous security testing across code, APIs, and live apps—ranked by proven exploitability.

ci-security-scanning-with-strix
AI-powered security scanning for CI/CD pipelines — block vulnerable code before merge.

valyu-best-practices
Complete Valyu API toolkit for real-time search, content extraction, and AI-powered research across web, academic, and financial sources.

drama-creator
创作竖屏短剧剧本,包括宏观建构、剧本创作、精准优化、创意发想。适用于从零开始创作短剧、优化现有剧本、设计故事大纲和悬念钩子