PluginBench
Skill
Review
Audit score 70

owasp-top-10-testing

usestrix/strix

Test applications against OWASP Top 10:2025 and API Security Top 10 with autonomous exploitation agents.

What is owasp-top-10-testing?

Strix agents autonomously attempt real exploits for each OWASP Top 10:2025 category and OWASP API Security Top 10 (2023) category, reporting only what they can actually prove with proof-of-concept. Use this when assessing OWASP compliance, conducting security reviews mapped to OWASP categories, or validating fixes after remediation.

  • Systematically tests all ten OWASP Top 10:2025 categories: broken access control (including SSRF), security misconfiguration, software supply chain failures, cryptographic failures, injection, insecure design, authentication failures, integrity failures, logging/alerting failures, and mishandling of exceptional conditions.
  • Covers OWASP API Security Top 10 (2023) categories including BOLA, broken object property-level authorization, and broken function-level authorization.
  • Reports only exploits that agents could actually prove, mapped back to the specific 2025 category with proof-of-concept evidence.
  • Accepts source code, running instances, and multi-privilege credentials for maximum coverage.
  • Groups findings by category and explicitly states what could not be assessed (e.g., logging/alerting requires pipeline review).

How to install owasp-top-10-testing

npx skills add https://github.com/usestrix/strix --skill owasp-top-10-testing
Prerequisites
  • Strix binary installed (via penetration-testing-with-strix skill or managed-cloud login).
  • A running instance of the application to test, or source code repository URL.
  • Credentials at two privilege levels (e.g., regular user and admin) for A01 (access control) testing.
  • For managed-cloud runs: `strix cloud login` and valid cloud credentials.
Claude Code
Cursor
Windsurf
Cline

How to use owasp-top-10-testing

  1. 1.Prepare credentials: obtain at least two user accounts (different privilege levels) and note their passwords.
  2. 2.Run Strix with OWASP Top 10:2025 instruction, specifying source and/or running instance: `strix -n -t <source_url> -t <running_instance_url> --scan-mode deep --max-budget 30 --instruction "OWASP Top 10:2025 assessment..."`.
  3. 3.Review the run output in `strix_runs/<run>/vulnerabilities/` grouped by category.
  4. 4.For each category, verify the proof-of-concept yourself and confirm the mapping to the 2025 category ID.
  5. 5.Note categories that could not be fully assessed (A09 logging/alerting always requires manual review; A03/A04/A06/A08/A10 may be partial).
  6. 6.Remediate findings using fix-security-vulnerabilities-with-strix and re-run to prove exploits are closed.

Use cases

Good for
  • Conduct a baseline OWASP Top 10:2025 assessment on a web application before production deployment.
  • Validate that security fixes for a reported vulnerability actually close the exploit path.
  • Gate pull requests in CI/CD with automated OWASP category coverage scanning.
  • Generate an auditor-facing technical report mapped to OWASP 2025 categories for compliance review.
  • Test API endpoints against OWASP API Security Top 10 (2023) for authorization and data-exposure flaws.
Who it's for
  • Security engineers and penetration testers conducting application assessments.
  • DevSecOps teams integrating security testing into CI/CD pipelines.
  • Compliance and audit teams validating OWASP Top 10 coverage.
  • Development teams validating security fixes before re-release.

owasp-top-10-testing FAQ

What is the difference between OWASP Top 10:2025 and 2021?

2025 is the current edition. Key changes: SSRF is now folded into A01 Broken Access Control, A03 Software Supply Chain Failures expands the old "Vulnerable and Outdated Components", A02 Security Misconfiguration moved from position 5 to 2, and A10 Mishandling of Exceptional Conditions is new. Always ask the user which edition they need; a report with the wrong label is misleading.

Can Strix test all ten OWASP categories equally well?

No. A01, A02, A05, and A07 have strong coverage via exploitation. A03, A04, A06, A08, and A10 are partial (need source or infra review). A09 (logging/alerting) is not testable from outside — it requires reviewing the logging pipeline. Always state what was attempted and what could not be assessed rather than claiming a clean sweep.

What credentials do I need to provide?

For maximum coverage, especially A01 (access control), provide at least two user accounts at different privilege levels (e.g., user in org 1, user in org 2, and an admin). Without a second account, A01 results will be incomplete — disclose this in the report.

What does a budget-capped run mean?

If the scan hits `--max-budget` before finishing, it is not a completed assessment. Check `run.json` status and cost; a partial run may have missed categories. Increase the budget and re-run to cover all ten categories systematically.

How do I use this in CI/CD?

Use the ci-security-scanning-with-strix skill to gate pull requests. Run Strix on each commit, fail the build if new exploitable vulnerabilities are found, and track remediation with fix-security-vulnerabilities-with-strix.

Full instructions (SKILL.md)

Source of truth, from usestrix/strix.


name: owasp-top-10-testing description: Test an application against the OWASP Top 10 with Strix — autonomous AI agents that attempt real exploits for each category of the current OWASP Top 10:2025 (broken access control including SSRF, security misconfiguration, software supply chain failures, cryptographic failures, injection, insecure design, authentication failures, integrity failures, logging and alerting failures, mishandling of exceptional conditions) and report only what they could actually prove, mapped back to the category with a proof-of-concept. Also covers the OWASP API Security Top 10 (2023). Use when the user asks for an OWASP Top 10 assessment, OWASP compliance testing, or a security review mapped to OWASP categories. license: Apache-2.0 metadata: author: usestrix homepage: https://docs.strix.ai

Test against the OWASP Top 10

The OWASP Top 10 is a taxonomy of risk categories, not a test suite — "OWASP Top 10 testing" means exercising each category against the real application and reporting what's actually exploitable. Strix's agents do the exploitation; this skill covers running it category-by-category and reporting coverage honestly.

Use the current edition: OWASP Top 10:2025 (8th installment, superseding 2021). Ask the user before targeting an older edition — some compliance checklists still reference 2021, and a report labelled with the wrong edition is misleading. Key differences from 2021: SSRF is folded into A01, A03 Software Supply Chain Failures expands the old "Vulnerable and Outdated Components", and A10 Mishandling of Exceptional Conditions is new; A02 Security Misconfiguration moved 5→2.

Install, LLM setup, and the managed-cloud alternative: penetration-testing-with-strix. For a run with no Docker and no LLM key, the same binary drives the managed platform: strix cloud login, then strix cloud scans start ... (details in managed-pentesting-with-strix).

What is and is not testable by an agent

Be straight with the user about this — claiming a clean sweep of all ten is misleading.

Category (2025)Coverage
A01 Broken Access Control (incl. SSRF)Strong — cross-user/tenant access, privilege escalation, IDOR, and SSRF (including blind, via out-of-band callbacks) are all exploit-validated. Needs two accounts plus a privileged one to prove the authorization half.
A02 Security MisconfigurationStrong — debug endpoints, verbose errors, permissive CORS, missing hardening, default credentials, exposed admin surfaces.
A03 Software Supply Chain FailuresPartial — version fingerprinting, and vulnerable/outdated dependency review when source is supplied. Build-system and distribution-infrastructure compromise (the broader half of this category) is out of scope for a runtime scan — pair with SCA plus build-provenance controls.
A04 Cryptographic FailuresPartial — transport config, unencrypted data in transit, secrets and tokens leaked in responses. At-rest crypto and key management need source or infra review.
A05 InjectionStrong — SQL/NoSQL/command/template injection and XSS, exploit-validated.
A06 Insecure DesignPartial — business-logic abuse (price/quantity tampering, workflow skipping, race conditions) is found where reachable; design intent still needs human review and threat modelling.
A07 Authentication FailuresStrong — auth bypass, weak session/token handling, password-reset and MFA flaws.
A08 Software or Data Integrity FailuresPartial — insecure deserialization and unsigned-update paths where reachable; CI/CD trust boundaries are not runtime-testable.
A09 Security Logging & Alerting FailuresNot testable from outside — requires reviewing the logging and alerting pipeline. State this rather than reporting it as passed.
A10 Mishandling of Exceptional ConditionsPartial — agents actively probe error handling and fail-open behavior (malformed input, forced errors, race and timeout conditions) and report what leaks or bypasses a control; exhaustive coverage of internal error paths needs source review.

For APIs, run the same exercise against the OWASP API Security Top 10 (2023) — API1 BOLA, API3 Broken Object Property Level Authorization (2019's excessive data exposure + mass assignment merged), API5 broken function-level authorization — using the api-security-testing skill.

Run it

Maximum category coverage comes from giving the agents both the source and a running instance, plus credentials at two privilege levels:

strix -n \
  -t https://github.com/org/app \
  -t https://staging.example.com \
  --scan-mode deep --max-budget 30 \
  --instruction "OWASP Top 10:2025 assessment. Cover every category systematically and map each finding to its 2025 category id.
Accounts: userA@example.com/<pw> (org 1), userB@example.com/<pw> (org 2), admin@example.com/<pw>.
Prioritise A01 (cross-org access, privilege escalation, SSRF), A02, A05, A07, A10.
Out of scope: /billing/*, outbound email."
  • --scan-mode deep matters here: systematically walking ten categories is not a quick scan.
  • Without a second account, A01 results are structurally incomplete — say so in the report rather than leaving it implied.
  • Need an auditor-facing PDF? Run it through the managed platform and pull the technical report (managed-pentesting-with-strix).

Report honestly

From strix_runs/<run>/, group vulnerabilities/*.md by category and state, per category: what was attempted, what was proven, and what could not be assessed (A09 always; A03/A04/A06/A08/A10 partially). Label the report with the edition used. Verify each PoC yourself before it goes in front of the user.

A 0 exit code means nothing exploitable was proven in what was analyzed — check run.json status and cost against --max-budget; a budget-capped run is not a completed assessment.

Then fix and re-test

Remediate with fix-security-vulnerabilities-with-strix and re-run to prove each exploit is closed. For ongoing coverage as the app changes, gate pull requests using ci-security-scanning-with-strix.