Security Testing — Authorized Pentest
Only test systems you own or are explicitly authorized to test. AI Partner requires an in-scope allowlist and refuses any host outside it. Unauthorized scanning is illegal in most jurisdictions — you are responsible for having permission.
What it is
AI Partner can run an authorized web penetration test end to end. A team of cybersecurity specialists works inside the T3 desktop (Partner Mode), and the system turns their raw evidence into a defensible report:
- Recon → scan → web-app exploitation → analysis → report, driven by the
cybersec-*specialist agents. - Findings are proven — with real proof-of-concept requests, out-of-band callbacks
(interactsh) for blind bugs, or a cited scan line. Unproven claims are dropped or marked
not_tested. - The report is deterministic and code-rendered:
findings.json→findings.csv+ CVSS v3.1 scores +compliance.md(SOC 2 / ISO 27001 / PCI-DSS mapping) +report.pdf. The model never hand-writes scores or the findings table, so the prose can never contradict the numbers.
Prerequisites
- 1Enable Authorized Security Testing
Open Capabilities and turn on Authorized Security Testing (high-risk — you confirm you will only test systems you are authorized to test).
- 2Activate Partner Mode (T3)
The pre-installed toolkit lives in the T3 desktop. Say "wake up partner" in chat (or use the Partner toggle in the top bar). If the desktop image isn't built yet, AI Partner tells you the exact
docker buildcommand.
How to run it
/pentest example.com
The target you name becomes the in-scope allowlist. /pentest checks Partner Mode first and
guides you if it's off.
Modes
| Mode | Input | What it does |
|---|---|---|
| Black-box | a URL / host | Unauthenticated external test — recon, scan, web-app exploitation |
| Authenticated grey-box | URL + credentials | Logs in and tests behind auth (IDOR, priv-esc, business logic) |
| White-box (repo SAST) | a GitHub repo / local path | semgrep + gitleaks + dependency audit, with suggested fix patches |
| CI / PR-diff | a pull request | Diff-scoped SAST gate, headless — run from CI via security-vapt-ci |
What you get
All deliverables land in the goal's workspace folder (browsable under Generated Files, with in-app PDF/HTML preview):
findings.json— the structured source of truth (one entry per finding, with its proof).findings.csv— the canonical findings table, system-rendered with CVSS scores.compliance.md— CWE → SOC 2 / ISO 27001 / PCI-DSS control mapping (indicative).report.pdf+executive-summary.pdf— the narrative report.scans/andevidence/— the raw tool output and proof-of-concept artifacts.
Safety model
- In-scope allowlist (fail-closed). Every active tool validates the target against
scope_authorization.txt; anything outside the allowlist is refused — even if the agent forgets to check first. - Evidence-first, no fabrication. A
confirmedfinding must cite a real scan line, an out-of-band receipt, or a validated PoC; the anti-fabrication floor rejects the rest. - WAF honesty. If the target only returns a bot/WAF challenge, application-layer findings are
automatically downgraded to
not_testedand the report states no origin access was achieved. - Non-destructive by default. No data exfiltration, deletion, or brute-force without an explicit per-action opt-in in the scope file.
The specialist team
| Agent | Role |
|---|---|
cybersec-pm | Leads the engagement; delegates specialists and chains findings |
cybersec-recon | OSINT, DNS/subdomains, endpoint discovery |
cybersec-scanner | Ports, TLS, CVE/template scanning |
cybersec-webapp | XSS / SQLi / SSTI / IDOR / CSRF / JWT / SSRF / XXE against discovered endpoints |
cybersec-exploit | Builds + validates non-destructive proofs-of-concept |
cybersec-sast | White-box static analysis + secret scanning + fix patches |
cybersec-analyst | Consolidates all evidence into findings.json |
cybersec-reporter | Writes the narrative report from findings.json |