Skip to main content

Security Testing — Authorized Pentest

Only test systems you own or are explicitly authorized to test. AI Partner requires an in-scope allowlist and refuses any host outside it. Unauthorized scanning is illegal in most jurisdictions — you are responsible for having permission.

What it is​

AI Partner can run an authorized web penetration test end to end. A team of cybersecurity specialists works inside the T3 desktop (Partner Mode), and the system turns their raw evidence into a defensible report:

  • Recon → scan → web-app exploitation → analysis → report, driven by the cybersec-* specialist agents.
  • Findings are proven — with real proof-of-concept requests, out-of-band callbacks (interactsh) for blind bugs, or a cited scan line. Unproven claims are dropped or marked not_tested.
  • The report is deterministic and code-rendered: findings.json → findings.csv + CVSS v3.1 scores + compliance.md (SOC 2 / ISO 27001 / PCI-DSS mapping) + report.pdf. The model never hand-writes scores or the findings table, so the prose can never contradict the numbers.

Prerequisites​

  1. 1
    Enable Authorized Security Testing

    Open Capabilities and turn on Authorized Security Testing (high-risk — you confirm you will only test systems you are authorized to test).

  2. 2
    Activate Partner Mode (T3)

    The pre-installed toolkit lives in the T3 desktop. Say "wake up partner" in chat (or use the Partner toggle in the top bar). If the desktop image isn't built yet, AI Partner tells you the exact docker build command.

How to run it​

/pentest example.com

The target you name becomes the in-scope allowlist. /pentest checks Partner Mode first and guides you if it's off.

Modes​

ModeInputWhat it does
Black-boxa URL / hostUnauthenticated external test — recon, scan, web-app exploitation
Authenticated grey-boxURL + credentialsLogs in and tests behind auth (IDOR, priv-esc, business logic)
White-box (repo SAST)a GitHub repo / local pathsemgrep + gitleaks + dependency audit, with suggested fix patches
CI / PR-diffa pull requestDiff-scoped SAST gate, headless — run from CI via security-vapt-ci

What you get​

All deliverables land in the goal's workspace folder (browsable under Generated Files, with in-app PDF/HTML preview):

  • findings.json — the structured source of truth (one entry per finding, with its proof).
  • findings.csv — the canonical findings table, system-rendered with CVSS scores.
  • compliance.md — CWE → SOC 2 / ISO 27001 / PCI-DSS control mapping (indicative).
  • report.pdf + executive-summary.pdf — the narrative report.
  • scans/ and evidence/ — the raw tool output and proof-of-concept artifacts.

Safety model​

  • In-scope allowlist (fail-closed). Every active tool validates the target against scope_authorization.txt; anything outside the allowlist is refused — even if the agent forgets to check first.
  • Evidence-first, no fabrication. A confirmed finding must cite a real scan line, an out-of-band receipt, or a validated PoC; the anti-fabrication floor rejects the rest.
  • WAF honesty. If the target only returns a bot/WAF challenge, application-layer findings are automatically downgraded to not_tested and the report states no origin access was achieved.
  • Non-destructive by default. No data exfiltration, deletion, or brute-force without an explicit per-action opt-in in the scope file.

The specialist team​

AgentRole
cybersec-pmLeads the engagement; delegates specialists and chains findings
cybersec-reconOSINT, DNS/subdomains, endpoint discovery
cybersec-scannerPorts, TLS, CVE/template scanning
cybersec-webappXSS / SQLi / SSTI / IDOR / CSRF / JWT / SSRF / XXE against discovered endpoints
cybersec-exploitBuilds + validates non-destructive proofs-of-concept
cybersec-sastWhite-box static analysis + secret scanning + fix patches
cybersec-analystConsolidates all evidence into findings.json
cybersec-reporterWrites the narrative report from findings.json