AI Bug Hunter — Autonomous Vulnerability Research Platform
Autonomous security testing tends to come in two extremes. On one end, automated scanners — fast, but they only flag potential weaknesses, with no proof, no human accountability, and no real control over what they touch. On the other end, a new wave of “autonomous AI hackers” that treat their own model output as permission — adaptive, but unsafe the moment the AI’s narrative becomes its own authority.
There was little in between. AI Bug Hunter is built to fill that gap — an authorization-first, Two-Brain autonomous vulnerability research platform where the AI proposes and interprets, but deterministic controls decide what may execute and what counts as proof.
What AI Bug Hunter Actually Is
AI Bug Hunter sits deliberately between a passive scanner and an ungoverned AI agent, combining adaptive research reasoning with a deterministic authorization-and-evidence core — so every action is authorized, and every finding is backed by verified evidence rather than model confidence.
- Authorization-First Execution — every executable action must traverse an authorization path the AI cannot modify or self-approve. The system fails closed when the authority to test is absent, stale, ambiguous, or inconsistent with the requested target.
- Two-Brain Adaptive Research — Brain A forms a hypothesis and proposes bounded experiments; Brain B interprets the deterministic outcome and recommends the next bounded step. Between them sits a non-AI evidence and policy core the model cannot rewrite.
- Deterministic Policy Engine — scope, exclusions, action class, rate limits, approval requirements, parallelism, and kill-switch state are all evaluated outside the model. Model preference is never treated as authorization.
- Single-Use Execution Grants — an approved action is transformed into narrowly bound, one-time execution authority, not general capability. Nothing is reusable or open-ended.
- Controlled Workers — execute only the exact authorized bounded operation; a grant can never be turned into arbitrary shell or network authority.
- Exact Evidence Lineage — every candidate finding is bound to the exact action, policy, approval, grant, job, worker, and test run. Model prose is not evidence; only verified worker evidence counts.
- Differential Semantics — a bounded outcome vocabulary (EXPECTED_DIFFERENCE, UNEXPECTED_DIFFERENCE, EQUIVALENT, INCONCLUSIVE, BLOCKED, ERROR). Missing or ambiguous evidence cannot be “upgraded” by model confidence.
- Database Isolation & Concurrency — PostgreSQL row-level security, atomic grant consumption, durable job leasing, and worker claim controls; production startup rejects unsafe runtime privileges such as superuser or BYPASSRLS.
- Human-in-the-Loop Approval — human approval is required wherever policy demands it, and every finding is reviewed by a human before it means anything.
- Release Assurance — the platform deliberately separates “the application runs and passes tests” from “production is genuinely ready,” keeping production fail-closed until external infrastructure is provisioned and verified.
Why AI Bug Hunter Is Different
It’s not a scanner. Scanners hand you a list of potential weaknesses with no proof and no governance. AI Bug Hunter authorizes each action, tests bounded hypotheses, and binds every finding to verified evidence linked to the exact action that produced it.
It’s not an ungoverned AI agent. The AI never becomes its own authority. It can propose and interpret, but it cannot approve itself, alter scope, mint or consume execution grants, release the kill switch, or obtain unrestricted execution. Those are separate control planes by design.
It’s not an AI black box. No match decision comes from the model. Scope, policy, approval, execution, and what counts as proof are decided by deterministic, non-AI controls — this is architectural, not a policy setting. The AI’s job is to hypothesize, adapt, and interpret; never to grant itself permission.
And it’s honest about what it isn’t. It doesn’t claim universal market exclusivity, and it doesn’t claim to out-discover other tools — recorded validation verifies that the controls work, not discovery superiority. Its production release deliberately remains NO-GO until mandatory external infrastructure is genuinely provisioned and verified. Every boundary is stated up front.
The Two-Brain Research Core
A controlled research pattern — not two unrestricted agents — that lets the platform stay adaptive without letting the AI decide what’s true or what’s allowed:
- Brain A (Research Formation) — forms a hypothesis, selects a bounded registered test, and proposes parameters. It has no direct execution.
- Deterministic Core — evaluates policy per side, enforces approval/grant/worker/evidence lineage, applies allowlisted normalization, and performs exact evidence matching. The AI cannot rewrite this comparison.
- Brain B (Interpretation) — interprets the deterministic outcome, refines or stops the hypothesis, and recommends the next bounded test.
- Minimum-necessary-proof stopping reduces unnecessary destructive experimentation, and pre-execution revalidation re-checks scope, policy, approval, kill-switch, and rate limits immediately before any action runs — because authority can change between proposal and execution.
Who AI Bug Hunter Is For
Security teams, VAPT providers, and researchers who want adaptive, AI-accelerated testing without handing an AI unchecked authority — and organizations that need governed, auditable, authorization-first security research with a complete evidence trail behind every finding. Where a scanner gives noise and an ungoverned agent gives risk, AI Bug Hunter gives governed autonomy with verifiable boundaries.
Research & Documentation
The full technical architecture, methodology, threat model, and an honest breakdown of production boundaries are published as a preprint:
