What actually changed after we put an LLM in the pipeline
Hunting is a coverage problem. There is a hard ceiling on how many endpoints one person can read carefully, and that ceiling is lower than most people assume. We put an LLM in the pipeline not to find new classes of bug, but to find known classes across more surface.
That distinction carries the whole post. The model has not found a vulnerability a human could not have found. It has surfaced candidates in places nobody had time to look. The workflow has earned over $10,000 in real bounties, and the cause of that number is throughput, not insight.
Shape of the pipeline
Three stages, each one’s output feeding the next.
[1] recon → endpoints, parameters, auth boundaries
[2] code audit → suspect locations + evidence + how to check
[3] variant search → where the confirmed pattern survives elsewhere
1. Recon — structure is the entire trick
Asking a model to “find vulnerabilities” produces output you cannot act on. So it is never asked to. It is given a classification job.
Input is proxy logs, OpenAPI specs, and routes pulled out of client bundles. Output is a fixed schema: endpoint, method, required auth level, which identifiers are attacker-controlled, whether this endpoint calls another service.
Forcing the output into a schema is the point. Free prose cannot be verified, and what cannot be verified cannot be scored. Schema output can be checked by a program: if a row claims an endpoint needs no authentication, send the request without a token and find out. Run that check automatically and close to half of the model’s claims die before a human sees them.
2. Code audit — demand the evidence
This pays off most on white-box targets, where source is available. The source does not go in wholesale. Starting from the handler the recon stage flagged, the pipeline walks the call graph and feeds it in slices.
The required output has three columns, always:
| Column | Content |
|---|---|
| Location | File, function, line |
| Evidence | Why this might be wrong, quoting the code |
| How to check | A concrete procedure that settles whether it is real |
Anything missing the third column is discarded. A finding whose author cannot describe how to test it is usually a model reciting generalities rather than reading the code in front of it. Requiring that one column cuts the review queue dramatically.
3. Variant search — where the money actually is
The most productive stretch begins after the first flaw is confirmed. The same mistake recurs within a codebase, because the same author is wrong in the same way twice.
Two patterns we have scraped repeatedly are named on the home page:
- The auth-proxy pattern. A front tier handles authentication and the back tier trusts the front tier. If any route reaches the back tier directly, the authentication in front of it is decoration. Once you have seen the shape, it is recognisable in unrelated services.
operation_idreplay. Where a work-item identifier is guessable or reusable, you can ride into someone else’s operation context.
In this stage the model’s job is “find code that is structurally the same as this pattern”. grep cannot do it — the names and the phrasing differ and only the shape matches. This is the task LLMs are honestly good at.
For binary targets, headless decompiler output goes through the same pipeline. Decompiled code is voluminous and unpleasant to read, so the step that used to be “pick somewhere promising by intuition” becomes an exhaustive sweep.
Where false positives die
This is the part worth taking away. The false positive rate is not low. It never became low. What changed is that most false positives now die before a human reads them.
Three sieves, in order:
- Schema validation. Output that breaks the required shape is dropped. In practice, a model that cannot hold the format is usually not holding the content either.
- Mechanical checking. Any verifiable claim gets confirmed in code. “This parameter is not validated server-side” is settled by sending the request. This sieve catches the most.
- Reproduction. Whatever survives, a human tries to reproduce. Only what survives that becomes a report candidate.
Skip the automation in step 2 and the whole pipeline is a net loss. Hand unverified model output to a person and that person stops reading code and starts auditing the model’s prose instead — which is slower than the thing it replaced. Almost every failed LLM adoption we have seen fails exactly here.
What the model gets wrong most reliably
- Reachability. It finds logic flaws inside a function well, and is frequently wrong about whether anything external can reach that function. It guesses, without the full call graph.
- Implicit defences. It reports things the framework or a middleware already blocks. Stating the framework’s default behaviour in the context reduces this; it does not remove it.
- Authentication versus authorization. It treats “only authenticated users can reach this” as safe. A missing authorization check looks exactly like that.
- Confident invention. It will produce function names and config keys that do not exist, fluently. This is why evidence must be a code quotation: if the quoted line is not in the file, the whole row is discarded.
Steps that stay human
The manual steps are written down rather than left to judgement in the moment.
Scope. People decide what may be tested. We do not hand a program’s scope document to a model and let it pick targets. Testing something out of scope cannot be undone.
Sending live requests. The pipeline produces candidates and procedures; a person executes them. State-changing requests — writes, deletes — are manual without exception. Running a model-authored procedure unreviewed against production data is not a discovery, it is an incident.
Impact. What the flaw actually enables, and how severe it is, is written by a person. Models overestimate severity consistently and in one direction.
The report. Reports are written by whoever confirmed the reproduction actually reproduces, and redaction is a human pass. That part has its own document.
When the AI product is the target
Run it the other way and LLM and agent systems are the target: prompt injection, tool misuse, privilege boundary bypass. The method does not change much.
In an agent system the interesting surface is not the model, it is the boundary between the model and its tools. Who validates the arguments when the model calls a tool, and when a tool’s output re-enters the model’s context, is it treated as data or as instructions? Confuse the second one and external input becomes command input. It reduces to an authorization boundary problem — the same class of mistake we read about in web applications, wearing different clothes.
Summary
- An LLM does not find new classes of bug. It finds known classes across more surface.
- Force output into a schema and check verifiable claims mechanically. Without that step the pipeline costs more than it returns.
- The false positive rate is unchanged. What changed is how many of them reach a person.
- Variant search is the highest-yield stage. Confirming the first flaw is where the work starts.
- Scope decisions, state-changing requests, impact assessment and report writing are not automated.