Most AI pentest tools generate findings.
Nemesis Red proves them.
One reasoning loop drives the full kill chain: recon, vulnerability assessment, live exploitation, and novel zero-day discovery. It runs the real Kali arsenal over SSH to a box you control, fuses every tool's output into the next decision, and reports nothing it cannot prove. Your targets, traffic, and findings stay in your environment.
A finding you cannot reproduce is not a finding
Autonomous scanning got cheap. Trustworthy autonomous scanning did not. A language model can describe a vulnerability that does not exist, cite a CVE that does not apply, and rank it critical. Run one of those tools against a real environment and the output is a queue of plausible claims a human still has to verify by hand. That is the work the tool was supposed to do.
Nemesis Red makes verification the gate, not an afterthought. Every finding that reaches a report has cleared a deterministic check, not a model's opinion. When nothing proves out, the report says so.
How the loop runs
The reasoning loop plans, drives the arsenal, and fuses every result into the next decision. Before anything reaches a report it passes the verifier gate. Proven findings ship. Unproven ones are dropped, not printed. No fixed script, and no unverified hypotheses in the deliverable.
What sets it apart
- Proof, not probability.Web and network findings are scoreboard-verified against the live target. Memory-safety findings climb a proof ladder from crash, to attacker-controlled primitive, to a working proof of exploitation. Nothing ships on a model's say-so.
- A loop that reasons, backed by research.The core is ReasonChain, a closed-loop architecture that plans, acts, reads the result, and replans on accumulated evidence. Cross-tool fusion is causally validated in a controlled ablation (Wilcoxon p<0.001). The paper and the code that regenerates every number are public.
- It drives the real arsenal. The standard offensive toolchain runs over SSH on a Kali box you own: nmap, nuclei, sqlmap, ffuf, and the rest, alongside purpose-built engines for API surface discovery, broken access control, and an authenticated browser agent. The value is the reasoning over the tools, not a count of them.
- Zero-day discovery, not CVE lookup. Point the discovery engine at the open-source components a target runs. It performs coverage-guided fuzzing and patch-seeded variant hunting, proves exploitability on the same proof ladder, and assembles a coordinated-disclosure packet.
- Security for the AI you ship. Products now ship language-model features, and the attack surface moved with them. Nemesis Red tests LLM applications against the OWASP LLM Top 10, including prompt injection, data exfiltration, and tool abuse, across API, chat, and browser.
Four ways to work, one platform
- Copilot console.A live multi-session terminal with an AI that reads every command's output, tells you why it matters, and hands you the next commands to run with one click.
- Autonomous vulnerability assessment. Give it a URL, host, or CIDR and a depth. It builds a live reasoning graph, fuses tool output into CVE-backed findings, and exports a signed report and remediation plan. Run it once, or schedule it on a cron.
- Autonomous pentest. Take assessment findings into live exploitation. It selects and fires exploits and captures real proof: credentials, hashes, database dumps, and shells.
- Zero-day scanner. Point the discovery engine at an asset or an open-source component. It fuzzes, hunts variants of known fixes, proves each candidate on the proof ladder, and produces a reproducing trigger, a root cause, and a suggested fix. For installed Windows desktop software (image, PDF, and office parsers), a separate native package, Nemesis Red Zero-Day for Windows, runs the same proof engine against closed
.exefiles, since a Linux container cannot open a Windows app.
Using Zero-Day for Windows: download the portable package (Windows 10/11 + Server, x64, self-contained), extract it, and point it at an app you own in a disposable VM:
nrzd.cmd --exe "C:\Program Files\Vendor\App.exe" --seeds seeds\sample.tga --suffix .tga --argv-template "{file}" --coverage offProof it works, not a demo
A novel bug, found and disclosed. The discovery engine found a previously-unreported heap out-of-bounds write in CWPack, a widely-used MessagePack library for C. A truncated MessagePack stream is treated as a successful read, the parser advances past the end of its buffer, and the next refill calls memmove with a negative size. The engine minimized a 74-byte trigger, root-caused it to the stream-refill handler, produced a short proof-of-vulnerability and a suggested patch, and reported it upstream under coordinated disclosure. Found by fuzzing the public source, with no third-party systems touched. This is novel-bug discovery with a reproducing trigger and a fix, not a match against a known CVE list.
Depth on a live target. A single autonomous run against an OWASP-class environment surfaced more than 300 CVE-backed findings, including multiple CVSS 9.8 issues (for example CVE-2024-38476, an Apache mod_rewrite SSRF), each reproducible from a clean checkout. That is roughly an order of magnitude more validated findings than a single-pass scan, measured in a controlled ablation, while precision on the labeled benchmark holds at 82 percent. The verifier gate is what keeps the noise out.
The architecture, cooperating modules
- reasonchain. The closed-loop planner. State is an evidence graph. Actions are tool invocations. It plans, runs, reads the result, and replans against everything learned so far, backed by the LLM you point at.
- engines. Kali tool adapters plus purpose-built engines (API surface discovery, broken access control, browser agent, LLM-application testing), each with a structured-input and structured-output contract. The loop calls by capability, not by name.
- verifier. Scoreboard proof-of-impact for web and network findings, and a proof ladder for memory-safety bugs. A returned shell, a dumped credential, an attacker-controlled write. No unverified hypotheses reach the report.
- discovery. The zero-day engine: coverage-guided fuzzing, patch-seeded variant hunting, and a disclosure-packet assembler.
- consent. Per-engagement scope enforcement. Out-of-scope actions are refused at the adapter, before the tool runs.
- audit and report. An append-only, hash-chained log of every reasoning step and tool call, and a deliverable generator with proofs, attack-chain narratives, and prioritization by exploit chain.
Bring your own key, or run fully local. Nothing leaves your network.
- Frontier LLM, your key. Anthropic, OpenAI, Gemini. Your enterprise key, your terms. No Nemesis Labs cloud in the path.
- Self-hosted via Ollama. Qwen, Llama, or your own fine-tune on a single workstation GPU.
- Air-gapped operation. Run offline on consumer hardware under Enterprise.
All evidence, tool output, intermediate reasoning, and final reports live in the customer-controlled Docker volume. For teams selling into defense, this is the architectural prerequisite that cloud-resident competitors cannot match without years of compliance work.
The economics
A senior pentester engagement runs $40K to $150K and happens three to five times a year. Continuous-pentest SaaS (Horizon3, Pentera) is $200K to $500K per year, cloud-hosted, and locked behind enterprise sales. Nemesis Red is autonomous and self-serve at the same time: a working pentester gets full autonomous penetration testing for $99 a month, self-hosted, data never leaving their own box. Business teams run continuous assessment for a few thousand a year, not a few hundred thousand. One avoided manual engagement pays for years.
Pricing
- Free1 target, VA + 3-run PT taste, local copilot$0
- Pro Pentester10 targets, full autonomous PT, scheduled scans$99 / mo · $990 / yr
- Business3 seats, 50 assets, continuous scanning, compliance, API$599 / mo + $99/seat + $10/asset
- Enterpriseunlimited scope, on-prem/air-gap, SSO, MSSP, SLAcustom, from ~$25K / yr
Every tier ships the Docker engine and a pre-baked Kali UTM VM, runs on your own hardware against your own Kali, and uses your own LLM keys or a local model. Invoicing available in EUR, GBP, AED, SGD, NGN, ZAR, KES on request.
Built for security researchers
Nemesis Red is built the way researchers work. Findings come with a reproducing trigger, a root cause, and a suggested fix. The zero-day engine assembles the coordinated-disclosure packet for you. The core architecture is public and MIT-licensed, and every result in the research paper regenerates from a clean checkout. You self-host it, you bring your own key, and your work stays yours. When the engine finds nothing, it tells you, which is the only honest output for a run that found nothing.
Honest framing
Nemesis Red is not a substitute for a senior human pentester on a zero-knowledge engagement against a mature blue team. It is a force-multiplier on recon, assessment, and initial access that frees human operators for chain-validation, OPSEC, and reporting. Autonomous tool use, not autonomous tradecraft.
It also will not run against a target without attested written consent. The engine gates every engagement on a consent attestation before it will touch a target, and the hash-chained audit log retains that attestation for the engagement's lifetime.
See it in action
The operator console, the autonomous vulnerability assessor, live exploitation with captured proof, and the AI browser agent. Click any shot to zoom.
Sources
- CWPack, MessagePack library for C. Heap out-of-bounds write in the stream-unpack refill handler, found by the Nemesis Red discovery engine and disclosed under coordinated disclosure.
- NVD entry, CVE-2024-38476. Apache HTTP Server backend selection via crafted URLs (example of a CVSS 9.8 surfaced on a live run).
- OWASP Top 10 for LLM Applications. The reference the LLM-application test suite is built against.
- Kali Linux tools catalog. The offensive toolchain Nemesis Red drives over SSH.