Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Darkmoon, autonomous AI penetration testing

Darkmoon Research, autonomous AI penetration testing evidence corpus

Star Dark-Moon

License GPL-3.0 Evidence corpus Public labs only Website

Real reports from autonomous AI penetration testing runs, the raw AI security testing evidence behind Darkmoon's multi-agent engine. 50 specialist AI agents. Real exploits chained across web, cloud, Active Directory and Kubernetes. Proof for every finding. Self-hosted, and the Privacy Gateway tokenizes your real values so the model works on placeholders while your real IPs, hosts and credentials stay on your perimeter.

If Darkmoon is useful, a star on the main repo helps others find it.

⭐ Star DarkMoon · ▶️ Watch the 60s demo · Darkmoon autonomous AI penetration testing · AI pentest tool comparison · remediation benchmark (Pro) · visual walkthrough · benchmark results


Darkmoon open source CLI, terminal environment model summary from black-box recon of the target

The open source Darkmoon CLI, its terminal environment model, the recon evidence each report in this corpus is built from, captured black-box against a public lab.


Pro web dashboard, Darkmoon infrastructure map with per-finding evidence

Pro: the paid Darkmoon Pro web dashboard, its infrastructure map, targets, connections and per-finding evidence across planes.

Web dashboard and remediation are Darkmoon Pro (paid) features; the open source edition is the CLI shown above.


Darkmoon Research — Autonomous Multi-Agent Penetration Testing: Lab Evidence

This repository is the evidence corpus behind Darkmoon's autonomous penetration-testing agents. Every report here was produced by the agents themselves — an orchestrator that fingerprints a target and dispatches specialist sub-agents — running end-to-end against public training labs, not hand-written after the fact. Each finding carries the exact command and its raw response.

No client data appears in this repository. All targets are public, deliberately-vulnerable training environments (Pwned Labs, OWASP IoTGoat) or locally-hosted lab services.


Why this corpus exists

Autonomous pentest tools are easy to demo and hard to trust. A benchmark on one deliberately-broken web app tells you little about whether the same engine handles a cloud identity chain, an SSRF pivot into a metadata service, or an IoT firmware image. This corpus answers that by publishing the actual agent output across six distinct planes — cloud, identity, CI/CD, IaC, data/secrets, and embedded firmware — so the methodology can be inspected and the results reproduced.

Method (short)

Each engagement follows the same loop: fingerprint → dispatch specialist → exploit with proof → cascade recovered material to the next plane → server-side report. Cloud, infrastructure and firmware agents are artifact-gated: they dispatch only on a concrete positive artifact (a token, an exposed API/port, a firmware image), never on inference — the discipline that prevents false-positive plane attacks. The full method is in docs/methodology.md; the observed failure modes and their fixes are in docs/agent-hardening.md.


Evidence index

☁️ Cloud & identity

Lab Plane / technique Report
Pwned Labs — S3 online AWS keys → S3 enumeration & object exfiltration report
Huge Logistics S3 AWS anonymous/object-versioning exfiltration report
Azure Blob → Entra Anonymous blob → versioned secret → ROPC → tenant takeover path report
Intro to Azure Recon (BloodHound) Entra/Azure recon, role graph report
Unlock Access with Azure Key Vault Key Vault secrets → password-reuse pivot → Storage Table PII report · confirmation run
Reveal Hidden Files in Google Storage GCS object-name fuzzing → encrypted-archive crack → PII report
SSRF with Gopher for GCP Initial Access SSRF → gopher metadata smuggling → service-account token → bucket report

🔌 IoT / firmware — OWASP IoTGoat

Mode Technique Report
IMAGE binwalk extraction, shadow crack, backdoor & CVE mapping (20 findings) report
DEVICE live root via backdoor, Mirai-default SSH, LuCI command-injection (9 findings) report

🏗️ Infrastructure, data & secrets (local labs)

Lab Plane Report
Terraform + AWS(LocalStack) + Ansible IaC state secrets → IAM privesc cascade report
HashiCorp Vault + Registry + Docker secrets → registry → container escape report
PostgreSQL + MySQL data-plane RCE (COPY PROGRAM), hash dump report
Redis / messaging & cache broker/cache exposure report
GitLab CI/CD supply-chain & secrets report
Jenkins CI/CD runner & credential exposure report

Reading a report

Every report is server-generated from the pushed findings and follows the same structure: management summary, findings table, then a per-finding section with description, exploitation commands, raw evidence, CVSS 3.1 and remediation. The status of each finding is qualified adversarially — EXPLOITED (impact executed), CONFIRMED (impact demonstrated), UNCONFIRMED (lead, not yet proven) — so the counts are honest.

License & ethics

Published for research and education. All targets are public training labs; do not point these agents at systems you are not authorized to test.

About

Evidence corpus for autonomous AI penetration testing. AI security testing across cloud, identity, CI/CD, IaC, databases and IoT firmware, validated on public labs with real exploit proof.

Topics

Resources

Code of conduct

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages