OWASP 2026 · MAESTRO · AIVSS

Threat modeling for the AI era that never forgets.

Point it at a repo. Maroon Elephant detects your AI, agentic, and MCP components, maps every risk to the OWASP 2026 standards, and hands you an enterprise governance verdict — deterministic-first, zero-dependency, local-first.

$ pip install maroon-elephantcopy Get started →
LLM Top 10Agentic ASI01–10GenAI Data Security MCP securityMITRE ATLASSARIF 2.1.0 CycloneDX AI-BOM0 runtime depsApache-2.0
See it run

A scan, start to verdict.

Deterministic findings with file:line evidence — then a governance verdict you can act on.

maroon — zsh
0
OWASP-2026 risks
0
frameworks mapped
0
runtime deps
0
tests passing
The problem

AI collapsed the data plane and the control plane.

LLMs read instructions and data from one flat token stream — the root cause behind prompt injection, data exfiltration, excessive agency, and memory poisoning. The market is full of runtime products, but there's no strong open-source, repo-driven, design-time threat modeler for AI systems. That's the gap.

🔍 Detect

Fingerprint LangChain, LangGraph, CrewAI, MCP servers, vector DBs, model SDKs — and build an AI inventory / AI-BOM.

🧠 Model

Infer the architecture, place components on MAESTRO's 7 layers, and produce a design-time threat model as code.

⚖️ Govern

Score with AIVSS and place the repo on the AT×L governance matrix — from "Well-governed" to "DO NOT DEPLOY".

Explain it like I'm five

What does it actually do?

🐘 Imagine a very careful elephant that reads your code and asks: “If this app uses AI, how could someone trick it or steal from it?”

It doesn't guess. It looks for exact, known-risky patterns — a password typed into the code, an AI that can run shell commands, a tool that auto-approves dangerous actions — and points at the exact line.

Then it writes you a report card: what AI you're running, what could go wrong (in plain standards everyone trusts), and one clear next step to be safer.

Your code never leaves your machine. The AI parts are optional — the careful elephant works entirely offline.

How it works

Deterministic first. The LLM is a garnish, never the meal.

Every finding comes from the deterministic engine with file:line evidence and a stable fingerprint. The optional LLM layer can only describe a finding that already exists — it can never invent one. That's what makes the output auditable.

01

Ingest

Local path, git URL, or uploaded zip — safely (no path traversal).

02

Detect

Component inventory → AI-BOM + the system's adoption-tier signals.

03

Analyze

Rules + AST taint + Lethal-Trifecta + Grey Panda, merged & deduped.

04

Score

AIVSS severity + the AT×L governance verdict.

05

Report

SARIF · CycloneDX AI-BOM · threat-model-as-code · governance.

CRITICAL GAP
Adoption tier AT6 — Externally-Extended (MCP, third-party tools)
Maturity L0 — Ad Hoc
→ Raise maturity: the highest-leverage missing control is dependency pinning.

Every finding carries a full cross-framework tuple: OWASP LLM/ASI/DSGAI · MAESTRO layer · AIVSS · NHI · MITRE ATLAS/ATT&CK · CWE · NIST · regulatory.

The ecosystem

Two animals, one mission.

🐼 Grey Panda

The calm, deterministic, in-the-file guardian — regex + AST, OWASP-tagged, ships its own MCP server. The ground truth.

+

🐘 Maroon Elephant

The orchestrating, design-time, multi-repo threat-modeler that stands on top — architecture, governance, GitHub, and a UI.

Quickstart

Running in under a minute.

CLI

# scan a repo, ranked report in seconds
$ maroon scan .
$ maroon scan https://github.com/org/repo
# CI-ready: SARIF + exit codes
$ maroon scan . -f sarif -o results.sarif --fail-on high

Enterprise dashboard

# local web UI — scan an org, worst-first
$ maroon serve
# its own MCP server, for agents & IDEs
$ maroon serve-mcp

GitHub Action

Drop-in workflow → SARIF in the Security tab + PR annotations. Public repos free.

VS Code

Inline diagnostics with OWASP tags and the governance verdict in your status bar.

BYOK & air-gapped

Bring your own key (or a local model). Keys never leave your machine; air-gapped mode makes zero external calls.

Docs & standards

Honest about what it can — and can't — do.

Static analysis, Python-first deep coverage, governance checks that say “needs attestation” rather than faking a pass. Read the research and the spec:

Findings

The research synthesis — the full AI threat corpus & competitive landscape. Read →

Build Plan

The engineering spec — architecture, determinism contracts, roadmap. Read →

Can & Cannot

The honest limits, written before the code. Read →