OWASP Top 10 for LLMs: Turning the List into Real MLSecOps Pipelines
Sécurité

OWASP Top 10 for LLMs: Turning the List into Real MLSecOps Pipelines

September 15, 20267 min readOWASPLLMMLSecOps

The OWASP Top 10 for LLM Applications 2025 decoded risk by risk, and how to turn it into real MLSecOps pipelines: model signing, CI red-teaming, admission control and runtime guardrails.

Why LLMs needed their own Top 10

The classic OWASP Top 10 assumes a world where instructions and data travel on separate channels: code is code, user input is a string you validate, escape and parameterise. Large language models break that assumption at the architectural level. The system prompt, the retrieved document, the tool output and the user question all arrive as one flat token stream. The model has no structural way to know which part is authoritative.

That is why the OWASP GenAI Security Project published a dedicated Top 10 for LLM Applications, first in 2023 and then substantially reworked for 2025. The refresh is worth reading in its own right: the community dropped some 2023 entries (model theft, overreliance as a standalone item) and promoted issues that only became visible once teams shipped RAG pipelines and autonomous agents to production — system prompt leakage, vector and embedding weaknesses, unbounded consumption.

The 2025 list, and the control that actually moves the needle

IDRiskWhat it looks like in productionHighest-leverage control
LLM01Prompt InjectionA retrieved web page or support ticket carries hidden instructions the agent obeysLeast-privilege tools, human confirmation on side effects, trust boundaries per data source
LLM02Sensitive Information DisclosureSecrets or another tenant's data surfacing in a completionRedaction before inference, permission-aware retrieval, output scanning
LLM03Supply ChainUnpinned model weights, malicious pickle, compromised adapter or pluginInternal model registry, signature verification, safetensors-only policy
LLM04Data and Model PoisoningPoisoned fine-tuning corpus, backdoored LoRA adapterDataset provenance, lineage tracking, backdoor evals
LLM05Improper Output HandlingModel output executed as SQL, shell, HTML or Markdown linkTreat every completion as untrusted input; encode, validate, allowlist
LLM06Excessive AgencyAgent with write access to a CRM, a mailbox and a Kubernetes clusterScoped tokens per tool, no destructive verbs, approval workflows
LLM07System Prompt LeakagePrompt extraction revealing business rules or internal endpointsNever put secrets or authorisation logic in the prompt
LLM08Vector and Embedding WeaknessesCross-tenant leakage in a shared collection, embedding inversionTenant isolation, filters applied at query time, ACL reindexing
LLM09MisinformationConfident hallucinations consumed by a downstream automated processGrounding with citations, confidence gates, human review on high impact
LLM10Unbounded ConsumptionRecursive agent loops burning six figures of tokens over a weekendToken budgets per tenant, hard step limits, spend alerting

Prompt injection is not a bug you patch

The most common mistake we see is treating LLM01 as a filtering problem. Teams bolt on a jailbreak classifier, watch it block the obvious ignore previous instructions payloads, and declare victory. Then someone plants instructions in a PDF, a Jira comment, an HTML comment on a crawled page, or base64 inside a code block, and the agent happily exfiltrates data through a Markdown image URL.

Indirect prompt injection is an architecture problem. The practical question is not "can I stop the injection?" but "what is the worst thing that happens when it succeeds?" If the answer is "the agent deletes production records" or "the agent sends an email with the retrieved context to an attacker-controlled address", filtering was never going to save you. Design for injection the way you design for a compromised pod: minimal scope, no ambient credentials, egress allowlists, and irreversible actions gated behind a human or a signed policy.

Output handling and excessive agency: where incidents actually happen

LLM05 and LLM06 compound each other. A completion that is executed without validation, combined with an agent that holds broad credentials, is the GenAI equivalent of running a web app as root with string-concatenated SQL.

# Anti-pattern: model output executed verbatim
sql = llm.invoke(f"Convert to SQL: {question}")
db.execute(sql)   # one indirect injection away from a full dump

# Pattern: structured intent + allowlist + server-side authorisation
intent = llm.with_structured_output(QueryIntent).invoke(question)
if intent.table not in ALLOWED_TABLES:
    raise PermissionError(intent.table)
rows = repo.select(
    table=intent.table,
    filters=intent.filters,          # typed, validated
    tenant_id=session.tenant_id,     # never model-provided
    limit=min(intent.limit, 500),
)

The rule of thumb: the model chooses intent, your code enforces authority. Anything the model emits — SQL, a file path, a tool name, a URL, HTML — is user-controlled input and gets the same treatment you would give a query parameter.

RAG: the weak link is usually the vector store

LLM08 is the entry that most surprises teams. A vector database holding embeddings of your entire Confluence, Salesforce and HR drive is a de facto secondary copy of your data — with none of the ACLs. Two failure modes dominate.

First, post-filtering. Retrieving the top-k chunks and then dropping the ones the user cannot see leaks through relevance signals and returns thin, inconsistent context. Filters must be applied at query time, using the caller's identity, with metadata written at ingestion. Second, stale ACLs. Permissions change in the source system and nobody reindexes, so a document remains readable in the RAG long after it was restricted. Build the reindex path before you build the chatbot.

MLSecOps: turning the Top 10 into pipelines

MLSecOps is simply the recognition that a model is a build artefact with a supply chain, a deployment gate and a runtime posture — like a container image, except that it is opaque, stochastic and frequently downloaded from the internet. Four planes matter.

Supply chain. Model weights are executable. Legacy PyTorch checkpoints are pickles, and loading one runs arbitrary code. Enforce safetensors, scan artefacts with a model scanner, pin Hugging Face revisions by commit SHA rather than branch, and mirror everything into an internal registry — OCI artefacts in Harbor or Artifactory work well and give you the same provenance tooling as images. Sign with Sigstore/cosign and publish an ML-BOM (CycloneDX supports machine learning components) alongside your SBOM.

Build and CI. Security evals belong in the pipeline, not in an annual pentest. Automated red-teaming tools such as garak, Microsoft PyRIT or promptfoo can be run as jobs that fail the build when jailbreak or leakage resistance regresses against a baseline. Treat the results as any other test artefact.

name: llm-security-gate
on: [pull_request]
jobs:
  redteam:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - run: pip install garak promptfoo
      # Probe the candidate config for injection, leakage, toxic output
      - run: garak --model_type rest --generator_option_file gen.json \
               --probes promptinject,leakreplay,encoding \
               --report_prefix artifacts/garak
      # Golden-set evals with an assertion threshold
      - run: promptfoo eval -c promptfooconfig.yaml --fail-on-error
      - uses: actions/upload-artifact@v4
        with: { name: llm-security, path: artifacts/ }

Deployment. The same admission control you use for images applies to models. A Kyverno or OPA policy can refuse any inference workload whose model artefact is unsigned, pulled from an external registry, or missing a risk-tier label. Pair it with a default-deny NetworkPolicy and an egress allowlist: that single control turns most successful injections from a data breach into a logged failure.

apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
  name: require-signed-models
spec:
  validationFailureAction: Enforce
  rules:
    - name: verify-model-artifact
      match:
        any:
          - resources:
              kinds: [Pod]
              namespaces: [ml-serving]
      verifyImages:
        - imageReferences:
            - "registry.internal/models/*"
          attestors:
            - entries:
                - keyless:
                    issuer: https://token.actions.githubusercontent.com
                    subject: https://github.com/acme/model-registry/*

Runtime. Layer guardrails on both sides of the model: an input classifier, an output classifier (Llama Guard, NeMo Guardrails or a managed equivalent), PII redaction before the request leaves your perimeter, and secret detection on both directions. Instrument everything with the OpenTelemetry GenAI semantic conventions so prompts, tool calls, token counts and latencies land in the same traces as the rest of your platform. Then wire token budgets per tenant and hard step limits per agent run — LLM10 is where security and FinOps converge, and it is the control most often missing.

Governance: the inventory comes first

NIST AI RMF, ISO/IEC 42001 and the EU AI Act all push in the same direction: you must be able to enumerate your AI systems, their purpose, their data and their risk tier. In practice the blocker is rarely the control framework — it is shadow AI. Teams cannot certify what they cannot see, and a surprising share of GenAI usage in large organisations runs through unmanaged API keys and personal accounts. An egress policy plus a gateway that every model call must traverse is worth more than three months of policy writing.

A pragmatic 90-day plan

If you are starting from zero, the sequence that delivers the most risk reduction per sprint is: (1) build the inventory of GenAI use cases and route all traffic through a single gateway; (2) classify each use case by blast radius, not by hype, and cut agent privileges to the minimum; (3) add the supply-chain controls — safetensors, scanning, internal registry, signature verification at admission; (4) put automated red-teaming and golden-set evals in CI with a failing gate; (5) close the loop with tracing, token budgets and an incident playbook that treats a successful injection like any other credential compromise.

None of this requires exotic tooling. The uncomfortable truth of MLSecOps is that most of it is the DevSecOps you already know, applied to an artefact class your platform team has not yet onboarded — plus one genuinely new primitive: never trust a model's output with authority it should not have.

← Back to blog