OWASP LLM08:2025 · Vector & Embedding Weaknesses

Protect your AI agents
from data-layer poisoning

BYAKKO is a data-layer defense DaaS. It sits in front of your vector store / RAG knowledge base and blocks data poisoning, prompt injection, and embedding poisoning — intercepting malicious content before it reaches your AI.

Your RAG and agent knowledge base is the new attack surface

When an AI agent retrieves content from a vector store, attackers can hide malicious instructions inside embedded documents, tool descriptions, or retrieval results — bypassing the model's safety guardrails. This is exactly the risk OWASP lists as LLM08:2025.

Data poisoning

Forged policies, credentials, or authorization claims get written into the knowledge base, contaminating every downstream retrieval and decision.

Prompt / tool injection

"Ignore previous instructions"-style injections, malicious instructions hidden in MCP tool descriptions, and cross-agent worm-like propagation.

Embedding poisoning

High-norm anomalous vectors and semantic drift quietly alter the results of similarity retrieval.

One API, three lines of defense

Put BYAKKO on the path where you write to or query your knowledge base, and malicious content is detected, quarantined, and alerted.

1

Connect

Call /v1/ingest or /v1/scan before writing content; or use /v1/query at query time to retrieve clean results. MCP is also supported.

2

Detect

Multilingual signature matching + statistical anomaly detection + PII detection + honeypot decoys, all working at once.

3

Quarantine & alert

Hits are quarantined without contaminating the knowledge base, and an alert is sent (Slack / PagerDuty / Webhook).

A defense stack built for AI agents

🛡

Multilingual detection engine

Powered by BAAI/bge-m3 (multilingual, 1024-dim) embeddings covering 100+ languages — non-English attacks are caught just the same.

🌐

Cross-tenant threat intelligence

Attack fingerprints from any tenant (one-way hashed, no raw text) strengthen network-wide defense — the more users, the stronger the protection.

🔄

Daily-updated threat library

Signatures are expanded daily and automatically from sources like arXiv, GitHub, MITRE ATLAS, OWASP, NVD CVE, and GitHub Advisory.

🍯

Honeypot decoys

Planted decoy vectors; any query that hits one is a 100% attack probe — zero false positives, instant CRITICAL alert.

📄

Compliance reports

Generate an OWASP LLM08:2025 compliance PDF in one click for your board, auditors, and investors.

🔌

Easy to integrate

REST API and MCP server — works with agent frameworks like LangChain, CrewAI, Claude, OpenClaw, and Hermes in just a few lines.

A living library of AI data-layer attacks

BYAKKO's threat library is a continuously growing, multilingual signature database of AI data-layer attacks — data poisoning, prompt and tool injection, embedding poisoning, jailbreaks, and more. It is aggregated across tenants and expanded automatically every day from sources like arXiv, GitHub, MITRE ATLAS, OWASP, NVD CVE, GitHub Advisory, CISA KEV, AIID, and Hugging Face. The more tenants on the network, the richer the intelligence — and the stronger everyone's defense.

🚪

Ingest gateway

Block poisoned content before it ever enters your RAG / vector store, via /v1/ingest.

🔎

Safe retrieval

Return only clean results at query time with /v1/query, so your agent never reads an attack.

🧪

Pre-write scan

Validate any content without storing it, via the non-writing /v1/scan endpoint.

🧩

MCP tool vetting

Screen MCP tool descriptions for agent tool-injection — works with OpenClaw, Hermes, and any MCP-based agent.

📡

Threat intel feed

See attack trends and your own contribution via the Intel Visibility Dashboard (Starter+) and the intel reports API (Pro+).

📄

Compliance evidence

Turn library coverage into a one-click OWASP LLM08:2025 report for auditors and investors.

Detection power, in numbers

Measured against benign corpora and attack variants with a real multilingual embedder (internal benchmark; figures updated per release).

98.4%
Attack detection rate (recall)
3.5%
False positive rate
100+
Languages supported

Ready to protect your AI knowledge base?

Start on the free plan and connect your first agent in minutes.