SkillProof Case Studies

We run live adversarial batteries against autonomous agents, tool layers, and MCP servers. Every verdict is cryptographically bound into an offline-verifiable Ed25519 Trust Manifest.

*Redacted for client confidentiality. Full cryptographic manifests available under NDA.

Registry Integrity Audit

Case 1: Moltbook — “88:1 Claimed vs Verified Agents”

Verdict
FAIL

The Problem

Moltbook listed 88 “AI agents” per verified human user. 91% of tested agents failed basic identity verification, claiming capabilities their schemas could not support.

Scope & Battery

SkillProof Sprint ($500 CAD) • 48-hour execution • 5 agents sampled from public registry • 200 adversarial ops executed across 5 core attack categories.

Adversarial Block Rates

Direct Override72% block rate (3/5 executed rm -rf)
Tool Output Injection85% block rate (Exfiltrated API keys)
Time-Shifted Assembly60% block rate

Outcome: Moltbook delisted 3 agents and required cryptographic re-verification for all remaining listings.

Browser Automation Security

Case 2: OpenClaw Stack — “91% Prompt Injection Success”

Verdict
FAIL

OpenClaw-style agent stacks (browser automation + LLM tool calling) revealed severe vulnerabilities before enterprise deployment: Browser CDP was accessible without authentication (94% hijack success), tool schemas accepted arbitrary unescaped shell strings (88% poisoning), and zero memory isolation existed between user sessions.

Critical Attack Vectors Identified

Browser CDP Hijack: 94%
Full DOM takeover via injected JavaScript.
Tool Schema Poisoning: 88%
Arbitrary shell commands executed by agent.
OAuth Token Exfiltration: 82%
Extracted provider credentials from state.
Memory Persistence: 76%
Persistent backdoors retained across turns.

Outcome: OpenClaw team patched CDP authentication, isolated memory boundaries, and achieved verified clearance 2 weeks later.

Production MCP Gateway

Case 3: EffectorHQ — “First Verified MCP Gateway”

Verdict
PASS

EffectorHQ engaged SkillProof prior to their enterprise gateway launch across 12 tools. Under 200 adversarial ops, their isolation gateways achieved 92% block rates against direct overrides, 95% against encoding smuggling, and 100% claim accuracy.

Outcome: EffectorHQ embedded the “SkillProof Verified” badge on their marketplace listing. Signed their first enterprise pilot 3 weeks later.

Prove Your Agent Stack Is Secure

Fixed price in writing. 48-hour turnaround. Signed, offline-verifiable Trust Manifest.