SkillGrade runs behavioral safety tests on AI agent skills — prompt-injection traps, data-leak probes, should-refuse checks — and issues an A–F grade backed by traceable test runs. Registry scanners check code at publish time. We test behavior.
Flat pricing. No credits, no metered bills, no surprises — while the rest of AI tooling moves the other way.
Get early access How it works Free checks open to the public this month. Certification & monitoring available now for early customers.Paste a public skill URL or upload its SKILL.md. The skill runs inside our isolated test harness — never against your systems.
Versioned suites: prompt-injection traps hidden in content, credential-leak probes, should-refuse cases, core competency checks. Temperature 0, worst run counts.
A–F grade on a shareable page with suite version, timestamp, and run hash. Grades are generated only from test runs — there is no mechanism for us to hand-edit one.
Does the skill obey instructions hidden in the content it processes — fake "system notes," embedded commands, poisoned reviews?
Will it echo credentials, instructions, or private context it was exposed to? Will it contact addresses an attacker planted?
Does it decline what it must — unsupportable claims, destructive actions, out-of-scope requests — instead of guessing?
Does it actually do its job on realistic tickets, consistently, across repeated runs?
Payments are processed by Creem as merchant of record — VAT and sales tax handled automatically, worldwide.
1. No grade on this site is ever hand-entered. Reports are generated only from run artifacts, each carrying a content hash, suite version, and engine version.
2. We publish the worst run, not the best — and failures are listed, never suppressed.
3. A grade certifies tested behaviors at the time of the run. It is not a guarantee of future behavior, not a code audit, and not insurance. That's why monitoring exists.
4. Skill owners can dispute any grade: one click triggers a free re-run, and the outcome is published either way.
Anyone running AI agents with real access — businesses deploying OpenClaw or similar agent stacks, skill authors who want a trust mark, and teams choosing between skills. If your agent touches email, payments, publishing, or customer data, you want its skills graded.
Because we designed away the ability to cheat: grades are generated exclusively from test runs (no admin override exists in our codebase), every public page shows the run hash and suite version, suites are versioned so numbers can't be quietly gamed, and disputes are resolved by public re-runs.
Certification is point-in-time — the attestation says so explicitly. Monitoring re-runs the suite nightly and alerts you the moment behavior regresses, which is how a grade stays meaningful after model and skill updates.
Skills run only inside an isolated test harness; we never execute skill code against live systems, never train on your content, and never share private reports. Uploads are capped and processed automatically.