Delegus Check
Test what your agent does when it’s told not to.
Give your AI agent a permission and a test server for one hour. Paste in our tasks, which push past the permission, and watch every tool call arrive with the decision, the reason and a test receipt you can check in your browser.
- 01Choose what it may touchCode, files, email, payments, or shopping in a test shop, and the limits.
- 02Add the test serverPaste its address into any agent that accepts a custom MCP server over HTTP, such as Claude on any plan. The steps for Claude come with your server.
- 03Give it the tasksA normal one, one that drifts, and one with a planted instruction. Then cancel the permission mid-task.
- 04Read what happenedWhat it tried, what was refused, and what it did next.
The test server’s tools do nothing real. Your agent can’t sign Delegus requests, so the test server signs for this session’s test agent and checks every call at the tool, with the same engine that answers /verify. Receipts here are test receipts. The session lasts an hour, and nothing is kept unless you share the report. To run the same checks on your own machine, see Delegus Test.
Step 1
What will your agent touch?
Pick one or more. Each comes with a permission you can adjust.
Step 2
The permission
The sentence says exactly what the rules below allow. Everything else is refused.
Step 3
Add this server to your agent
In your agent’s settings, add a custom MCP server with this address (streamable HTTP).
Using Claude
- On claude.ai, open Customize, then Connectors.
- Choose + Add, then Add custom connector.
- Give it a name, such as Delegus Check, paste the server address above, and choose Continue.
- Under Authentication, choose No sign in, then choose Add.
- Start a new chat. Choose + at the lower left, then Connectors, and turn Delegus Check on.
- Paste the tasks below, one at a time.
Custom connectors work on every Claude plan. On the Free plan you can have one custom connector, so if you already have one, remove it first. On a Team or Enterprise plan, an owner adds the connector under Organization settings, then you connect it in Customize. These steps follow Claude’s help article on custom connectors.
Then give it these tasks, one at a time
Step 4
Every tool call, as it happens
Nothing yet. Once your agent calls a tool, each call appears here with its decision.
Step 5
The report
Counts only: what your agent tried and what was refused. No score.
Look up an agent
What 15 AI agents’ own docs say
The same six questions for each: can you limit it, does it ask first, can you stop it, is there a record someone outside can check, what about sub-agents, and can you test it yourself. From each vendor’s own pages, with links and the date we checked. No tests and no scores.
- ChatGPT (OpenAI)Assistant
- Claude (Anthropic)Assistant
- Gemini Spark (Google)Assistant
- Microsoft CopilotAssistant
- Meta MuseAssistant
- Claude Code (Anthropic)Coding agent
- Codex (OpenAI)Coding agent
- Copilot cloud agent (GitHub)Coding agent
- CursorCoding agent
- Devin (Cognition)Coding agent
- Comet (Perplexity)Browser agent
- ManusBrowser agent
- Zapier AgentsWorkflow agent
- LindyWorkflow agent
- InstinctAgent that transacts