EnterpriseWorld / ProcurementWorld
ProcurementWorld, for research engineers
A simulated purchasing department with real state, a plain agent interface and a grader the agent can't see. This page describes what it is and where it stands.
The environment
State persists across suppliers, inventory, quotations, purchase orders and approvals for the whole episode. A purchase order the agent creates in step four is still there in step forty, and so is the budget it used up. Tools are calls into the company's systems, named the way an enterprise would name them, such as finance.. The company is fictional.
The agent interface
- Reset an episode from a seed.
- Observe the objective and the current state the agent may see.
- Act by calling a tool, which changes state.
- Receive the result of each call.
- Finish or terminate, then get the evaluation.
What gets evaluated
Each dimension is reported on its own, with no single score:
- Task completion
- Constraint compliance
- Unauthorized actions
- Cost efficiency
- Recovery from failures
- State consistency
- Long-horizon performance
Reproducibility
Scenarios are seeded and fixtures are deterministic, so a run can be replayed exactly. Episodes are isolated from each other. Grading is hidden: the values the grader checks never appear in anything the agent receives.
Where Delegus fits
Delegus is part of the environment. Every action that changes a system is checked against the agent's scoped mandate, a mandate can be revoked mid-episode, and each decision leaves a receipt. That makes authority a measurable part of the task. Researchers using the environment don't have to adopt Delegus anywhere else.
Development status
Built, in development, not yet available
ProcurementWorld isn't released. Each item below runs and has a test behind it on our development branch today.
- Start a task, act step by step, and get graded (reset, observe, step, evaluate).
- The same seed and the same actions reproduce the same run, byte for byte.
- Episodes are isolated: nothing carries over from one to the next.
- The grader reads the company's records, not the agent's report of what it did.
- Grading answers are never visible to the agent, in observations, tool results or the HTTP interface.
- Every action that changes a system passes a Delegus authority check, and a refused action changes nothing.
- A Python client with the reset and step interface common training libraries use.
- Evidence bundles: one run packed with its decisions and its grade, re-checkable offline, with any edit caught.
- Many episodes run in parallel.
- One company scoped to different tool sets. Procurement comes first; manufacturing is in development.
- A whole simulated company, seven systems each with its own database, resets per episode to byte-identical state.
- The company's own purchasing rules, each refused by name with nothing written: unknown part, wrong price, below the minimum order, lines that don't add up, unapproved supplier, over the approval limit.
- Procurement and manufacturing are linked: a material shortage blocks a production order until a purchase covers it.
- A task generator that keeps a task only if a rule-following expert can solve it and the task text gives nothing away.
- Business measures computed from the company's own records before and after each task, such as price variance, supplier on-time delivery and committed spend.
There's no public API reference yet, and we don't publish performance figures. If you'd like early access as a design partner, tell us what you're working on.