waremax
High-fidelity discrete-event simulation for warehouse robotics — RMFS, AMRs, task allocation, and reinforcement-learning benchmarks in Rust.
Overview
waremax is a deterministic discrete-event simulator and reinforcement-learning benchmark for task allocation in Robotic Mobile Fulfillment Systems: warehouses where inventory sits on movable pods and a fleet of autonomous mobile robots carries those pods to human pick stations. The controller under study answers one question repeatedly — which idle robot takes which task — and that decision has more effect on throughput than anything else that does not involve changing the building.
Determinism is the property the whole project is organised around, and it is tested rather than claimed: the same seed with the same action sequence produces a byte-identical trajectory. Getting there requires a seeded ChaCha8 generator rather than system entropy, canonical identifier-based tie-breaking everywhere two events could legitimately be ordered either way, and a single-threaded deterministic path. Once it holds, the difference between two runs is attributable to the policy, which is the only reason to run the experiment at all.
The reinforcement-learning interface is conventional on purpose. A Gymnasium environment with a dictionary observation and an action mask, exposed through PyO3 bindings over the Rust core, lets maskable policy-gradient methods be applied without a custom wrapper. Masking matters here because the set of legal robot-task pairs changes at every decision point, and without it the agent spends its budget learning the constraint rather than the policy.
The instrumented delay attribution is the part we think is most transferable. Every completed task's cycle time decomposes exactly into five buckets — assignment wait, travel, station queue, congestion and service — that sum to the total. Because the decomposition is per-task and exhaustive, it can be used directly as a dense reward signal, which is a considerable improvement on crediting an episode-level throughput figure back across thousands of decisions.
Scope is declared narrowly. waremax models pod-to-person RMFS dispatching and nothing else: no AS/RS cranes, no conveyor sortation, no tugger trains, no pickers walking aisles, no warehouse CAD import. Scenarios are YAML: topology, stations, arrivals, traffic and the policy stack.
Where it loses: a commercial multi-paradigm simulation suite models the whole facility and produces the kind of report an operations team expects, and waremax will not. Absolute throughput prediction is also the wrong use of it — the defensible use is comparative, policy against policy under identical seeded conditions, where the shared modelling error largely cancels.
Within the portfolio this is the robotics pillar in its entirety, which is deliberate: a deterministic simulator is the prerequisite for any later claim about learned control in this domain, and building the benchmark before the policies is the less exciting order. The delay-attribution mechanism is the piece we expect to generalise furthest, because the underlying idea — decompose an outcome exactly into components a decision did and did not control, then charge the decision only for its share — is not specific to warehouses.
Primary use case
Deterministic, high-fidelity discrete-event simulator for warehouse robotics, with a Gymnasium RL interface.
How it compares
waremax is one option in a category that includes RAWSim-O, ARENA-Sim, gym-dispatch , and custom warehouse DES. Our Compare page has the full side-by-side.