Donut Shops Inventory

Reorder stock across three stores. Press start and watch an untrained planner become a sharp forecaster.

Simulated single store replenishment for learning not a forecast or procurement recommendation for your live network.

Donut Shops Inventory2DIllustrative preview
episode 1 · step 0 · 0 fps

This is a live preview while the agent learns. Head to Replay to watch the fully trained replenishment policy in action.

Now Managing: — idle —Day: 0Run Profit: $0.00Fill Rate: 76%
S1Downtown · 1.0× traffic
0 / 1870 units
Glazed
MISS 3
0
Sprinkle
PICK 20
0
Chocolate
PICK 21
0
Boston
PICK 22
0
Jelly
MISS 7
0
Bacon
PICK 10
0
Sausage
PICK 11
0
Turkey
PICK 12
0
Coffee
MISS 11
0
Coffee
PICK 6
0
Coffee
PICK 2
0
Hash
PICK 7
0
Blueberry
MISS 15
0
Orders in transit
3 live
Glazed Donut+24
Sprinkle Don…+20
Chocolate Do…+21
S2Suburban · 0.8× traffic
0 / 1637 units
Glazed
MISS 5
0
Sprinkle
PICK 20
0
Chocolate
PICK 21
0
Boston
PICK 22
0
Jelly
MISS 9
0
Bacon
PICK 10
0
Sausage
PICK 11
0
Turkey
PICK 12
0
Coffee
MISS 13
0
Coffee
PICK 7
0
Coffee
PICK 3
0
Hash
PICK 8
0
Blueberry
MISS 17
0
Orders in transit
3 live
Glazed Donut+24
Sprinkle Don…+20
Chocolate Do…+21
S3Airport · 1.2× traffic
0 / 2666 units
Glazed
MISS 7
0
Sprinkle
PICK 23
0
Chocolate
PICK 24
0
Boston
PICK 25
0
Jelly
MISS 11
0
Bacon
PICK 12
0
Sausage
PICK 13
0
Turkey
PICK 14
0
Coffee
MISS 15
0
Coffee
PICK 8
0
Coffee
PICK 4
0
Hash
PICK 10
0
Blueberry
MISS 19
0
Orders in transit
3 live
Glazed Donut+24
Sprinkle Don…+23
Chocolate Do…+24

Inventory Movement Replay Timeline

On-handDemandSold0 days
0.40.81.1Waiting for replay days…

Same idea for retail and supply chain teams

Static reorder rules leave money on the table through stockouts or spoilage. This playground trains an agent to order against stochastic demand — scored on service level, waste, and margin.

Frequently Asked Questions

How the inventory playground works, what the KPIs mean, and how this connects to production RL.

A single-store, multi-SKU replenishment shop (default: 13 donut-shop SKUs). Each hourly step the agent decides how much to order per SKU while demand, spoilage, and capacity interact. The goal is high service level with low waste and solid gross margin.

Discrete MultiDiscrete order quantities fit this action space. Compatible algorithms come from the RLX discovery API for env=inventory — typically on-policy methods such as PPO/SAC depending on registry support.

The shipped six: margin, service_level, waste_rate, plus summed sales / lostSales / waste across SKUs. Values come from kpi_summary after eval — do not recompute client-side.

No — this is a simulated playground. You can optionally upload a SKU catalogue via the inventory-datasets API; otherwise training uses the default 13-SKU donut shop set.

Ready to explore inventory RL in production?

We help teams move from simulation to governed deployment with monitoring, guardrails, and continuous optimization.

Classic CSV replenishment POC