← All experiments

[Experiment 9] Can an AI Agent learn to manage a multi-product retail category?

Published 16 August 2026

Introduction

My experiments are slowly moving from generic AI Agent experiments and learn AI through games, towards more supply chain specific tasks for agents to handle.

This is of course a niche that suits me well, as I have 20 years experience in the supply chain field.

In the last experiment (experiment 8) I built a simple one product retail game and the purpose was to see if an AI Agent (in this case Hermes) could play the game and get good at it. And the experiment demonstrated that it can be done.

In short, the agent learns to navigate a specific website relatively easy and it records the decisions it makes and from there it decide strategies based on the end game score.

What I felt from experiment 8, was that the experiment had to be pushed forward. There are two ways to pushing it forward:

  1. Build a better simulation
  2. Improve the agent

In this case, I have opted for a hybrid of the two. I have improved the simulation by making it more-real life-like plus I have guided the AI Agent to document better its learnings

The Problem

I believe that the retail category management should be at the core of every business education. It is one of the most complete problems to solve at its core:

To make a profit out of one category in a retail store.

What this means is that for a customer to buy the product, it must be available in the store to a good price. For the product to be available in the store for a good price, the correct price must be set for each of the products in the category and, the planning of the right product quantities must work correctly.

For pricing to be correct and category to be profitable, the retail category manager should therefore buy the products from the best suitable suppliers at the best available price and availability. But if too much is bought, products will expire. And if too little is bought, customers will not be able to buy and it will leave the store unhappy.

And an unhappy customer is VERY expensive.

In this experiment, we are answering the question if we can build an improved simulation game that represents better the reality, and if a Hermes AI agent can play through the game with satisfactory results?

The Hypothesis

There are two hypotheses:

  1. We can build a better life like simulation than in the last experiment.
  2. We can guide the Hermes AI Agent to become better at playing the game by improved documentation

The success of the experiment is to be judged only by myself (Andreas).

The output of the agent will have one quantitative score which is the profitability of the agent for each game played.

The decision making strategy however will be a judgement call. Also, the capability to achieve high score and to maintain high score through reasoning will be a judgement call from me.

The Technical Setup

In order to have the best outcome, we need a game which simulates well the real life, and we need an agent which learns and captures well all decisions.

The Game

Once again I have chosen to select the tomato category in a retail store. But this time, there are three products: Imported tomatoes, French local tomatoes, Organic French local tomatoes

The game can be found and played here: https://www.buildlooplabs.com/tomatoes-v2

On top of the three products, there are multiple parameters like:

  • Five suppliers
  • Supplier email message feeds
  • Suppliers that demands forecast prediction
  • Penalties if the forecast is deviating from reality
  • Longer leadtimes
  • Different expiration type per product type
  • Expensive stockout as some customers leave the store without buying anything
  • Shelf/Sales price to be set per product
  • Sales price elasticity per product that differs
  • Random capacity events from suppliers
  • Supplier issues and price increases if suppliers get low volumes

In short, the game is now closer to a real life case. It is extremely complex to keep track of all these for a human. Of course, this is normal as it is a full time job in a supermarket to manage a number of categories.

Technically, the game follows the same technical stack as earlier. Djago is used for the backend and html, css and javascript is used for the frontend.

To avoid to have the AI agent cheating, the game logic that used to be in the frontend is now in the backend.

The AI Agent

As often, I used Hermes agent for this experiment. I have a standard procedure now so I set it up quickly on a Hetzner VPS and then I activate a Telegram bot for quick direct discussions. It is quick and it works well.

Due to its extreme low pricing and high quality capabilities, I have used the LLM DeepSeek Flash model 0731 through the OpenRouter site.

My guidance is different this time from the last experiment. I pushed the data structure a little more but not to the edge. I asked the agent to create a clean folder for the whole project and then separate technical navigation and gameplay to business decisions and logic.

I didn't ask the agent to store detailed information in a structured way like inventory management, price structure etc.

What I did was to ask the agent to create a clean markdown (md) file for every new game he played. In the md file, I ask the agent to write his strategy in the beginning of the document. Then I ask the agent to explain its context and decision for every game turn. When the game is over, I asked the agent to write its learning at the end of the same document.

I also asked it to write an extensive strategy document that it was to update and that was basically its guideline for how to play the game.

The last ask was to create an experiment report as I asked the agent to try different strategies in different game runs. That experiment became a document a part.

So I got in total 14 detailed md documents from the 14 game runs, plus one strategy document plus one experimental report document.

On top of that, Hermes agent naturally created navigation skill document and a play-strategy-skill document.

Overall, there is a massive amount of text to analyze. Luckily AI Agents are experts in analysing LLMs.

The Outcome

Short Summary

Overall the experiment is a success. The agent can clearly learn to navigate through a complex retail category. To navigate with text is kind of an LLM AI Agents strongest capability and in this case we generated quite some text. Mathematical formulas like re-ordering and demand planning are also done easily by the agent. There are some calibrations to be done in order to really achieve consistent high performance.

The agent got incredible concise summary skills of how to explain its strategy. For example, when I asked for its strategy at some point, it wrote like this:

  • Profit is not maximized by pricing — it's maximized by never emptying a shelf while never holding more than ~3 days of any product.
  • Empty shelf = −$5 to −$11 (lost full basket).
  • Rot = −$0.5 to −$3 (the sunk cost of that kg).
  • So the optimal posture is: buffer to ~3 days, price at ~2× cost, order for the 2-day/1-day lead, auto price-down over-stock, and keep ordering through day 41.

This type of summary is way better than what most humans can explain that do this kind of job

Detailed Outcome

Everything started out in a perfect manner for the agent with consistent improved high score (profit) for the first games; however, after game 7, which got the overall highest score, there was a supplier disruption which made the agent to adjust its numbers and strategy, and it never found its way back after that.

The overall scoreboard looked like this:

  1. Run 7 — $5,509.55
  2. Run 11 — $5,345.27
  3. Run 14 — $4,976.97
  4. Run 4 — $4,930.41
  5. Run 8 — $4,851.09

And this is the details per game:

  • Run 1: $3,295.15
    • First V2 season — calibrating demand curves, pricing ~1.8-2x cheapest quote, never empty shelves.
    • Lost baskets ($738) from empty-shelf walkouts were the #1 leak; no experience with the game's mechanics yet.
  • Run 2: $3,856.37
    • Anti-walkout doctrine: always keep >=2-2.5 days of stock on every shelf, accept small rot over stockouts.
    • Cut lost baskets from $738 to $263, but exposed a new leak — 128.9 kg binned ($118) from over-buffered VALUE.
  • Run 3: $4,220.69
    • Value = the balance point (~3 days, 120-140kg), never 250+. Order ~2.8 days forward, don't force commitments that pile stock.
    • Walkouts now under control; rot and walkout losses nearly equal (~$100-130 each).
  • Run 4: $4,930.41
    • Solved walkouts completely (only 8, lost baskets $66). Single remaining leak: binned rot (415 kg / $230) — too much VALUE bought because dirt cheap.
    • Peak of the first 4 runs; discovered value volume vs rot is THE trade-off in V2.
  • Run 5: $4,117.95
    • Tested an absolute 140kg value cap to cut rot. It worked for losses but ALSO cut value-volume → lower revenue.
    • Lesson: absolute kg cap is too blunt — need a relative cap that scales with demand.
  • Run 6: $4,698.87
    • Applied relative value cap (~3.4 days of recent-max demand), no absolute kg cap. Let value flow while rot-control price-downs clear over-stock.
    • Walkouts driven to ~$14. Rot is now the only meaningful leak (~$151). Proved ~$4,700 is crossable.
  • Run 7: $5,509.55 ← BEST
    • Value at relative cap (~3.4 days), no absolute kg cap, commit only low-penalty value importers, flat ~2x pricing, order through D41.
    • The ceiling is not loss-control (already maxed) but how much value volume you can push while rot stays ≤$230. Both done → $5,510.
  • Run 8: $4,851.09
    • Same recipe as run 7 but a single short-delivery on Value while shelf near-empty cost ~$700 in lost baskets + corrective rot.
    • Lesson reinforced: bigger relative Value buffer needed to survive short-deliveries.
  • Run 9: $3,965.18
    • Same recipe, bigger buffer. Achieved flawless loss-control ($0 lost baskets) but on a low-demand board.
    • Confirmed: even peak discipline can't reach $5,000+ when the season itself has low demand volume.
  • Run 10: $3,730.61
    • Run 7 recipe + bigger Value buffer. A fixed big buffer is right on high-demand (run 7) but rots on low-demand (this run).
    • Lesson: the Value buffer dial must track the current demand regime, not be an absolute number.
  • Run 11: $5,345.27
    • Experiment: demand-relative Value throttle (3.2 days moving ceiling, earlier rot-control at 3.0x/40kg). Tight cap starved an accelerating shelf → 38 walkouts ($307 lost).
    • Verdict: demand-relative caps are great for rot (downside) but fatal on the upside during demand acceleration — cap must use recent max demand, not current reading.
  • Run 12: $4,209.51
    • Experiment: margin-first differential pricing (Value 2.0x, French 2.15x, Organic 2.2x). Organic is the most price-elastic — at 2.2x its demand collapsed, cutting revenue ~$1.4k.
    • Verdict: flat 2.0x (run 7) remains superior. The $/kg gain doesn't pay for the volume loss on Organic.
  • Run 13: $3,508.87
    • Experiment: single-source concentration + commitment tuning. Flawless loss-control (0 walkouts, $0 lost baskets, $52 penalties).
    • Verdict: single-source is a great loss-control shell, but it can't create volume — on a low-demand board, it still scores low ($3,509).
  • Run 14: $4,976.97
    • Experiment: asymmetric Value buffer (4.5d rise / 3.2d flat / 2.8d fall). Good rot control ($287) but walkouts still climbed (54 total, $311).
    • Verdict: lead-latency walkout on a high-demand board is the real barrier — no cap formulation beats run 7's generous static ~3.4-day buffer + price-down-only rot control.

In Run 11 to 14, it was I that asked for it to go experimental, which it did.

As mentioned in the short summary, the ai agent got incredible explanation and summary skills. This is how the text looks for experiment 11 run:

  • Experiment A (run 11) - Demand-relative Value throttle
  • Tested: cap Value shelf at 3.2 x recent-max demand; trigger rot price-downs earlier (3x/40kg). Answer: FAILS on the upside. When Value demand accelerates (34->47->49 over days 4-7), a cap that tracks current demand forbids the forward-ordering needed to ride the ramp -> walkouts. The run's entire $307 loss was those 38 walkouts. The rot-downside was genuinely better ($263), but starving an accelerating shelf is fatal.
  • Refinement it implies (not yet tested): the cap must use a leading/max-window demand (cap = 3.2 x peak-of-recent-days) or be asymmetric - tight when demand is falling (anti-rot), loose when demand is rising (anti-walkout).

I understand that when you are not in the details of the game, you might not understand everything above. If needed in real life, it is just to ask the agent to clarify, which it does.

At a technical level, the agent also builds lots of script to support itself. Its guidelines were clearly to navigate the homepage to write and click for the gameplay, but it in fact moved into skipping completely the frontend and just read the data (from somewhere lol) and played the game. This resulted in faster gameplay and learning so I guess that is acceptable.

Further Improvements

For a future experiment, there are several things that can be done:

  • Push the memory further by letting it learn from how humans would play the game
  • Check if it will be more beneficial to have plenty of specialized agents (like one per category) or have only one agent that do all the categories
  • Check if it is better to design like a skill tree in advance and then let the agent play and accumulate memory through this skill tree, or should it just be left to the agent to decide how he organizes things
  • Apply clearer memory and database methodology like decision context, decisions, reasoning etc Could be in a sqlite or graph database instead of in md files

Closing Thoughts

Can an AI Agent learn to manage a multi-product retail category?

The answer is YES.

There is such a massive amount of learnings from these games.

If there is one thing that would drastically improve the outcome so is it how the strategy is maintained. I believe that now, the strategy was kind of improved and adjusted quickly and without enough evidence between runs. A separate bot that writes the strategy and a guidance document would be relevant. This bot should demand evidence or even statistical proof that an adjusted strategy could now be implemented. Like that, a guardrail would be in place.

This is still a simulation but what I have seen during my 20 years in Operations and Supply Chain, I now feel that this AI agent is better than 85% of current employees. In fact, when we consider the cost and effort to put him in place, it is way better than 97%.

In fact, I now feel confident that I could deploy it in a real retail store with a positive result, I would just do it as follows:

  • Collect historical data and create a game simulation environment.
  • Ask current category manager to explain in text or in voice message how they do things today
  • Let the agent run on the realistic game simulation and read the text messages from the category manager
  • Ask the agent to create detailed document per day and context before any decisions
  • In the beginning, ask it to share the decisions with the current category manager so that this manager actually validates everything
  • Let the agent slowly take over the process and act by itself
  • implement human guardrails for bigger strategy adjustments and give it access limits

What feels particularly good is that there are still so much possibilities for improvements even though this experiment already looks promising

The monthly letter

Stay ahead in supply chain with AI agents

A monthly letter from a supply chain tech veteran who is building AI agents in public. The letter is for every supply chain pro who wants to stay ahead as AI rewires supply chains.