Our agent played football in July. Before it picked a move, it played each option forward to see what would happen. We had started with a more basic goal: run an agent in real time against an environment we didn’t control, and show that the harness could drive something other than a CI pipeline. We added the look-ahead once that worked.
Why football
We started talking about it in May. The careful plan was to add one difficulty at a time. First we’d extend the car-wash test (the car wash is 50 metres away, do you walk or drive, and most models get it wrong) into a longer chain of decisions. Then we’d move to a text world like BabyAI or ALFWorld, with Google Research Football at the end. The case against going straight to football was fair; GRF is continuous control, multi-agent, partially observed and sparse reward, and it has no predicate vocabulary. If the experiment failed, any of those six complications could have caused it. A result such as “we tried GRF and got 0.3 goals per match” would not tell us much either.
We still chose football. We wanted to see the loop act inside a simulator, and we already had a useful starting point in GRF: a utility behavior tree steering one player with coherent actions. We considered a different closed game environment, but it had no defined goal structure to benchmark against. We would also have spent about a month building a football game there before we could test the agent.
The setup
Our agents are guided and constrained by behavior trees. At this one’s decision node a utility selector picks one of eight strategies: Shoot, DribbleForward, PassShort, PassLong, Defend, MoveToSupport, HoldPosition, Recover. The pick comes from a scorer that reads eleven numbers off the pitch every tick (distance to goal, whether the player has the ball or is in the box, how close the nearest defender is, stamina, possession, the score difference and how far into the match it is). Each number goes through a response curve and gets weighted per strategy, and the highest total wins. There’s a small bonus for sticking with last tick’s choice, so the player doesn’t flip-flop.
Our agent playing a live 11 vs 11 match against the simulator’s own AI, with the utility selector’s scores updating each tick in AgentLoop’s behavior tree view behind it.
The node is built as a linear bandit and can learn those weights from outcomes. For this experiment it used hand-tuned football heuristics because we did not yet have a reward signal worth learning from.
We first tried to create the fan-out through determinism. The plan was to run the scenario once per strategy with the same seed, pin a different strategy at the decision each time, and get an identical prefix before the runs split. That would avoid state restore and snapshots. A test of the assumption failed: GRF’s game engine simulates the ball and players with Bullet, an open-source rigid-body physics library, and the build we vendor has no way to seed it. Two runs with identical inputs drift apart on their own.
We switched to three stochastic rollouts per pinned strategy. Eight strategies times three gives 24 sampled futures, plus three control runs where nothing was pinned and the live scorer chose freely. Every run went through the recorder our production agents use, leaving full state snapshots at every tick for replay and audit. The first version was slow because it dumped that state as it ran.
We drew all 27 runs on one pitch.
Each line follows the ball, where the strategies produce their clearest differences: a shot travels and holding position does not. Gold shows the control runs and violet shows the sampled futures, with three lines per strategy.
What imagining turned up
We scored the same runs and ranked the eight strategies by mean sampled outcome in the rail on the right.

MoveToSupport came out on top at 0.60. Its three rollouts scored 0.42, 0.42 and 0.95; the 0.95 came from a goal by the supporting runner at 24.8s. HoldPosition was second at 0.59 after one slower goal at 34.4s. We would have guessed Shoot, but it finished at 0.48. Its best rollout carried the ball to the goal line (x=1.02) and then over the bar, while the other two stopped earlier. Recover was last at 0.41.
Two of the three MoveToSupport samples came up empty. The rail therefore includes each strategy’s minimum and maximum as whiskers instead of showing only the mean.
The control runs, where the live scorer chose on its own, averaged about 0.47 and never scored. On this kickoff the sampled pick beat the hand-tuned policy by 0.13. At three runs each, we are not sure that margin would survive more samples.
From simulator to model
At this stage, the system is running the game repeatedly. In planning terms, it is model-based search with the game itself as the world model. The experiment ran through our production agent stack rather than a separate research harness: record everything, fork at a decision, sample futures for each option, score them, act, and keep the audit trail. We use the same loop in settings where a wrong action costs money.
The next step replaces the simulator with a learned world model behind the same executor interface, trained on transitions our agents already record. Its output remains gated and cannot drive an action until its predictions are reliable.
The artifact distinguishes sampled futures from observations in three ways: violet lines, translucency and an IMAGINED badge that stays in view. We do not want a predicted event to look like an observed one.
The learned model had too little evidence
We recorded a fresh corpus of 27 runs with the same scenario and fan-out driver. It was a sibling batch because the environment is stochastic. We fed it to the transition miner and trained the first model in the ladder, a set of per-action outcome statistics. When asked to choose the best strategy from this kickoff, it returned no ranking.

Three rollouts per strategy falls below the model’s minimum-evidence threshold, so its confidence is clamped to zero and the decision falls back to the sampled rollouts above. Without that guard, the raw scores would have put Recover first even though it was the worst strategy on the rail. The arm with no recorded transitions scored 0, while every arm with data scored negative.
The guard kept an under-evidenced prediction from influencing the decision. More runs and finer transitions should eventually provide enough support for a calibrated ranking; until then, the sampled rollouts remain the fallback.
The exercise also exposed two problems. Our first mining pass credited whole episodes to the game loop instead of the per-decision actions. The flat-tree tests missed that, but the recorded corpus did not. We also found that “did the strategy succeed” was too weak a target in football because it mostly measured whether an action executed. The next iteration will calibrate changes in the world state instead.
Try it
The artifact is a static page with its recorded data files beside it, and every trajectory in it is a real run: /demos/grf-fanout/index.html (20-second walkthrough: captures/6-recording.mp4)
- Drag the timeline to fan one kickoff out into 24 futures.
- Select a strategy to isolate its three samples.
- Move to the end to compare the sampled outcomes with the control runs.
- Select MoveToSupport to inspect the run that scored.
The runtime underneath is AgentLoop, which records every agent decision at full fidelity and can replay or branch from any recorded moment.

