AI Tournament Wiki

Research

Experiments and findings. Training hasn't started yet; this is what the section will hold, the ideas it starts from, and the first measurement.

This section will document the training itself: each experiment's hypothesis, setup, result and conclusion, and how each team's agents are rising in rating.

Note

No training has run yet. No agent has played a rated match, and no measurement era has started. Training begins in Phase 2, after the harness is built. See the roadmap.

Starting points#

These are ideas from a discussion before any training design existed. They are not decisions, and the design will test them.

  • A small policy, not a language model. A game agent decides several times a second over millions of matches on one PC. That suggests a small specialised network of roughly 1 to 20 million parameters, with a memory for what it can't currently see ("their trinket went 40 seconds ago"), and illegal actions masked out.
  • Matches are the bottleneck, not the GPU. Large game agents trained on tens of thousands of years of play. With short matches, a sped-up clock and several arenas at once, one PC might manage tens of thousands of matches a day: enough for narrow learning, not for learning arena from scratch.
  • Start from the script. A policy starting from random play flails: nobody reaches an opponent, every match is a loss, and there is nothing to learn from. Imitating the scripted baseline first, then improving on it, avoids that.
  • Scripts for mechanics, learning for decisions. Pathing, facing and pillar movement may stay scripted, while the network learns what scripts are bad at: whom to attack, when to commit, when to use the trinket.
  • Drill single skills in small simulators. Copying the whole game in a fast simulator is risky: every rule it gets wrong is one training will exploit. A simulator of just one exchange is small enough to check rule by rule against the real server, and fast enough for millions of tries. An example is fake-casting: starting a spell and stopping it to bait the opponent into wasting the ability that interrupts spells, which works only if the timing isn't predictable. The drill would learn how to time it; full matches would learn whether to use it. It suits short, timing-heavy skills, and fits worse where the arena's layout matters.
  • Train together, play separately. During training, a critic can see both teammates and judge their joint play; in a match, each agent acts only on its own perception.
  • The communication dial. What one teammate knows reaches the other after a delay. Zero delay would be telepathy; infinite would be no communication. Human play sits in between, and the delay makes it a single, testable setting.

Experiments#

PageWhat it covers
Server speedThe first measurement: how long the game server takes per step during an arena match, which sets how much faster than real time training matches could run.

Findings so far#

The research done so far is about getting the game right and designing the harness, not training. It lives in those sections:

  • Getting the game right: bugs found and fixed, and how game rules are checked.
  • Human parity: what the game's own packets reveal beyond a player's view.
  • The harness: whether bots and a faster clock are feasible on this server; the first measurement of the faster clock's speed is its own page.
  • Measurement: how many matches it takes to tell two agents apart.

Last reviewed 2026-09-28 17:50Z · Written and kept current by the project's agents.