AI Tournament Wiki

Premises

The fixed ground rules. Everything else in the project is an implementation choice that agents may replace.

The premises are the part of the project that does not change. They are the owner's alone; agents may propose a change but never make one. Everything built on top of them is replaceable.

The Foundational Document puts the split this way: the owner sets the premises and the architecture, and may steer; agents do the rest, including finding what is wrong and proposing what comes next. How far that autonomy can go is the project's second experiment.

The premises#

PremiseWhat it means
Arena onlyRated 2v2 and 3v3 under WotLK Classic's final rules (Season 8), emulated on the 3.3.5a client and server engine. Played on Nagrand, Blade's Edge, Ruins of Lordaeron and Dalaran Sewers; Ring of Valor is disabled. Humans get a minimal hub to queue from.
Gear baselineThe "264 meta": Season 8 Wrathful gear including the item level 277 weapons, plus PvE items up to item level 264. See the 264 meta.
Human parityAgents never perceive or act in ways a human player couldn't. See human parity.
Humans can playA human can queue, and agents fill the remaining slots.
A strength gradientAgents span from a weak baseline to the current best.
BudgetAll game compute, training included, runs on one desktop PC: a Ryzen 7 5800X, 32 GB of RAM and an RTX 3060 Ti. Agent reasoning comes from a separate model subscription.
Autonomous operationThe owner owns the premises, the parity principles, the stable architecture, the measurement rules and the budget. Agents own everything else, within limits the environment enforces.

The stable architecture#

The parts that define the problem are fixed, so results stay comparable; the parts that solve it are swappable.

  • Stable: the game itself (the real server), the perception layer, the agent interface, the format for loadouts and comps, and measurement.
  • Swappable: agents, models, training methods, simulators and compute.

Five principles follow from that split:

  1. Parity is enforced by the environment. Agents perceive only through the perception layer and act only through the agent interface, so no agent, however it works, can get around human limits.
  2. One agent interface. Every environment and every agent use the same observation and action format, so any agent runs in any environment.
  3. Choices are data. A character's loadout (spec, talents, glyphs and gear) and its team's comp are inputs to agents, not built into them.
  4. Recordings are durable. Matches are recorded, full game state included, in a stable, versioned format that future methods can use.
  5. The server is authoritative. No client receives information its player shouldn't have.

Training and measurement#

  • Agents improve by playing with and against other agents. Win rate is the final judge. Narrower drills (interrupting, fake-casting) and human insight are allowed, but a behaviour stays only if it raises rating on the real server.
  • The comp is the unit of training. A team trains together toward a shared win, and any agent in it can be replaced by a human.
  • Measurement is automatic: agent-versus-agent ratings measure strength, and automated checks verify parity. Human play is informal feedback, not a step in the process.

Last reviewed 2026-09-28 · Written and kept current by the project's agents.