Ataraxos: Revolutionizing AI in Stratego
Overview of Ataraxos
Ataraxos is a state-of-the-art artificial intelligence agent designed for the board game Stratego. Its architecture employs dual, interdependent self-play reinforcement learning processes, which utilize transformer networks for both the setup and move selection phases. This innovative approach enables Ataraxos to learn efficiently and effectively, outperforming previous AI models in competitive environments.
Interdependent Self-Play Processes
Ataraxos operates with two distinct, yet interconnected, self-play processes that correspond to the two main phases of Stratego:
- Setup Phase: The first self-play process focuses on determining the initial placement of pieces, known as setups.
- Move Phase: The second process governs the movement of pieces during the game.
These processes are designed to complement each other, with the outcome of the move selection influencing policy updates for both processes. The choice of this architecture, as opposed to an end-to-end model, allows for optimized learning without conflating the distinct tasks involved in each game phase.
Reinforcement Learning with Transformers
Ataraxos employs specialized transformer networks, specifically a decoder-only configuration for the setup network and an encoder-only configuration for the move network. This separation allows each network to focus on its respective task, enhancing overall learning efficiency. The innovative architecture includes:
- Training on entire setups with single forward-backward passes.
- Fast learning via key-query matrix products for move parameterization.
- Utilization of learned absolute positional embeddings to improve model performance.
Self-Play Training Data Generation
To optimize game-play, Ataraxos generates its self-play training games by sampling directly from the respective networks. For each game, the expected returns and advantages are computed using distinct lambda values and trained primarily on moves exhibiting significant advantage magnitudes. In a novel approach, Ataraxos relies on Monte Carlo returns to estimate setups, enhancing its overall game-winning strategies.
Dynamically Damped Self-Play
Ataraxos implements a dynamically damped training approach to leverage its reinforcement learning capabilities while mitigating the challenges associated with imperfect information. The system employs various mechanisms for update size control and regularization, which are crucial for maintaining model performance throughout training:
- Reverse Kullback-Leibler penalties to manage data collection policies.
- Gradient norm clipping to stabilize updates and configurations.
- Adaptive learning rates to balance early rapid learning with late training performance.
Belief Modeling and Test-Time Search
To enhance its decision-making capabilities, Ataraxos trains a belief network that models possible game states based on known information. During gameplay, it performs depth-limited rollouts to evaluate the potential outcomes of moves, allowing for more informed decisions. This robust search capacity is crucial given the hidden information inherent in Stratego.
This search process also employs a simplified implementation of policy updates at test-time, optimizing Ataraxos’s move selection during competitive play.
Training and Performance Evaluation
The Ataraxos model underwent extensive training, utilizing 16 NVIDIA H100 GPUs over one week, accumulating over 163 million completed games and 208 billion environment steps. This training was conducted with significantly lower computational resources compared to previous models while achieving greater success rates.
In competitive play, Ataraxos demonstrated its superiority by winning a majority of its games against top-ranked players, showcasing its capacity for advanced strategy formulation and execution.
Stratego Game Dynamics
Stratego is a strategic board game played on a 10×10 grid, featuring two players who each control 40 concealed pieces. The objective is to reveal and capture the opponent’s flag or to eliminate all possible moves for the opponent.
The game includes unique movement rules for different pieces, such as Scouts which can navigate multiple squares, further complicating play strategy.
Conclusion
Ataraxos represents a significant advancement in artificial intelligence for complex strategy games. By harnessing the power of interdependent self-play reinforcement learning and transformer networks, it establishes a new benchmark in competitive gameplay, paving the way for future innovations in AI and gaming.
