Papers – Abstract

AuthorAbstract
Jonathan Lamontagne-KratzUndergraduate journal paper proposing Rainbow Zero, a hybrid of MuZero/Muesli-style search with Rainbow/Dreamer-style distributional RL and world models. Primary testbed is Catan (reduced-complexity randomized boards), with planned ablations vs Catanatron baselines. Framed as a proxy for stochastic real-world planning. Work is explicitly ongoing; full Catan results not yet reported.
Visakh MadathilA 48-game, 4-player tournament of prompt-only LLMs on Catanatron (no fine-tuning, MCTS, or RL): Gemini 3 Flash Preview, Claude Sonnet 4.5, Claude Haiku 4.5, Gemini 2.5 Flash. Gemini 3 Flash won 35/48 (72.9%, Elo 1696); others 16.7% / 8.3% / 2.1%. Traces show specific failure modes (VP-card mechanic hallucination, endgame paralysis, scratchpad calcification). Domestic trade acceptance ~8%; games averaged 163 turns vs ~70 for humans. Tournament cost ~$25 / ~13k API calls.
Anthony van TilburgA hub for Settlers of Catan enthusiasts: sharing and explaining Python code to play against AI players and improve these AI players.
Ali W. BekheetCustom full-rules engine (1,017 discrete actions, 90-dim features, action masking) comparing DQN, Ray RLlib PPO, TorchRL SAC, plus an experimental LangGraph LLM agent. Reward shaping and curriculum/self-play. Reported SAC ~28% vs random in 4p (barely above 25% parity); PPO ~23% after 5k episodes. Web UI for human vs AI. Companion Substack: “I Trained a Bot to Play Catan. It Barely Learned Anything.”
Nikolas Belle and othersWe address the long-horizon gap in large language model (LLM) agents by en abling them to sustain coherent strategies in adversarial, stochastic environments. Settlers of Catan provides a challenging benchmark: success depends on balanc ing short- and long-term goals amid randomness, trading, expansion, and block ing. Prompt-centric LLM agents (e.g., ReAct, Reflexion) must re-interpret large, evolving game states each turn, quickly saturating context windows and losing strategic consistency. We propose HexMachina, a continual learning multi-agent system that separates environment discovery (inducing an adapter layer without documentation) from strategy improvement (evolving a compiled player through code refinement and simulation). This design preserves executable artifacts, al lowing the LLM to focus on high-level strategy rather than per-turn reasoning. In controlled Catanatron experiments, HexMachina learns from scratch and evolves players that outperform the strongest human-crafted baseline (AlphaBeta), achiev ing a 54%winrate and surpassing prompt-driven and no-discovery baselines. Ab lations confirm that isolating pure strategy learning improves performance. Over all, artifact-centric continual learning transforms LLMs from brittle stepwise de ciders into stable strategy designers, advancing long-horizon autonomy.
Cole Miller and othersAlphaZero-style agent (dual policy/value net + neural-guided MCTS, self-play) on a custom headless Catan engine. Simplifications: no player trading, no ports, random discards on 7s. 300 self-play iterations x 20 games. Iter-300 win rates in 4p games: 66% vs 3 random, 50% vs iter-50, 38/30/28% vs iter-100/200/250. Placement net later added (60% vs random). Compute-limited vs classical AlphaZero.
Wiktor PiszczekThe thesis aims to find the difference it makes for a simple agent in the game of Settlers of Catan to be able to communicate in game chat using natural language and persuade their opponents to make worse moves. In the experiments three types of models were tested on the colonist website against human players in a ranked game.From those models the agent that always performed the move which assumes that opponents will make a mistake fared the best. The model with and without chat capabilities performed similarly.The first one ended with slightly worse score due to unfulfilledp erceived affordances stemming from presence in chat leading to overestimating its strength in the eyes of human opponents and annoyance resulting from artificial sounding language.
Ozzy Simpson and othersUCB MCTS on a simplified pygame Catan: no P2P trades, no ports, no knights/robber, random opening placements, three random opponents. Stochastic transitions sampled rather than chance nodes. Backup: VP margin vs best opponent (+10 if win). Known bugs (negative resources, loops). No quantitative win rates.
G.J.B. RoelofsThis article describes the abstraction required for a framework capable of playing multiple complex modern board games, and provides a proof-of-concept implementation of Settlers of Catan within said framework. Monte Carlo Tree Search (MCTS) is chosen as the basis for our framework, because of its lack of reliance on domain knowledge. The domain implementation is validated, and used to experiment on two proposed structure changes to MCTS for non-deterministic games with a high branching factor. The first technique proposed is a simplification of the Chance Node model as seen in Expectimax. The second technique outlines the introduction of move groups within the tree structure of MCTS to significantly reduce its branching factor. We conclude that both techniques are equivalent to MCTS in terms of play ing strength when a sufficient amount of simu lations can be guaranteed.
Andrea Martinenghi and othersHuman language interactions involve complex processes beyond pure information exchange, for example, actions aimed at influencing beliefs and behaviors within a communicative context. In this paper, we propose to investigate the dialogue understanding capabilities of large language models (LLMs), particularly in multi-party settings, where challenges like speaker identification and turn-taking are common. Through experiments on the game-based STAC dataset, we explore zero and few-shot learning approaches for dialogue act classification in a multi-party game setting. Our intuition is that LLMs may excel in tasks framed through examples rather than formal descriptions, influenced by a range of pragmatic features like information presentation order in prompts and others. We also explore the models’ predictive abilities regarding future dialogue acts and study integrating information on dialogue act sequences to improve predictions. Our findings suggest that ChatGPT can keep up with baseline models trained from scratch for classification of certain dialogue act types but also reveal biases and limitations associated with the approach. These insights can be valuable for the development of multi-party chatbots and we try to point out directions for future research towards nuanced understanding and adaptation in diverse conversational contexts.
Matthew Whelan and othersThis paper presents an in-depth analysis of various Reinforcement Learning (RL) models applied to the strategic board game, Settlers of Catan. The study investigates the efficacy of different Deep Reinforcement Learning architectures including Deep Q-Learning (DQN), Double DQN (DDQN), and Dueling Networks (DN) across several model sizes and with distinct reward policies. By simulating thousands of games, we generate data to train our models, focusing on optimizing their decision making capabilities in a game characterized by a large action space and stochastic elements without resource trading. Key findings suggest that the algorithmic choice significantly impacts performance, with Dueling Networks showing superior results. The study further highlights the importance of reward function design, revealing that simpler terminal-based rewards often perform as effectively as more complex schemes. Additionally, the effectiveness of Monte Carlo Tree Search (MCTS) in improving strategic decision-making in RL is demonstrated, albeit at a high computational cost. This paper contributes to the understanding of RL in board games with complex strategic interactions and provides a foundation for future explorations into advanced RL applications in similar environments.
Henry CharlesworthReinforcement Learning (RL) has produced impressive results in games like Go, DOTA 2, and StarCraft II, but these successes demand massive teams and compute resources, limiting most RL research to simple benchmarks. To explore more interesting domains with modest resources, the author previously built a successful RL agent for the four-player card game Big 2. Inspired by that experience and the 2021 UK lockdown, this project applies deep RL to Settlers of Catan—a popular 4-player board game with complex state/action spaces, hidden information, resource trading, and many decisions per game. The work details the custom simulator, key challenges (multi-agent dynamics, partial observability, trading), the PPO-based training approach, and progress achieved: agents show clear learning and promising performance, though not yet superhuman. Earlier Catan RL efforts either used hand-crafted heuristics, restricted play to 2 players without trading, or relied on rule-based agents (e.g., JSettlers) for most actions. This appears to be one of the first attempts to learn the full 4-player game with trading end-to-end from scratch, offering practical insights for applying RL to similarly rich, real-world-inspired environments.
Brahim Driss and othersCatan is a popular multiplayer board game that in volves multiple gameplay notions: stochastic ele ments related to the dice rolls as well as to the the initial placement of resources on the map and the drawing of development cards, strategic notions for the placement of the cities and the roads which call upon topological and shape recognition notions and notions of expectation of gains linked to the proba bilities of the rolls of the dice. In this paper, we de velop a policy for this game using a convolutional neural network. The used deep reinforcement learn ing algorithm is Expert Iteration (Anthony, Tian, and Barber 2017) which has already given excellent re sults for Alpha Zero and its descendants.
Bryan CollazoDuring the pandemic, I started playing online Settlers of Catan. I quickly realized there is more skill involved than one may think and this made the game beautiful for me. At the same time, I was amazed at the recent success of the AlphaGo team at making a superhuman player in the games of Chess, Shogi, and Go, with a seemingly simple algorithm. These two interests prompted me to take a shot at making a superhuman artificial intelligence player for Catan. The purpose of this post is to share these attempts so that others can take these findings further.
Bryan Collazo
Lauren NagelMarkov chains are stochastic models characterized by the probability of future states de pending solely on one’s current state. Google’s page ranking system,financial phenomen a such as stock market crashes, and algorithms to predict a company’s projected sales are a glimpse into the array of applications for Markov models. Board games such as Monopoly and Risk have also been studied under the lens of Markov decision processes. In this research, we analyzed the board game”The Settlers of Catan” using transition matrices. Transition matrices are composed of the current states which represent each row i and the proceeding states across the columns j with the entry (i,j) containing the probability the current state i will transition to the state j. Using these transition matrices, we delved into addressing the question of which starting positions are optimal. Furthermore, we worked on determining optimality in conjunction with a player’s game play strategy. After building a simulation of the game in python, we tested the results of our theoretical research against the mock run throughs to observe how well our model prevailed under the limitations of time (number of turns before winner is reached).
Gabriël van der KooijSettlers of Catan is a complex board game that makes a great Reinforcement Learning research environment because of its high strategic nature and a good mixture of competition and cooperation. In this research, Settlers of Catan was digitized using the Unity game engine and Python. Utilizing the versatility of the Actor-Critic methods, Advantage-Actor-Critic a.k.a A2C was used to try tackle a simplified version of Catan by trying to play Catan on a human level, beating Random-Agents on the way. It was found that the A2C agent learns to spend its resources and to avoid no-ops / passing. This was revealed by decreasing loss functions and inspection of individual episodes. However, because of several biases, it does not perform well enough against even the Random-Agents that just perform random legal moves. This is mostly due to the agent not learning how to prioritize specific locations for buildings and the difficulties surrounding the sparse rewards in this game. Promoting curiosity and decentralizing Actors could be utilized in the future to improve results of an RL-powered agent in Settlers of Catan.
J. DormansIn the previous chapter …
John Spencer and othersIn our project, we created a simplified version of the board game Settlers of Catan that involves only one player and then trained a AI agent to win the game as quickly as possible. We trained this AI agent using Q Learning with a epsilon-greedy exploration strategy. With a suitable number of iterations during simulation and training, the AI agent began to learn high level strategies in our complex game space and outperform a random agent.
Chris C. Kim and othersReinforcement Learning (RL), particularly Deep RL, excels at solving complex, dynamic problems with limited prior knowledge by learning optimal actions through trial-and-error rewards in Markov Decision Processes. Popularized by superhuman performance in games like Go and Poker, Deep RL combines neural networks with RL to handle large state spaces efficiently. This project develops a Deep Q-Network (DQN)-based agent to outperform humans in Settlers of Catan, a 4-player resource-trading board game distinguished by dice-driven randomness, player-to-player trading that allows direct intervention during turns, and mixed independent/competitive decision-making—features that make it significantly more complex and “loosely regulated” than traditional turn-based games like Chess or Go. The agent uses a custom Catan simulator to extract a rich feature vector (including player points, resources, buildings, dice rolls, board layout, development cards, etc.) which is fed into the DQN to select actions such as building, trading, buying cards, or passing. The work highlights Catan’s unique multi-agent trading dynamics as an exciting testbed for evaluating Deep RL in environments with heavy stochasticity and social/strategic interaction, aiming to assess how well purely learned models can compete with human players in such loosely structured settings.
Masoud Masoumi MoghadamThis three-part series by Masoud Masoumi Moghadam explores Monte Carlo Tree Search (MCTS) as a powerful, domain-independent reinforcement learning approach for building intelligent real-time game agents in complex strategic board games. Part 1 introduces reinforcement learning fundamentals (e.g., Q-learning challenges in large state spaces) and positions MCTS as an effective solution for time-constrained decision-making, balancing exploration and exploitation without requiring extensive domain knowledge. Part 2 delves into core MCTS mechanics: iterative cycles of selection (via the UCT formula: $ UCT = \frac{W_i}{N_i} + C \sqrt{\frac{\ln t}{N_i}} $), expansion, simulation (random rollouts to terminal states), and backpropagation of outcomes. It emphasizes MCTS suitability for high-branching-factor games under strict time limits, with references to a Python implementation (later optimized in Cython) and conceptual handling of challenges like memory usage, early unreliability, and stochastic uncertainty—though the example game is Hex rather than Catan. Part 3 delivers a complete, modular Python implementation of MCTS for the two-player perfect-information connection game Hex (hexagonal grid, edge-connecting objective). Key components include: a GameState class with NumPy board representation and custom Union-Find for efficient win detection; a Node class tracking visits (N) and value (Q); a UctMctsAgent orchestrating the four-phase search with time budgeting; and a tkinter-based GUI for human vs. agent play. The agent selects moves based on the most-visited child after bounded simulations, demonstrating practical real-time performance. While the series focuses on Hex (not Settlers of Catan), the author highlights MCTS’s generalizability to more complex environments, acknowledging scalability issues (e.g., convergence speed, memory) and suggesting future enhancements like RAVE or heuristics. No empirical results (win rates, baselines) are reported; the emphasis is educational—providing clear theory, code structure, and a working framework (GitHub: masouduut94/MCTS-agent-python) for learners interested in applying MCTS to strategic games, including potential adaptations for stochastic, multi-player settings like Catan.
Masoud Masoumi MoghadamThis three-part series by Masoud Masoumi Moghadam explores Monte Carlo Tree Search (MCTS) as a powerful, domain-independent reinforcement learning approach for building intelligent real-time game agents in complex strategic board games. Part 1 introduces reinforcement learning fundamentals (e.g., Q-learning challenges in large state spaces) and positions MCTS as an effective solution for time-constrained decision-making, balancing exploration and exploitation without requiring extensive domain knowledge. Part 2 delves into core MCTS mechanics: iterative cycles of selection (via the UCT formula: $ UCT = \frac{W_i}{N_i} + C \sqrt{\frac{\ln t}{N_i}} $), expansion, simulation (random rollouts to terminal states), and backpropagation of outcomes. It emphasizes MCTS suitability for high-branching-factor games under strict time limits, with references to a Python implementation (later optimized in Cython) and conceptual handling of challenges like memory usage, early unreliability, and stochastic uncertainty—though the example game is Hex rather than Catan. Part 3 delivers a complete, modular Python implementation of MCTS for the two-player perfect-information connection game Hex (hexagonal grid, edge-connecting objective). Key components include: a GameState class with NumPy board representation and custom Union-Find for efficient win detection; a Node class tracking visits (N) and value (Q); a UctMctsAgent orchestrating the four-phase search with time budgeting; and a tkinter-based GUI for human vs. agent play. The agent selects moves based on the most-visited child after bounded simulations, demonstrating practical real-time performance. While the series focuses on Hex (not Settlers of Catan), the author highlights MCTS’s generalizability to more complex environments, acknowledging scalability issues (e.g., convergence speed, memory) and suggesting future enhancements like RAVE or heuristics. No empirical results (win rates, baselines) are reported; the emphasis is educational—providing clear theory, code structure, and a working framework (GitHub: masouduut94/MCTS-agent-python) for learners interested in applying MCTS to strategic games, including potential adaptations for stochastic, multi-player settings like Catan.
Quentin Gendre and others
Martin L. Altenburg and othersDifferent AA228 2020 project from Spencer et al. 2-player simplified Catan (no P2P trade, later no dev cards); online forward search / planning vs RL. Authors report strong results vs baselines on building and resource management. Built on PyCatan.
Peter McAughan and othersOur goal for this project was to use this DQN algorithm to improve upon the JSettlers agent and show that a DQN can be successfully applied to the game mechanics of Settlers of Catan. Our original proposal had plans to implement both mechanics of the game: building and negotiating. Given some time and resource constraints, however, we decided to only implement negotiation decisions in our agent. Building an agent that utilizes DQN to play the entire game is outside the scope of this work; we are simply building upon the JSettlers agent by replacing rule-based decisions by DQN output where applicable.
Mihai Sorin DobrePhD thesis on planning plus learning from tiny human Catan corpora (~60 games vs thousands). Contributions: action-type hierarchy for MCTS rollouts, environment-model decoupling for data-efficient supervised learning, and a belief-MDP MCTS extension for partial observability (the last did not beat SOTA). Strongest agent combines decision-time planning with priors from sparse human play and defeats the then state of the art by a large margin.
Konstantia Xenou and othersIn this paper we propose a novel Deep Reinforcement Learning (DRL) algorithm that uses the concept of “action-dependent state features”, and exploits it to approximate the Q-values locally, employing a deep neural network with parallel Long Short Term Memory (LSTM) components, each one responsible for computing an action-related Q-value. As such, all computations occur simultaneously, and there is no need to employ “target” networks and experience replay, which are techniques regularly used in the DRL literature. Moreover, our algorithm does not require previous training experiences, but trains itself online during game play. We tested our approach in the Settlers Of Catan multi-player strategic board game. Our results confirm the effectiveness of our approach, since it outperforms several competitors, including the state-of-the-art jSettler heuristic algorithm devised for this particular domain.
Márton Attila BodaThe study analyses the attitude of players in a board game called Catan. In Catan, we are basically handling the players as opponents, but this does not rule out the possibility of cooperation. In a game with three players, in order to increase the chances of winning, it is worth acting together against the lead player. Cooperation has several possible modalities. In the article, the focus is on blocking situations which can lead to revenge. The primary objectives of this study were to examine how different types of thinking can cause revenge situations and which are the successful strategies among players. Strategies (as a mathematical solution to a decision problem) are examined in the study via computer modeling. To help the model, some kind of behavior of human Catan players was studied which enables profiling gaming styles used in the model. The winning chances of the player who was not involved in revenge have improved considerably, by 43%. To avoid being involved in revenge situations, the best solution is to accept other players’ opponent ranking methods.
Mihai Sorin Dobre and othersWe present a suite of techniques for extending the Partially Observable Monte Carlo Planning algorithm to handle complex multi-agent games. We design the planning algorithm to exploit the inherent structure of the game. When game rules naturally cluster the actions into sets called types, these can be leveraged to extract characteristics and high-level strategies from a sparse corpus of human play. Another key insight is to account for action legality both when extracting policies from game play and when these are used to inform the forward sampling method. We evaluate our algorithm against other baselines and versus ablated versions of itself in the well-known board game Settlers of Catan.
Gabriel Rubin and othersSettlers of Catan is one of the main representatives of modern strate gic board games and there are few autonomous agents available to play it due to its challenging features such as stochasticity, imperfect information, and 4-player structure. In this paper, we extend previous work on UCT search to develop an automated player for Settlers of Catan. Specifically, we develop a move pruning heuristic for this game and introduce the ability to trade with the other players using the UCT algorithm. We empirically compare our new player with a baseline agent for Settlers of Catan as well as the state of the art and show that our algorithm generates superior strategies while taking fewer samples of the game.
Mihai S. Dobre and othersTwo-step MCTS rollouts: sample action type then instance. Cuts bias toward high-cardinality trades, shrinks depth (random games 11,639 to 420 ply). TypedMCTS with 10k rollouts: 33.8% vs 3 STAC heuristic agents (full game with trades), 36.4% vs SmartSettlers (no trades); 54.3% vs untyped MCTS. First purely online learner covering the full legal action set. Afterstates help a lot.
Simon Keizer and othersHuman-vs-bot study on online Catan with chat trading. Compares rule-based, persuasion (Guhe and Lascarides), random-forest, and DRL (Cuayahuitl et al.) negotiators, plus original vs improved game policies. Persuasion and DRL both raise win rate vs humans over rule-based and supervised baselines; DRL also beats human players in the reported setup. Trained on STAC human-human logs.
Mihai S. Dobre and othersSupervised nets that score a variable legal-action set (no fixed output head). Mixture-of-experts by game phase, plus transfer from data-rich to data-poor phases. Trained on synthetic AI games then human Catan logs. Improves data efficiency where some phases have very few examples.
Emmanouil KaramalegosFull-rules MCTS Catan agent including inter-player negotiation. Compares VPI (value of perfect information) vs UCB and Bayesian UCB. First (per the author) complete-ruleset agent with all-player trading, plus a human-behavior-based initial placement. VPI beats bandit methods even with far fewer sims; human-like placement helps; simulation depth must be tuned. Benchmarked vs JSettlers.
Ioannis Efstathiou and othersRL trading agent in 2-player Catan that infers opponent preferences (accepted/rejected trades; also CP-nets). Tracking positive and negative prefs +8.3% trading success vs no model; CP-nets ~+8% vs no-preference RL. Best trader (NPRLA) matches hand-crafted experts in full 4p Settlers. Argues bilateral training can transfer to multilateral trading.
Heriberto Cuayáhuitl and othersArtificially intelligent agents equipped with strategic skills that can negotiate dur ing their interactions with other natural or artificial agents are still underdeveloped. This paper describes a successful application of Deep Reinforcement Learning (DRL) for training intelligent agents with strategic conversational skills, in a sit uated dialogue setting. Previous studies have modelled the behaviour of strate gic agents using supervised learning and traditional reinforcement learning tech niques, the latter using tabular representations or learning with linear function ap proximation. In this study, we apply DRL with a high-dimensional state space to the strategic board game of Settlers of Catan—where players can offer resources in exchange for others and they can also reply to offers made by other players. Our experimental results report that the DRL-based learnt policies significantly out performed several baselines including random, rule-based, and supervised-based behaviours. The DRL-based policy has a 53% win rate versus 3 automated players (‘bots’), whereas a supervised player trained on a dialogue corpus in this setting achieved only 27%, versus the same 3 bots. This result supports the claim that DRLisapromising framework for training dialogue systems, and strategic agents with negotiation abilities.
Mihai Sorin Dobre and othersWe propose a hybrid model for automatically acquiring a policy for a complex game, which combines online learning with mining knowledge from a corpus of human game play. Our hypothesis is that a player that learns its policies by combining (online) exploration with biases towards human behaviour that’s attested in a corpus of humans playing the game will outperform any agent that uses only one of the knowledge sources. During game play, the agent extracts similar moves made by players in the corpus in similar situations, and approximates their utility alongside other possible options by performing simulations from its current state. We implement and assess our model in an agent playing the complex win-lose board game Settlers of Catan, which lacks an implementation that would challenge a human expert. The results from the preliminary set of experiments illustrate the potential of such a joint model.
Takuya Hiraoka and othersIn this paper, we apply reinforcement learning (RL) to a multi-party trading scenario where the dialog system (learner) trades with one, two, or three other agents. We experiment with different RL algorithms and reward functions. The negotiation strategy of the learner is learned through simulated dialog with trader simulators. In our experiments, we evaluate how the performance of the learner varies depending on the RL algorithm used and the number of traders. Our results show that (1) even in simple multi-party trading dialog tasks, learning an effective negotiation policy is a very hard problem; and (2) the use of neural fitted Q iteration combined with an incremental reward function produces negotiation policies as effective or even better than the policies of two strong hand-crafted baselines.
Hassan Alsibyani and othersSettlers of Catan is a popular multiplayer board game that involves dice rolls, drawing cards from shuffled decks, secrecy, and agreements. This paper will first outline a fair protocol for playing Settlers of Catan with no trusted third party or shared resources, and then discuss a possible Python implementation. The protocol assumes that players are connected to each other via the internet, and that the connection is reliable. The protocol minimizes the effect of collusion between players, is able detect cheating and cheaters, and guarantees the privacy of players’ hands.
Markus Guhe and othersWe present an empirical framework for testing game strategies in The Settlers of Catan, a complex win–lose game that lacks any analytic solution. This framework provides the means to change different components of an autonomous agent’s strategy, and to test them in suitably controlled ways via performance metrics in game simulations and via comparisons of the agent’s behaviours with those exhibited in a corpus of humans playing the game. We provide changes to the game strategy that not only improve the agent’s strength, but corpus analysis shows that they also bring the agent closer to a model of human players.
Markus Guhe and othersWe present a method for obtaining a useful sym bolic model of persuasion in a complex game, where players’ preferences over the outcomes of negotiations over resources are incomplete, uncertain, and dynamic. We focus on the problem of identifying the stage in the game where successfully persuading an agent to perform a particular action will have the most impact on one’s chances to win. Our approach exploits empirical data from game simulations, where the set up allows us to investigate individual aspects of the policy in suitably controlled ways. We demonstrate the effectiveness of the approach within the domain of The Settlers of Catan and present some specific lessons that can be learned for this particular game, e.g. that a tipping point in the game occurs for persuasion moves that are made when the leading player reaches 7 Victory Points.
Konstantinos-Panagiotis PanousisMCTS (plus learning) Catan agents evaluated against each other and against existing bots including the strongest heuristic of the time. Reports MCTS can match existing autonomous Catan players despite less compute. Precursor to Karamalegos 2016 at the same lab.
Anais Cadilhac and othersPredicts which Catan trades actually execute from STAC chat. Builds evolving partial preference models (CP-nets) from dialogue, then game-theoretic equilibrium trades. Beats four baselines; both preference tracking and equilibrium reasoning matter. Does not play the full game, but is core ML for Catan negotiation.
Markus GuheWe collected and analysed a small dialogue corpus of people playing The Settlers of Catan. Dialogues are trading negotia tions where Gricean maxims of cooperativity often break down as players adopt conflicting intentions in their attempt to win the game. This has consequences for what information players are sharing and for the sincerity of their contributions. In this paper, we motivate and describe a two-level scheme for analysing non-cooperative dialogues, where both levels are in terdependent. Each dialogue move is a move in the game (e.g., an offer to trade), and a coherent contribution to the dialogue so far, connected to a prior segment with a coherence relation, such as indirect answerhood or rejection. Parsing and gen erating coherence relations is computationally feasible (e.g., Baldridge & Lascarides, 2005), and here we’ll argue that their semantics help to identify the game move, even when it is im plicated rather than linguistically explicit
István Szita and othersGames are considered important benchmark tasks of artifi cial intelligence research. Modern strategic board games can typically be played by three or more people, which makes them suitable test beds for investigating multi-player strategic decision making. Monte-Carlo Tree Search (MCTS) is a recently published family of algorithms that achieved successful results with classical, two-player, perfect-information games such as Go. In this paper we apply MCTS to the multi-player, non-deterministic board game Settlers of Catan. We implemented an agent that is able to play against computer-controlled and human players. We show that MCTS can be adapted successfully to multi-agent environments, and present two approaches of providing the agent with a limited amount of domain knowledge. Our results show that the agent has considerable playing strength when compared to existing heuristics for the game. We conclude that MCTS is suitable for implementing a strong Settlers of Catan player.
Jeroen Geuze and othersAn Intelligent Tutoring Agent (ITA) for the board game Settlers of Catan (SoC) is introduced. It uses CLIPS knowledge bases, connected by JCLIPS to a JAVA implementation of SoC. It is founded on a new theoretical framework that describes the development of negotiation skills in children. Using this framework, the ITA helps children in developing negotiation skills through play, which makes it unique in its kind.
Robert Shaun ThomasDissertation that introduced JSettlers and the classic heuristic/plan-based Catan robot (decision tree over time-to-10-VP). Java Settlers is the research environment still used by STAC, Szita, Dobre, and others. The related 2002 IUI paper is a 2-page demo of the multi-agent negotiation testbed.
Michael PfeifferIn this paper we study the application of machine learning methods in complex computer games. A combination of hierarchical reinforcement learning and simple heuristics is used to learn strategies for the game Settlers of Catan (© 1995 by Kosmos Verlag, Stuttgart) via self-play. Since existing algorithms for function approximation are not well-suited for problems of this size and complexity, we present a novel use of model trees for state-action value prediction in a sophisticated computer game. Furthermore we demonstrate how a priori knowledge about the game can reduce the learning time and improve the performance of learning virtual agents. We compare several different learning approaches, and it turns out that, despite the simplicity of the architecture, a combination of learning and built in knowledge yields strategies that are able to challenge and even beat human players in a complex game like this.
Luca Branca and othersSettlers of Catan is a board game where the main goal is to collect victory points by building a society of set tlements, cities and connecting roads. We have constructed a multi-agent system solution able to play the game and evalu ated it against other available bots for the game. Our initial re sults show that even if the proposed solution does not beat the best monolithic solutions, its strategic game play is promising enough to encourage further research in the area.
Scroll naar boven