Chess, Go and poker have all served as important testing grounds for artificial intelligence. But poker differs from the other two in one fundamental way: it is a game of imperfect information.
In chess and Go, the complete state of the game is visible. In poker, a player — and likewise a poker AI — knows its own cards, the community cards, pot size, stack sizes and previous actions, but cannot see the opponents' private cards.
Even an opponent's actions do not reveal that hidden information with certainty. A large bet may represent a strong hand or a bluff. A check can indicate weakness, but it can also conceal a monster. Different hands can take the same action, while the same hand does not necessarily have to be played the same way every time.
A poker AI therefore cannot simply try to "guess" what cards an opponent is holding. It must reason about distributions of possible hands and strategies, and make decisions based on those probabilities and potential future outcomes.
That is precisely what made poker such a valuable challenge for AI research.
Why is poker a different AI problem from chess?
In chess, the entire position is known at any given moment. One of the main difficulties comes from the enormous number of possible future sequences that can develop from that position.
Poker adds another problem: part of the current game state itself is hidden.
On the river, for example, an AI may know the exact board, pot size, effective stacks and complete betting history while still having no direct knowledge of its opponent's two hole cards. Instead, it must reason about a range of possible hands and the probability associated with each of them.
Then there is mixed strategy.
An optimal poker strategy does not necessarily say that the same hand should always take the same action. Betting may be correct at one frequency and checking at another.
Bluffing is part of the same strategic system. If a player only bets strong hands, their actions reveal too much information and the strategy becomes exploitable.
Poker therefore confronts AI with several problems simultaneously: hidden information, an enormous decision space and strategic opponents actively trying to exploit its decisions.
Poker attracted AI researchers long before the online poker boom
The relationship between poker and artificial intelligence predates the explosion of online poker.
Researchers were interested in the game because it combined probability, game theory, hidden information and strategic interaction within a formal system with clearly defined rules.
Growing computing power and the ability to simulate poker digitally dramatically expanded what researchers could attempt.
In 2006, the Annual Computer Poker Competition was launched, allowing poker programs developed by universities and research groups to compete against each other. The University of Alberta and Carnegie Mellon University became particularly important centers of poker AI research.
The real question went far beyond whether researchers could build a strong poker bot.
It was: How can an artificial intelligence develop an effective strategy in an environment where it does not have complete information?
2015: A form of Hold'em was essentially solved
One of the biggest breakthroughs arrived in 2015.
Michael Bowling and researchers at the University of Alberta announced that their Cepheus system had essentially weakly solved heads-up limit Hold'em.
A key part of the research involved Counterfactual Regret Minimization and its improved CFR+ variant.
In simplified terms, the system repeatedly evaluates how much better it could have performed by choosing alternative actions. By minimizing this accumulated "regret" over enormous numbers of iterations, the algorithm moves toward an equilibrium strategy.
The result is not necessarily a rule such as "always raise in this situation."
Instead, a strategy might specify raising with a particular hand at one frequency and calling at another.
That approach will sound familiar to anyone who has studied modern GTO poker. From an AI perspective, however, the breakthrough demonstrated something much broader: extremely large imperfect-information games could be approached with strategies that were extraordinarily difficult for an opponent to exploit.
2017: Libratus defeats professional poker players
The next spectacular milestone came from Carnegie Mellon University's Libratus.
Created by Tuomas Sandholm and Noam Brown, Libratus tackled heads-up no-limit Hold'em — a vastly more complex game than heads-up limit poker.
In January 2017, it played 120,000 hands against four strong heads-up professionals: Jason Les, Dong Kim, Daniel McAulay and Jimmy Chou.
Libratus won decisively.
But its importance was not simply that a computer had beaten professional poker players.
Libratus was not primarily trying to imitate how elite human players approached the game by learning from human hand histories. Starting from the rules of poker, it computed its own strategy.
It was not learning how humans played poker.
It was trying to determine how the game itself should be played.
DeepStack: What happens when you cannot calculate everything?
Another major breakthrough arrived around the same time with DeepStack, which also demonstrated superhuman performance in heads-up no-limit Hold'em.
DeepStack combined deep learning, game-theoretic reasoning and real-time search.
The complete game tree of no-limit Hold'em is far too large to calculate exhaustively. Instead, DeepStack used neural networks to estimate the value of future game states without having to calculate every possible continuation all the way to the end of the hand.
That addresses a much broader problem in artificial intelligence:
If you cannot calculate every possible future, can a system accurately estimate the value of states it has not fully explored?
The question extends far beyond poker.
2019: Pluribus takes on six-player poker
Heads-up poker is a strategic contest between two participants. Six-player poker introduces an entirely different level of complexity.
In 2019, Noam Brown and Tuomas Sandholm unveiled Pluribus, which demonstrated superhuman performance in six-player no-limit Hold'em against elite human players.
The significance again extended beyond poker.
Many real-world strategic problems involve multiple decision-makers. Auctions, negotiations, markets and economic competition can all involve outcomes that depend on the actions of several independent participants.
Poker therefore became an important experimental environment for multi-agent decision-making.
What does any of this have to do with ChatGPT?
This is where an important distinction is necessary.
Large language models such as ChatGPT are not direct technological descendants of poker AIs such as Cepheus, Libratus or Pluribus.
Transformers, large-scale language pretraining and next-token prediction emerged from a different research lineage.
The connection becomes more interesting in areas such as reinforcement learning, self-play, search and reasoning.
There is also a very concrete link between the two eras of AI research.
Noam Brown, one of the researchers behind Libratus and Pluribus, later joined OpenAI, where his work shifted toward improving the reasoning capabilities of large language models. He also played an important role in the research behind OpenAI's o1 reasoning model.
That does not mean that Libratus' poker algorithms were simply transferred into a language model.
But there is continuity in some of the underlying research questions.
From the poker table to reasoning models
One major area of modern AI research is inference-time, or test-time, compute.
The basic idea is that computing power does not only have to be spent during the training of a model. A system can also use additional computation while working on a particular difficult problem.
Game-playing AI had already demonstrated the value of a related idea.
Instead of calculating and storing a complete solution for every conceivable situation in advance, an AI can perform additional search and computation when it encounters a particular decision.
Modern reasoning models operate using very different technology, but the conceptual parallel is clear:
How can an AI use more computation on a specific difficult problem to produce a better solution?
That question connects some of the research lessons from game-playing AI with one of the most important areas of modern AI development.
Why was poker such a useful AI laboratory?
Poker offers something close to laboratory conditions for studying difficult decision-making.
Its rules, actions and payouts are precisely defined, yet the game simultaneously contains private information, randomness, bluffing, mixed strategies, an enormous decision space and intelligent opponents.
Digital poker provides another major advantage: the game state can be represented as structured data — hole cards, board cards, position, pot size, stack sizes, bet sizes and action history.
Computers can also simulate enormous numbers of hands against each other. With self-play, they do not even need human opponents to generate experience.
Modern artificial intelligence did not "grow out of" commercial online poker.
Rather, computer-simulated poker became one of AI research's most useful laboratories for studying strategic decision-making under uncertainty.
Poker's real legacy in artificial intelligence
In chess, a machine must find the right move from a fully observable position within an enormous search space.
Poker gave researchers a different challenge:
How should a machine make decisions when part of the game state is hidden, the same opponent action can represent many different hands, and the optimal strategy itself may require bluffing and randomized actions?
That is a much broader question than how to play a particular river spot.
Poker became one of the major experimental environments for studying how machines can make rational decisions in partially observable, adversarial and strategic settings.
In chess, the machine could see everything.
At the poker table, it had to learn how to reason about what it could not see.
And decades later, some of the researchers and research questions that emerged from that challenge have found new relevance in the development of today's reasoning-focused AI systems.

















0 comments