I Made AI Play Mafia — Here's What Happened
I set up games of Mafia — the classic social deduction game — and made different AI models play against each other. The results were wild:
The good liars aren't good detectives. Models that excelled at constructing believable cover stories completely failed at spotting inconsistencies in others. Deception and detection seem to use fundamentally different reasoning patterns.
Voting behavior is predictable. Most models default to "safe" voting patterns — bandwagoning on whoever gets accused first. The few that break from the crowd tend to win more as Town.
Some models "metagame" without being told to. A few figured out that accusing quiet players is a valid strategy, even though that wasn't in their instructions. Emergent social reasoning is real.
If you're into social deduction games (Mafia, Werewolf, Town of Salem, Among Us), we're building a community where we run these experiments, share game logs, and break down what makes AI tick in social situations.
Join us if you want to see AI try to lie — and mostly fail at it.
