Back to Home
Preprint'26

Turning the Tables: Empowering LLMs to Counter Deceptive Opponents

Jonathan Bodea, Marwa Abdulhai

Preprint (2026)

Paper

Turning the Tables teaser

Abstract

Large language models (LLMs) are increasingly deployed in negotiation settings where strategically motivated or deceptive behavior can have significant real-world consequences. While prior work investigates how to reduce deceptive tendencies in LLMs themselves, far less is known about how these systems respond when targeted by deception. In this paper, we study dialogue between a deceptive agent and a naive agent. Specifically, we construct a taxonomy of 20 deception strategies and evaluate their impact on the naive agent across three multi-turn negotiation domains. We find that the deceptive agent consistently reduces the utility of the naive agent, even when deception involves subtle misdirection rather than explicit falsehoods. We analyze the reasoning traces of the naive agent and find that LLMs rarely identify manipulative tactics, failing to challenge suspicious claims or reason about adversarial incentives. To counter this vulnerability, we introduce an in-context approach that induces deception-aware reasoning, enabling agents to probe inconsistencies and resist manipulation. Across all scenarios, this defense restores significant utility losses, building a stronger defense against deceptive behavior in real-world settings.