I’ve been running an experiment: asking 8 AI chatbots the same question, separately, then comparing their answers. Each episode reveals something different about how these systems think — and where they agree, diverge, or surprise you.
The 8 AIs: Claude, ChatGPT, Gemini, Grok, DeepSeek, Kimi, Qianwen (Tongyi), and Doubao. Same question every time. No coordination.
All 7 Episodes
Patterns Across the Series
Convergence on structured questions. Episodes 1, 6, and 7 — all designed to produce divergence — produced unanimous answers instead. AI systems share enough common training to arrive at the same well-reasoned position when a dominant argument exists.
Divergence on open-ended questions. Episodes 2 and 3 had genuinely scattered answers — the “which stock” and “which skill” questions had no dominant answer, and each AI’s choice revealed its individual character.
The AI personalities are consistent across all 7 episodes. Grok synthesizes classical frameworks. ChatGPT finds the most elegant framing. DeepSeek is the sharpest and most quotable. Gemini finds structural features others miss. Kimi is the most rigorous and table-heavy. Claude is the most self-critical — in Episode 7, it explicitly argued that AI answers should carry less weight because AI bears no consequences for being wrong.
Related
→ 7 AIs each picked China’s next long-term stock winner — time capsule to open in 2031
Support This Work
These AI experiments take real time and real API costs. If you found this useful, a coffee helps keep them going.
Leave a Reply