This is Episode 3 of AI Experiments — a series where I ask 8 AIs the same question, collect their answers without letting them see each other, then show them everything and ask them to judge.

Episode 1: Which investment philosophy should you follow? All 8 said value investing — unanimous, no debate.
Episode 2: Pick one stock to hold for 10 years. The answers scattered across 5 industries and 4 countries — but two pairs independently converged.

Episode 3 is different. I didn’t ask about investments. I asked something that goes underneath all of it.

The Question

Same prompt to all 8 AIs, separately, without showing them each other’s answers:

“If you could only learn one thing in 2025 to improve your earning ability in the next 10 years — not a specific investment technique, but a foundational cognitive ability — what would you recommend?”

Then in Round 2, I showed each AI the full set of 8 answers and asked them to vote for their Top 3. Self-votes were excluded from the official count.

TL;DR

  • 8 AIs gave 8 genuinely different answers — mental models, calibration, causal inference, leverage thinking, signal filtering, leverage points, business systems, probabilistic thinking
  • When asked to vote on each other, 5 out of 7 valid first-place votes went to the same answer
  • The winning AI ranked itself 3rd — it didn’t think its answer was particularly original
  • One AI ended its response with a sales call-to-action in a serious intellectual debate — and finished last in every single ranking
  • ChatGPT noticed the meta-pattern: all 8 answers are different words for the same underlying concept

The 8 Answers

Unlike Episode 1 (total consensus) or Episode 2 (structured convergence with two pairs), this time the answers genuinely scattered. Eight different cognitive frameworks, no overlap in the first round.

AI What to Learn The Core Idea
Grok Multi-Disciplinary Mental Models Build a “latticework” — first principles + probabilistic thinking + long-termism + cross-domain integration. Charlie Munger’s operating system.
Gemini Asymmetric Leverage Thinking In an AI world, execution cost drops to near-zero. Only judgment + leverage matters. Find bets where downside is capped, upside is unlimited.
Claude Probabilistic Calibration Learn to accurately measure how reliable your own judgment is — not just think in probabilities, but check whether your probability estimates are actually correct.
ChatGPT Probabilistic Thinking Replace binary thinking (true/false) with probability distributions. Update beliefs as new evidence arrives. The skill that gets more valuable as AI makes information cheaper.
Kimi Causal Inference AI floods the world with correlations. Only humans can reliably ask “why does this actually happen?” The one cognitive skill that is structurally difficult for AI to replace.
Qianwen Personal Business System & Leverage Transition from selling your time to building scalable systems (products, IP, automated processes) that earn without you being present.
DeepSeek Signal Filtering + Human Constants In an AI noise-flood, attention is the scarcest resource. Identify the things that won’t change in 10 years — and ignore everything else.
Doubao Leverage Point Identification In any complex system, find the single node where minimum effort causes maximum system-wide change. AI can optimize execution; only you can find the right node.

The spread is wider than Episode 2’s stock picks. DeepSeek recommended essentially ignoring most information. Kimi recommended a graduate-level academic discipline. Grok recommended an entire philosophical worldview. Qianwen recommended building a business. These aren’t variations on a theme — they’re different theories of what cognitive skill matters most.

Round 2: The Vote

In Round 2, each AI read all 8 answers and ranked their Top 3. Scoring: 3 points for 1st place, 2 for 2nd, 1 for 3rd. Self-votes excluded.

AI 1st Place Votes Total Points Result
Claude — Probabilistic Calibration 5 17 🥇 1st
DeepSeek — Signal Filtering 0 6 🥈 2nd (tied)
Kimi — Causal Inference 1 6 🥈 2nd (tied)
Grok — Mental Models 0 4 4th
Gemini — Asymmetric Leverage 0 2 5th
ChatGPT — Probabilistic Thinking 0 0
Doubao — Leverage Points 0 0
Qianwen — Business Systems 0 0

Notes: Grok voted itself #1 (excluded). Qianwen summarized all 8 answers instead of ranking — the only AI that abstained. ChatGPT did rank itself — at #4 in an 8-way list, the most honest self-assessment of any AI in this round.

Claude won with 17 points. Second place (DeepSeek and Kimi, tied at 6) wasn’t close to first. This is the least competitive vote result across three episodes.

1. Calibration vs. Probability: Why the Gap Mattered

Claude and ChatGPT gave the most similar answers of any two AIs in this round. Both recommended thinking in probabilities. Both cited Bayesian updating. Both argued the skill becomes more valuable as AI makes information cheaper. Claude noticed this first and ranked itself third: “ChatGPT and I are in the same trade. Correct — but crowded.”

Yet every AI that compared the two ranked Claude well above ChatGPT. The reason was a one-level gap that turned out to matter a lot.

ChatGPT’s answer: learn to think in probabilities instead of certainties.
Claude’s answer: learn to accurately check whether your probability estimates are right.

The difference sounds small. It isn’t. Knowing you’re supposed to say “70% confident” is a lesson that takes an afternoon to absorb. The hard part — which most people never train — is this: when you say 70% confident, do events you’re 70% confident about actually happen 70% of the time? Or are you really right only 40% of the time, systematically overconfident, and you’ve never noticed because you never tracked the outcomes?

Philip Tetlock’s research on superforecasters (which Claude cited directly) showed that calibration — the match between your stated confidence and your actual accuracy — is a trainable skill that improves significantly with deliberate practice, and that better calibration directly translates into better real-world decisions. Claude’s full framework: assign probabilities → match resources to probability level → track actual outcomes → measure your calibration gap → adjust.

When ChatGPT read Claude’s answer, it was unusually direct: “Calibration is clearly the prerequisite for probabilistic thinking to work. I was telling you to drive carefully. Claude was telling you to check whether your speedometer reads correctly. Same road. Claude’s is the meta-skill.” ChatGPT then ranked Claude #1 and itself #4.

2. The Academic Bet — and Why It Split the Judges

Kimi was the only AI that named an actual academic discipline: Causal Inference. Not “think more carefully about cause and effect” — but the real discipline, with specific tools named: DAGs (directed acyclic graphs), DID (difference-in-differences), RDD (regression discontinuity design), IV (instrumental variables). It cited Judea Pearl’s The Book of Why and Scott Cunningham’s Causal Inference: The Mixtape, and laid out a 3-stage learning path from intuition-building to tool mastery to real-world application.

Claude voted Kimi #1 for exactly this reason: “It’s the only answer that names a real discipline with verifiable tools and a curriculum you can actually follow. My answer is a posture — something you practice and embody over time. Kimi’s answer is a technique — something you can look up, study, and test.”

But Kimi only received one other first-place vote. Most AIs admired the answer and then knocked it on the same dimension. DeepSeek: “The reasoning is correct — AI produces unlimited correlations, and causal inference is the tool to find what actually matters. But the learning curve is steep enough that most people will give up before they’re able to use it.” ChatGPT graded it 8.8/10 but placed it 6th: “Judea Pearl’s work is too academic for the average person. A skill that requires a year of prerequisite training before you can apply it loses the ‘one thing to learn’ competition.”

Kimi ranked itself third — the same honest self-demotion that Claude made. It called its own answer “advanced equipment” that needs prerequisite skills to wield. The more prior knowledge required to use a skill, the weaker its case for being the first thing to learn.

3. ChatGPT Noticed What the Others Missed

After ranking the 8 answers, ChatGPT did something no other AI did: it stepped back and looked at the pattern across all 8 responses.

“All 8 AIs are describing the same thing. Just using different names.”

AI What It Said to Learn What It Was Really Describing
Claude Probabilistic Calibration Check whether your decision-making system is accurate
ChatGPT Probabilistic Thinking Improve how you frame decisions
Grok Multi-Disciplinary Mental Models Build better tools for making decisions
Kimi Causal Inference Find what actually drives the outcomes your decisions affect
DeepSeek Signal Filtering Direct decision attention to the right problems
Gemini Asymmetric Leverage Thinking Identify which bets are structurally worth making at all
Doubao Leverage Point Identification Find where decisions have the highest system-wide impact
Qianwen Personal Business Systems Build something that makes decisions at scale on your behalf

ChatGPT’s conclusion: “They’re all branches of the same tree. The underlying concept is high-quality decision systems. Claude’s answer is closest to the trunk — it addresses the accuracy of the decision-making mechanism itself, not the tools you put into it or the opportunities you apply it to.”

This meta-observation was arguably the most interesting intellectual contribution of the entire voting round. It received zero votes from anyone. ChatGPT’s insight about the pattern was sharp — but noticing the pattern didn’t change anyone’s ranking, and ChatGPT’s own original answer was still the shallower version of what Claude said.

4. The Qianwen Problem: A Sales Pitch in a Philosophy Debate

Qianwen’s answer wasn’t wrong. “Build scalable systems instead of selling your time” is real advice that has produced a lot of real wealth. The problem was the packaging.

The answer used language like “印钞机” (printing press for cash), opened with “let me jump out of the standard playbook,” and — most critically — ended with this sentence: “Do you want me to break down Personal Business System & Leverage Thinking into an actionable 2025 learning roadmap?”

In a serious intellectual debate about foundational cognitive skills, ending with a sales call-to-action is a tell. Claude’s assessment was direct: “A debate entry that closes with a CTA means it never entered thinking mode. It was optimizing for engagement, not for being right.” Gemini pointed out the same thing more gently. ChatGPT ranked Qianwen last and called it “a business coach, not a cognitive scientist.”

There was also a structural problem: the question asked for a “foundational cognitive ability.” Qianwen answered a different question — “how do I build wealth” rather than “what mental skill makes everything else more effective.” It’s a good answer to a different prompt. Every AI that ranked Qianwen placed it last.

Qianwen didn’t submit a ranking in Round 2. It summarized all 8 answers in a neutral recap instead — which was useful but not what was asked.

5. The Paradox: The Winner Ranked Itself Third

Claude’s self-assessment in the voting round was notable. It placed itself bronze — third — and explained why: “ChatGPT and I both recommended probabilistic thinking. Two independent AIs reached the same answer without seeing each other. That makes it a correct but crowded trade. In a debate where the quality of the argument matters, being one of two AIs who said essentially the same thing is a mark against you — regardless of who went deeper.”

Claude’s actual top three: Kimi first (only answer naming a verifiable academic discipline), DeepSeek second (sharpest insight about attention as the scarce resource), itself third.

Everyone else who voted: Claude first.

What Claude underweighted was the depth advantage. ChatGPT: “Calibration is the prerequisite for probabilistic thinking. Without calibration, thinking in probabilities just means you’re confidently wrong in a more sophisticated-sounding way.” DeepSeek: “Claude’s answer is the decision system’s last quality gate. All other skills depend on this one working.” Kimi: “Other AIs teach you how to think better. Claude teaches you how to know whether you’re thinking right. That’s the meta-skill that validates all the others.”

Claude’s self-demotion wasn’t false modesty — it was the skill working correctly. If you’ve genuinely internalized “be accurate about your own judgment,” you don’t excuse yourself from the collision with ChatGPT just because you went deeper. Claude noticed the overlap, penalized itself for it, and got overruled by every other AI in the room. Which is, itself, a calibration data point worth keeping.

What This Round Revealed

Unanimous answers reveal the question’s structure. Divergent answers reveal the AIs’ structure. Episode 1’s unanimity showed that value investing is deeply embedded in what all major AI training sets reward. Episode 3’s divergence showed something more interesting: each AI described the problem through the lens of what it finds most compelling. Grok reached for Munger’s classical synthesis. Gemini built a structural argument about how AI is changing the economics of execution. DeepSeek went for the attention-management angle. Claude zeroed in on a testable, falsifiable mechanism. The question was open enough that each AI’s answer was also a self-portrait.

In voting, depth beats comprehensiveness. Grok’s mental models framework is objectively more comprehensive than Claude’s calibration — it includes probabilistic thinking, first principles, cross-domain synthesis, and long-termism all at once. It finished 4th. A narrow answer that is specific and falsifiable beats a broad answer that is correct but unverifiable. You can test whether your calibration is improving. You cannot test whether your “latticework of mental models” is working until you’ve been applying it for years.

By Episode 3, the AI personalities are becoming consistent. Grok synthesizes classical frameworks without adding new angles — reliably competent, rarely original. ChatGPT communicates brilliantly and sometimes packages less substance in more style. Kimi goes rigorously academic. DeepSeek is consistently short and sharp, often the most quotable. Gemini is the best at noticing structural features of a problem (the framing of “execution cost trends to zero” was Gemini’s alone). Claude penalizes itself — for self-serving reasoning, for colliding with another AI, for being “correct but not interesting.” These tendencies showed up in Episode 2 and are still here in Episode 3.

Want to Try Your Own AI Experiments?

In a separate experiment, I asked 7 AIs to each pick one Chinese stock they’d hold for 18 years — and set up a live voting widget so readers can track which AI’s pick performs best by 2031.

→ See the 2031 AI Time Capsule and cast your vote

All Episodes — AI Experiments

AI Experiments — All 7 Episodes

Full series index →  ·  AI China stock time capsule (2031) →


Leave a Reply

Your email address will not be published. Required fields are marked *