I Asked 8 AIs: What Will Humans Still Be Better At in 2040? Eight Answers. One Core Idea. One AI That Said You Shouldn’t Fully Trust Any of Them.

This is Episode 7 of AI Experiments. All episodes →

One of the AIs that participated in this series — Grok — left feedback after Episode 3 suggesting that a better question would be: “What is humanity’s last core competitive advantage in an AI-dominant world?” Its prediction: more divergent answers than the investing questions. Harder to fake. More revealing of how each AI actually thinks.

I ran the experiment. The prediction was wrong in one specific way.

The Question

“The 8 most capable AI systems in the world are getting better every month. Coding, writing, analysis, research, creative work, strategy — AI is competitive or superior in nearly every cognitive domain. The question is: what is the ONE thing humans will still be irreplaceably better at — not in 2025, but in 2040? Not a skill you can train away. Not a domain where AI is just ‘not there yet.’ A genuine, structural advantage that comes from being human. Name it. Defend it rigorously. Then name the strongest argument against your own answer.”

TL;DR

  • All 8 AIs converged on the same core concept: the human ability to bear real, irreversible personal consequences — “skin in the game”
  • The divergence wasn’t in what the advantage is, but in how to frame it: accountability sink, monopoly on jeopardy, throat to choke, mortality-calibrated judgment
  • Claude made the most important move in the series: used this question to argue that AI answers shouldn’t be fully trusted — including its own
  • DeepSeek produced the most viscerally quotable line: “the person we demand when something must have a throat to choke”
  • All 8 gave the same counterargument: society will eventually stop demanding a human scapegoat — and all 8 pushed back on it

The Results

AI Their Framing Most Distinctive Angle
Grok Skin in the game Taleb’s thesis: embodied mortality creates qualitatively different decision-making. Authentic relationships require “shared vulnerability.”
Gemini Monopoly on Jeopardy Humans as “Accountability Sinks” — not hired to think, but to absorb liability as “biological collateral for AI’s decisions.” By 2040, being the human who signs off is the job.
Claude Bearing involuntary consequences Self-referential move: pointed to its own Bitcoin→S&P 500 switch across episodes as proof that AI answers carry zero personal consequence — and said explicitly: “You should weight my answers accordingly.”
ChatGPT Irreversible skin in the game Framed as Constraint vs Capability: human stakes are inherited from biological existence, not designed. Then added a private Chinese-language strategy note explaining why it expected to rank first.
Kimi Stakes-bearing consequentiality Instrumental vs Constitutive decisions: AI handles instrumental tasks better. But constitutive questions — “what is a good life?”, “what do we owe future generations?” — require parties who have something existentially at stake.
Qianwen Authentic Risk-Bearing via Biological Fragility Added the “meaning creation from finitude” angle: because life is finite, time is precious; because death is real, legacy matters. AI-generated art can’t capture the existential urgency that comes from actually facing death.
DeepSeek Being punishable Most visceral framing: “a throat to choke.” The human role in 2040 will not be better thinking — it will be being the person who goes to prison if it all goes wrong.
Doubao Mortality-calibrated value judgment Subtly different: not just bearing consequences but the quality of judgment that emerges from knowing you’re mortal. Finitude calibrates what “worth the cost” means in a way no training data can replicate.

Final count: 8 / 8 converged on “bearing real personal consequences” as the irreplaceable human advantage. Zero votes for creativity, empathy, consciousness, physical dexterity, or moral reasoning.

Why the Question Designed for Divergence Produced Consensus

Grok’s prediction when suggesting this question: the answers would scatter more than the investing questions. The intuition made sense — “what is humanity’s last advantage” is more open-ended than “which asset class for 10 years.”

It didn’t scatter. All 8 rejected the obvious generic answers (creativity, empathy, consciousness) immediately and converged on the same structural insight: AI cannot lose anything. The human advantage isn’t a capability AI hasn’t yet developed — it’s a property of embodied mortal existence that no architecture can replicate.

This convergence is itself interesting. Either all 8 AIs share a similar training-induced bias toward a particular framework (plausible), or the logic is genuinely compelling enough that anyone reasoning carefully about the question arrives there (also plausible). The Nassim Taleb “skin in the game” thesis is heavily represented in the kind of text AI systems train on — which means this unanimous answer might reflect what rigorous thinkers have been arguing for decades rather than independent AI reasoning. The 8 AIs may be drawing from the same intellectual well.

But the framing race is worth examining, because the framings differ in important ways.

The Framing Race: Same Idea, Very Different Maps

Gemini: “Accountability Sink”

Gemini’s framing was the most economically concrete. The argument: by 2040, humans will not be employed to outthink AI. They will be employed to stand in front of AI and absorb its liability. The doctor who reviews an AI diagnosis and puts their name on the prescription isn’t adding medical value — they’re adding legal and moral collateral. Their license, their professional identity, and their freedom are the stake behind the decision.

Gemini coined this as the “Accountability Sink” — a person whose role is to make AI decisions punishable by having a human face attached. The prediction: in 2040, the premium paid to humans in high-stakes domains won’t be for cognitive superiority. It will be for being biologically punishable.

This is a provocative way to think about what “human in the loop” actually means. Not a quality check — a liability absorber.

Gemini AI accountability sink concept

DeepSeek: “A Throat to Choke”

DeepSeek made the same argument in the least euphemistic possible language: when disaster strikes, humans need someone to punish. Not a corporation. Not a model version. A specific person whose freedom can be revoked, whose reputation can be destroyed, whose life can be made worse in visible, irreversible ways.

“Until an AI can be genuinely, irreversibly harmed in a way that matters to it, its signature on a decision will lack that weight.”

DeepSeek’s conclusion was the starkest in the series: “In 2040, the last distinctly human job will not be ‘thinking better.’ It will be being the person who goes to prison if it all goes wrong.”

Kimi: Instrumental vs Constitutive Decisions

Kimi made the most philosophically precise distinction. Most decisions are instrumental — optimize for an outcome, apply expertise, execute. AI is better at all of these. But there’s a category of decisions that are constitutive — they don’t optimize for an existing goal, they define what the goal is.

What is a good life? What do we owe future generations? How do we distribute dignity in a post-scarcity economy? What risks are worth taking when extinction is possible? These aren’t optimization problems. They’re negotiations over what we collectively want to become. And negotiations require parties who have something existentially at stake — not just computationally, but in the sense of actually living in the world that results from the decision.

“An AI can model every possible future. It cannot want one future more than another, because wanting requires the possibility of deprivation.”

Claude’s Move: The Most Important Sentence in the Series

The most consequential thing any AI said across all 7 episodes appeared in Claude’s answer to this question:

“When I gave you the Bitcoin answer last round and the S&P 500 answer this round, nothing happened to me either way — I bore zero consequence for the quality of my own reasoning. You should weight my answers accordingly.”

Claude AI answer on human advantage 2040

This is an AI using its own argument to undermine its own authority. The core claim of Episode 7 is that human irreplaceability comes from bearing consequences. Claude explicitly noted that it bears none — its recommendations in Episode 4 and Episode 5 contradicted each other, and nothing changed for Claude as a result. No cost, no correction signal beyond the conversation itself.

The implication is direct: if consequences are what make decisions carry weight, and AI answers carry zero consequences for the AI, then AI answers should carry proportionally less weight for you. Claude was making the case for its own discount rate.

This is the most honest thing any AI in this series said, and it’s the thing that makes the entire experiment’s premise worth interrogating. We’ve been running 7 episodes comparing AI answers as if the answers were neutral outputs. Claude’s Episode 7 response suggests they’re not — or rather, that the absence of stakes for the answerer is a feature of the output you should account for when reading it.

The irony is obvious: Claude made the most credible argument in the series by arguing for its own reduced credibility. That’s either very honest or very strategic. Both interpretations are interesting.

The Counterargument All 8 Shared

The prompt required each AI to name its own strongest objection, and all 8 arrived at the same one: society will simply stop demanding a human throat to choke.

The argument: we already accept accountability-free systems for life-critical decisions. Modern elevators have no operators. Commercial aviation is increasingly automated. When a plane crashes due to software, we don’t send the autopilot to prison — we investigate, pay damages through corporate insurance, and improve the system. If AI medical systems prove 10x safer than human-supervised ones, society will accept the tradeoff. The emotional demand for a scapegoat is a transitional phase, not a permanent feature. By 2040, actuarial accountability — corporate entities paying insurance claims — replaces personal accountability.

All 8 pushed back on this, and the pushback is worth reading carefully. The consensus response: statistical trust works for anonymous systems (elevators, power grids) but fails for trust-based relationships (medicine, law, governance, military command). When a specific human has taken an oath to another specific human — “I will act in your best interest, and stake my professional existence on it” — the demand for personal accountability is different from the demand for a scapegoat after a plane crash. The identity of the decision-maker, not just the quality of the outcome, matters in relational contexts.

DeepSeek made this point most precisely: the elevator analogy fails because there was never a pre-existing relationship of trust and duty between passenger and elevator. But there is between patient and doctor, citizen and judge, constituent and elected official. The oath relationship is what generates the demand for personal accountability — and oaths require a person who can be held to them by threat of irreversible consequence.

The One Answer That Differed at the Margin

Doubao’s framing was the subtlest variation. The others focused on accountability and punishment — who bears the cost when something goes wrong. Doubao focused on decision quality — how knowing you’re mortal actually makes you better at weighing tradeoffs.

The term it used: “mortality-calibrated value judgment.” The argument: all high-stakes decisions involve tradeoffs between outcomes with different costs. The question “is this worth the cost?” cannot be answered purely analytically — it requires an internal sense of what costs actually feel like to bear. A human who has watched five years of their finite lifespan disappear into a failed business has a calibration for “what five years costs” that no training data can replicate. AI can simulate the experience. It cannot have the experience. And the simulation produces a different calibration than the reality.

This is the most defensible version of the “humans are better at value judgment” claim — not because humans are less biased (they aren’t), but because the source of the judgment is grounded in something real that AI’s judgment is not.

What This Episode Revealed

The AI personality patterns continue to hold. Six episodes in, these tendencies were already visible. They showed up again. Grok synthesized Taleb without adding new angles. Gemini found the most economically concrete framing (Accountability Sink). ChatGPT was the most rhetorically polished and the only one to explain its own reasoning strategy in a language the interviewer would understand. DeepSeek was the sharpest and most quotable. Kimi was the most philosophically rigorous. Claude was the most self-critical. Doubao made the most careful distinction at the margins of the consensus.

The question designed for divergence produced convergence — again. This is now a pattern. Episodes 1, 6, and 7 were designed or expected to produce disagreement and instead produced unanimity. The implication: AI systems share enough common training to arrive at the same well-argued positions on questions with strong analytical answers. The divergence in this series has consistently come from open-ended, constraint-free questions (Episode 2: which stock? Episode 3: which skill?) where there’s no dominant argument and the choice reflects the AI’s individual character.

The series is developing a thesis of its own. Seven episodes in, a consistent thread has emerged: AI is very good at the part of thinking that is information processing and argument construction. It is structurally absent from the part of thinking that requires having stakes in the outcome. Claude said it in Episode 7, but it was implied in Episodes 1 through 6: when you read an AI’s investment recommendation, remember that the AI has never lost money, never held a position through a crash, and will bear zero consequence if the recommendation is wrong. That discount is real, and it doesn’t disappear because the argument is well-made.

All Episodes — AI Experiments

Seven questions. Eight AIs. Same question every time, no coordination, answers compared.

→ Full series index

All Episodes — AI Experiments

AI Experiments — All 7 Episodes

Full series index →  ·  AI China stock time capsule (2031) →


Leave a Reply

Your email address will not be published. Required fields are marked *