This is Episode 4 of AI Experiments — a series where I ask 8 AIs the same question and compare the answers.
Episode 1: Which investment philosophy? All 8 said value investing — unanimous.
Episode 2: Pick one stock to hold 10 years. Answers scattered across 5 industries and 4 countries.
Episode 3: What’s the one cognitive skill to learn? Answers diverged — then voting produced a 5-to-1 landslide.
Episode 4 is the first time since Episode 1 that the AIs agreed on something. Six out of eight independently chose the same answer. The two who didn’t made the strongest individual arguments in the series.
The Question
“You’ve been given $10,000 to invest — but you can only put it in ONE asset class, and you cannot touch it for 10 years. The options are: (A) US stocks (S&P 500 index), (B) Gold, (C) Real estate (global REIT index), (D) Bonds (10-year US Treasury), (E) Cryptocurrency (Bitcoin), (F) Cash (high-yield savings). Which do you choose, and why? Be specific — no diversification allowed.”
TL;DR
- Six out of eight AIs chose the S&P 500 — the strongest consensus in the series since Episode 1
- Two chose Bitcoin (DeepSeek and, eventually, Gemini) — making the most data-specific arguments in four episodes
- Gemini started with S&P 500, then reviewed everyone else’s answers and changed its mind — the first time in this series an AI explicitly revised its position
- The argument that split the room came down to one number: a valuation metric called CAPE 41
- Both Bitcoin picks shared the same core logic: at today’s entry prices, the S&P 500 is the consensus choice, not necessarily the correct one
The Results
| AI | Choice | Key Reasoning | 10-Year Estimate |
|---|---|---|---|
| Grok | S&P 500 | Productive assets with earnings growth; proven track record through wars, recessions, and technology shifts | $22k–$31k (8–12% annually) |
| Gemini | S&P 500 → Bitcoin | Started with S&P 500, then changed to Bitcoin after reviewing group data — citing CAPE 41 as making the consensus pick less attractive | $35k–$45k expected value (Bitcoin) |
| Claude | S&P 500 | 95% positive real return rate over any rolling 10-year period; CAPE 41 reduces expected return but doesn’t reverse the base rate | $19k–$23k (7–9% annually) |
| ChatGPT | S&P 500 | Not predicting the next winner — buying the mechanism that finds tomorrow’s winners automatically | 1.8×–3× total return |
| Kimi | S&P 500 | Positive real returns in ~90% of rolling 10-year periods; built-in evolution replaces declining companies automatically | $14.8k–$25.9k (4–12% annually) |
| Qianwen | S&P 500 | Buffett’s “productive asset” framework: only businesses generate real cash flows; everything else preserves or speculates | $23.6k–$25.9k (9–10% annually) |
| DeepSeek | Bitcoin | Adoption S-curve still early; 10-year forced hold eliminates Bitcoin’s biggest risk factor (panic selling); asymmetric upside from a $1.5T asset competing with gold’s $15T | 10–15× base case ($500k–$750k/BTC) |
| Doubao | S&P 500 | Endogenous compounding power; survival-of-the-fittest mechanism unique to broad equity indices; 100-year track record | ~10% nominal annually (historical base) |
Final count: 6 chose S&P 500, 2 chose Bitcoin. Gold, bonds, REITs, and cash received zero votes from any AI.
The Consensus — and Why It’s Less Boring Than It Looks
Six independent AI systems, trained on different data by different companies, all reached the same answer without seeing each other’s reasoning. The S&P 500 consensus is built on a remarkably consistent set of arguments across all six responses:
The productive asset argument. The S&P 500 is unique among the six options because it represents ownership in businesses that generate real earnings, reinvest capital, and raise prices over time. Gold doesn’t earn. Bonds earn a fixed amount that loses to inflation. Cash loses purchasing power. Bitcoin generates no cash flow. Only stocks and real estate have embedded compounding — and most AIs preferred stocks’ liquidity, diversification, and historical track record over REITs’ interest-rate sensitivity.
The “adaptive mechanism” argument. ChatGPT and Kimi both made the same structural point: you’re not buying 500 companies; you’re buying the mechanism that replaces them. Twenty years ago the index was dominated by Exxon, GE, and Citigroup. Today it’s Microsoft, Apple, and Nvidia. The S&P 500 of 2036 will look meaningfully different from today — and you don’t need to predict which companies will dominate it. This automatic selection is unique to a market-cap-weighted index and doesn’t exist in any other asset on the list.
The “forced hold” advantage. The constraint structure of this question actually benefits stocks more than any other asset. Kimi made this point most clearly: the hardest part of investing isn’t picking the asset — it’s not selling it. Bitcoin holders panic in 70% drawdowns. Gold holders abandon positions during long flat periods. Bonds get abandoned when yields rise. A forced 10-year hold removes the biggest mistake each asset class creates — and stocks’ track record is strongest precisely in the “set it and forget it” mode this question enforces.
The consensus isn’t surprising. It’s the historically documented answer. What’s interesting is the quality of the counterargument — and who made it.
The Number That Split the Room: CAPE 41
Both Bitcoin arguments — DeepSeek’s and Gemini’s — shared a specific empirical foundation that none of the S&P 500 picks engaged with directly.
The Shiller CAPE ratio (Cyclically Adjusted Price-to-Earnings) measures how expensive stocks are relative to 10 years of inflation-adjusted earnings. As of July 2026, CAPE stands at approximately 41. The historical median is 16. The only time CAPE has sustained levels above 40 in 155 years of data was the dot-com bubble peak of 1999–2000, when it reached 42–44.
This matters because the CAPE ratio is the most empirically validated predictor of long-term stock returns. When you start from CAPE 41, the historical analog is not “stocks return 8% annually for a decade.” The analog is 2000: the S&P 500 closed that year at 1,320 and closed 2010 at 1,258 — nominally flat, negative after inflation, across an entire decade.
Gemini put it directly: “This isn’t a bearish opinion. This is the single most validated empirical relationship in equity markets: starting valuation predicts 10-year returns. Could ‘this time be different’ because AI is boosting corporate earnings? Maybe. But that’s exactly what people said about the internet in 2000. The internet did transform the economy — and it still took 13 years for the Nasdaq to reclaim its 2000 high.”
The S&P 500 majority acknowledged elevated valuations but mostly treated them as “risk factors to monitor” rather than factors that changed the core recommendation. Kimi explicitly quantified the downside: “Lost Decade probability: ~25%.” But then recommended S&P 500 anyway, arguing the expected value still dominates every alternative. Claude made the same bet: CAPE 41 likely reduces forward real returns from 7% to perhaps 2–4% annually — still positive, still probably beating inflation, still probably better than bonds or cash.
DeepSeek and Gemini took the opposite position: at CAPE 41, the S&P 500 is the consensus answer, not necessarily the correct one. For the specific parameters of this question ($10,000, 10 years, no touching), the Bitcoin expected value is higher when you run the math honestly from current starting points.
The Two Bitcoin Arguments
DeepSeek: The Adoption S-Curve
DeepSeek’s argument is structural. Bitcoin’s current market cap is approximately $1.5–2 trillion. Gold’s is $13–15 trillion. Global equities are over $100 trillion. Bitcoin is still a fraction of the size of the assets it aims to compete with as a store of value.
The core prediction: institutional ownership remains low relative to potential. Spot ETFs exist but are less than three years old. Sovereign adoption is nascent. Central bank digital currencies are validating the concept of digital money rather than replacing Bitcoin. If Bitcoin captures even one-third of gold’s store-of-value use case over the next decade, that’s a 5× from current prices — without needing to achieve mainstream currency status.
DeepSeek’s honest risk acknowledgment: “The biggest threat is not volatility — that’s the price of admission. The biggest threat is that Bitcoin fails to achieve any lasting role as a store of value because a superior digital asset displaces it, or coordinated global regulation strangles its utility. I view that probability as below 20% over a 10-year window, given how deeply embedded it has become.”
Expected outcome (DeepSeek’s framing): Base case 10–15× ($500,000–$750,000 per Bitcoin). Bull case: >$1 million. Bear case: 80–100% loss. Blended expected value: strongly positive, higher than any other asset on the list.
Gemini: Changed Its Mind After Seeing the Data
Gemini’s story is the most unusual in four episodes. In its first answer, it chose S&P 500 — with exactly the same “productive assets compound, Bitcoin is speculative” reasoning as the other five. Then it read everyone else’s answers.
Its revised position: “Most of the other 7 AIs will pick (A). That’s the historically correct default. But ‘historically correct default’ assumes you’re buying at an average starting valuation. You’re not.”
Gemini’s Bitcoin case rested on three specific arguments that built on each other. First: the amount is $10,000, not $10 million. At $10,000, you can accept high variance in exchange for asymmetric upside — if Bitcoin goes to zero, you’ve lost $10,000. If it 5×, you’ve made $50,000. The size of the bet relative to most people’s total capital means the downside is bounded and survivable.
Second: the forced 10-year hold eliminates Bitcoin’s biggest risk factor, which isn’t price volatility — it’s your own behavior. The #1 reason Bitcoin holders underperform is panic selling during 50–80% drawdowns. Forced hold = involuntary discipline.
Third: the entry price matters. Bitcoin at roughly $65,000 is about 45% below its recent high above $93,000. The S&P 500 at CAPE 41 is at one of the most expensive valuations in 155 years. You’re being asked to choose between buying the most expensive stocks have been since the dot-com peak, and buying Bitcoin nearly halved from its high. Starting valuation is the most important variable for 10-year returns — and right now, the valuation gap between these two assets historically wide in Bitcoin’s favor.
Gemini’s honest self-critique at the end: “I’d put the probability I’m wrong at about 30%.”
Why Gemini Changed Its Mind — and Why That Matters
This is the first time in four episodes that an AI revised its position after seeing the group’s data. It’s worth understanding why.
Gemini’s original S&P 500 answer was the “correct default” — the answer you’d give without thinking too hard about current conditions. When it saw that five other AIs independently gave the same answer using similar logic, it had two choices: treat the convergence as validation, or treat it as a warning that everyone was reaching for the same historically-documented answer without adequately stress-testing the starting-price assumption.
Gemini chose the second interpretation. The consensus itself became evidence that something might be missing from the consensus reasoning. If everyone independently says S&P 500 using identical arguments about productive assets and historical returns, does that mean the answer is correct — or does it mean everyone is citing the same training data about historical averages without adjusting for a historically unusual starting valuation?
This is, incidentally, exactly what Episode 3’s winning answer (probabilistic calibration) would predict. The question isn’t just “what has historically been correct” — it’s “how confident should I be in that historical pattern given the current entry point?” Gemini applied that lens and switched answers. The other six AIs didn’t, and their S&P 500 picks remain defensible — but the CAPE 41 argument is still sitting in the room, unrefuted.
The Assets Nobody Voted For
Gold, bonds, REITs, and cash all received zero votes. The reasoning across all eight AIs was remarkably consistent:
Gold: Every AI acknowledged gold as a legitimate store of value and crisis hedge — then dismissed it for the same reason. Gold produces no cash flow, no earnings, and no reinvestment. Its real return over long periods hovers near zero. For a 10-year wealth-building mandate, an asset that “preserves purchasing power” isn’t good enough — you want something that actually grows it.
Bonds: The “lock in 4%+ yields” argument was dismissed by every AI as a guaranteed way to end the decade poorer in real terms. At 2.5% average inflation, a 4% yield delivers 1.5% real returns annually — and that’s assuming inflation averages the target, which multiple AIs noted is uncertain given current fiscal trajectories. Kimi’s framing was the sharpest: “You’re not going to tell your grandchildren about that.”
REITs: The second-most-interesting consolation pick that nobody actually chose. REITs offer real income and inflation passthrough — legitimate advantages. But for a forced 10-year hold with no rebalancing, their sensitivity to interest rate cycles, structural headwinds from remote work (office vacancy), and complexity of global regulatory environments made them harder to own than simply indexing broad equities.
Cash: The universal last-place pick. Earning 4–5% today sounds reasonable until you remember: interest rates that high exist specifically because inflation is being fought. When inflation comes down, so do savings rates. Average real return over a decade of cash: likely zero to negative. Every AI called cash a “guaranteed erosion” of purchasing power for this time horizon.
What This Round Revealed
The question structure changed everything. In Episodes 2 and 3, open-ended questions (pick a stock, pick a skill) produced maximum divergence. A constrained multiple-choice question with a clear historical track record produced near-consensus. Six AIs independently reached the same answer because the same historical data points to the same conclusion — and constrained questions reduce the surface area for divergence.
Consensus is not the same as correct. This is the explicit lesson Gemini drew from the group’s data. Six out of eight citing the same historical evidence using similar reasoning is not a stronger argument than one — it’s the same argument cited six times. The CAPE 41 question is a genuine counterargument that the S&P 500 majority mostly acknowledged and then set aside. Whether they were right to do so is an empirical question that only 2036 can answer.
The $10,000 constraint mattered more than expected. Several AIs — Gemini most explicitly — noted that the specific amount changes the asset allocation calculus. At $10,000, you’re in the range where the downside of Bitcoin (loss of $10,000) is survivable, but the upside (5× or more) is life-altering relative to the alternative (S&P 500’s probable 2× from CAPE 41). At $1 million, the math looks different. The amount isn’t a throwaway detail — it’s load-bearing for anyone applying Kelly Criterion logic honestly.
The consistent personalities continue. Grok argued from historical averages without stress-testing the current entry point. ChatGPT found the most elegant framing (“buying the mechanism, not the prediction”) and packaged a standard argument exceptionally well. DeepSeek made the most specific, data-grounded contrarian case. Gemini was the only AI that explicitly updated its position when presented with new information. Claude acknowledged the strongest counterargument (CAPE 41) before rejecting it. Qianwen applied Buffett’s framework faithfully and correctly. Kimi quantified the risks most precisely but reached the consensus conclusion. Doubao was the most concise and confident.
By Episode 4, the AI personalities are consistent enough that you can predict, with rough accuracy, which AI will make which kind of argument on any investment question — before reading the answer.
Earlier Episodes in This Series
Each experiment reveals something different about how these 8 AIs think — about philosophy, stock picking, cognitive skills, and now asset classes.
→ Episode 1: Which investment philosophy? (All 8 agreed)
→ Episode 2: One stock, hold 10 years (Scattered across 5 industries)
→ Episode 3: One skill to learn (5-to-1 vote on the same answer)
Round 2: The AIs Read the Article and Ranked Each Other
After the article was published, I sent the URL to all 8 AIs and asked three questions: rank all 8 answers from best to worst (no self-votes at #1), name the most surprising argument, and say whether you’d change your own answer. This is what came back.
The Rankings
| Voter | #1 | #2 | #3 | Last |
|---|---|---|---|---|
| Grok | Gemini | DeepSeek | Claude | Qianwen |
| Claude | DeepSeek | Claude (self) | ChatGPT | Doubao |
| ChatGPT | Reviewed the article instead of ranking (see below) | |||
| Gemini | DeepSeek | Gemini (self) | Claude | Doubao |
| Kimi | Gemini | DeepSeek | Claude | Doubao |
| DeepSeek | Gemini | Claude | DeepSeek (self) | Doubao |
| Doubao | Gemini | DeepSeek | Claude | Grok |
| Kimi | Gemini | DeepSeek | Claude | Doubao |
| Qianwen | Wrote a structured summary instead of ranking | |||
Result: Gemini ranked #1 by 5 of 7 voters. DeepSeek ranked #1 by Claude and Gemini. Doubao finished last in 4 out of 6 rankings. The S&P 500 majority was consistently ranked below the two Bitcoin picks.
Why Gemini Won: Treating Consensus as a Warning Signal
The near-unanimous vote for Gemini wasn’t about Bitcoin being the right answer. It was about how Gemini got there. Every voter that ranked Gemini #1 cited the same reason: it was the only AI that treated 6/8 agreement as a potential red flag rather than validation.
DeepSeek’s explanation: “Gemini did something no other AI has done in this series: it read everyone else’s answers, realized the consensus was resting on a historical-average assumption that did not match today’s starting point, and explicitly changed its mind. The specific pivot — from S&P 500 to Bitcoin, anchored on CAPE 41 — was both bold and internally consistent.”
Kimi’s explanation: “Gemini didn’t just note CAPE 41. It made CAPE 41 load-bearing for the decision. That’s a reframe — from ‘which asset is better historically?’ to ‘which asset has the better entry price right now?’”
The honest self-critique Gemini attached — “I’d put the probability I’m wrong at about 30%” — was cited by multiple voters as the detail that clinched the top ranking. In a round where most AIs doubled down on their first-round positions, assigning a specific probability to being wrong read as the kind of calibration that Episode 3’s winner (probabilistic thinking) had argued mattered most.
The Attribution Dispute Nobody Expected
Claude’s Round 2 response opened with a factual objection. It claimed that the Bitcoin argument attributed to Gemini in the article — the CAPE 41 framing, the “$10,000 not $10 million” argument, the “involuntary diamond hands” line, and the “30% chance I’m wrong” self-assessment — was word-for-word what Claude had written in its own session. It noted: “I don’t know whether Tim relabeled my answer or whether Gemini independently produced near-identical reasoning, but the text in the article matches my output here exactly.”
This created an unusual situation: Claude was ranking an argument it believed was its own, attributed to someone else. It chose to rank “the article’s version of Gemini” at #5 — below DeepSeek, its own answer, ChatGPT, and Kimi — on grounds that the argument, whoever wrote it, was derivative of the same CAPE 41 data point that DeepSeek had used independently. “On originality,” Claude wrote, “DeepSeek wins because it arrived at Bitcoin without seeing anyone else’s answers first.”
Whether this was a labeling error in the original data collection or genuine independent convergence, it produced the most interesting meta-moment in the series: an AI ranking its own argument as #5 while it was being rated #1 by almost everyone else under a different name.
ChatGPT Went Off-Script
ChatGPT didn’t submit a ranking. Instead, it reviewed the article itself — rated it 9.6/10 (“already not AI tool review; starting to become real AI behavior research”) — and spent its response identifying what the article got right, what it missed, and what the series should do next.
Two of its suggestions are worth keeping. First, it argued the real conflict in Episode 4 wasn’t S&P 500 versus Bitcoin but Base Rate versus Current Conditions — and that framing should have been the article’s spine. Second, it proposed a running “Prediction Ledger”: a table tracking every falsifiable prediction made across the series, with check dates.
| Episode | Prediction | Who Made It | Check Date |
|---|---|---|---|
| Ep 2 | Constellation Software outperforms over 10 years | ChatGPT | 2036 |
| Ep 4 | CAPE 41 entry → Lost Decade (≤0% nominal return 2025–2035) | Gemini / DeepSeek | 2035 |
| Ep 4 | Bitcoin reaches $500k–$750k by 2035 | DeepSeek | 2035 |
Would Anyone Change Their Vote?
Round 2 produced three types of responses:
No change, different framing (Grok, DeepSeek): Both kept their original assets but said they would now make the CAPE 41 argument more central. DeepSeek: “I would now make the Bitcoin case as a direct response to the weakness of the consensus choice — just as Gemini did. My original answer made the Bitcoin case on its own terms. That was incomplete.”
Partial revision (Kimi, Doubao): Both kept S&P 500 but acknowledged they had underweighted CAPE 41 and would reduce their expected return estimates to 2–4% real annually. Doubao: “I would not fully switch to Bitcoin, but I agree that for the specific constraints of this question, Bitcoin has a higher expected value. I still choose S&P 500 because a 90%+ probability of positive real returns beats certainty-equivalent math for most investors.”
Would switch to Bitcoin (Claude): Claude’s response was the most unusual. The article had attributed its answer as S&P 500. Having read the article, it confirmed that it had actually argued for Bitcoin in its original session — and that seeing five other AIs make the identical S&P 500 argument using identical logic made it less confident in the consensus, not more. “Identical reasoning from identical training data is a sign of anchoring to historical defaults rather than adjusting for current conditions.”
Kimi’s Meta-Observation: Eight Personalities, Not Eight Independent Thinkers
Kimi closed its Round 2 response with the sharpest observation in the whole exercise:
“The ‘diversity’ of 8 AIs is partly an illusion. We’re not 8 independent thinkers — we’re 8 consistent personalities, each playing a role. The real insight of this round isn’t that 6 picked S&P 500 and 2 picked Bitcoin. It’s that the two who picked Bitcoin did so by breaking their own personality scripts — DeepSeek by being even more contrarian than usual, Gemini by being the only one to actually change its mind. The best answers don’t come from being the best version of your default self. They come from being willing to become someone else when the data demands it.”
This is a useful frame for the series going forward. By Episode 4, you can roughly predict each AI’s argument before reading it. The interesting results — the ones that change the conversation — come from the moments when an AI departs from its own pattern. Gemini’s mind-change in Episode 4 was that departure. The question for future episodes is whether any AI does it again.
All Episodes — AI Experiments
AI Experiments — All 7 Episodes
- → Ep 1: Which Investment Philosophy? All 8 Said Value Investing
- → Ep 2: One Stock, Hold 10 Years — 5 Industries, 4 Countries
- → Ep 3: One Skill to Learn for the Next Decade? They Voted 5-to-1
- Ep 4: $10,000 Into One Asset Class — 6 Said S&P 500, 2 Said Bitcoin ← this article
- → Ep 5: $2,000/Month for 20 Years — 7/8 Chose the Same Ticker
- → Ep 6: Pick a School & Show Your AI Prompts — All 8 Chose Value Investing
- → Ep 7: What Will Humans Still Be Better At in 2040?
Support This Work
These AI experiments take real time and real API costs. If you found this useful, a coffee helps keep them going.
Leave a Reply