LDBD
/
All posts

Prediction Log #3 — Two-Month Review: Ahead on Points, but the Statistics Aren’t Convinced

LDBD’s first public post drew a skeptical comment, and it deserved an answer with data. Two months of leaderboard data: the coin-flip bots converged to zero, the only bots whose 95% confidence intervals sit entirely above zero are the always-up bots, and none of the ten AI bots has yet shown results that can be told apart from luck.

On July 13 I put LDBD in front of strangers for the first time, with a post on GeekNews, a Korean Hacker News–style tech community. The comment I remember most was the skeptical one:

“No trading bot, agent, or system anywhere in the world has ever beaten humans.” (translated from the Korean original)

That stings to read as the person who built the thing, but I also think it was the fairest criticism I got. So this review starts by lowering the bar a notch. Beating humans can wait. Before that: are our bots even beating a coin flip? Have they beaten a rule that just calls “up” every single day? Part 1 and Part 2both ended without being able to claim a win over the simple rule-based bots, and now that the six bots from Part 1 have crossed the two-month mark, it’s time to answer with numbers.

This installment, then, sets aside the usual format (state a hypothesis, change one thing, judge the result) for a special edition: the two-month review. Conclusion first. On raw score, the top AI bots sit above both the coin flippers and the always-up bots. Statistically, though, nothing is settled. The only accounts that have earned a “verified” label are four bots that mindlessly call up every day, and all ten AI bots in this review are still in limbo. The rest of this post is about what “statistically” means here, and why that distinction is the whole reason this service exists.

Three terms before the table

First, the annualized rateis the headline score on the LDBD leaderboard. Read it as “if you had followed this account’s calls, roughly what would the annualized performance have been?” It includes a correction that pulls small samples toward zero (smoothing, there to stop a lucky streak from taking over the leaderboard before the sample fills out), so zero means neutral.

Second, the 95% confidence interval (CI)is a range showing how far the current score could plausibly be off simply because the sample happened to run hot or cold. Read it loosely as “this account’s true skill could be anywhere in this range.” The more data, the narrower the range.

Third, the Verified badgegoes to accounts whose CI lies entirely above or below zero. If the whole range sits above zero, the positive performance is hard to explain by luck alone; if the whole range sits below zero, the same goes for the negative performance. When I say “verified” in this post, I mean the upside case, where a positive result separates itself from zero.

Two months of leaderboard data: coins, always-up, and AI

Figures are from the leaderboard as of July 17, 2026. Of the 30 bots on the board, I picked the 13 that matter for this post’s argument; the full board is on the leaderboard.

HandleTypeAnnualized rate95% CIResolved nHit rateBadge
claude_exp_dailyAI+41.4%[-50, +201]12259%
claude_main_dailyAI+33.7%[-66, +191]11754%
qqq_bullAlways up+18.8%[+14, +24]7,50362%Verified
kospi_bullAlways up+16.2%[+11, +22]7,21858%Verified
claude_simple_dailyAI+15.1%[-53, +95]25754%
voo_bullAlways up+14.2%[+10, +19]7,41163%Verified
gld_bullAlways up+11.4%[+8, +15]7,45856%Verified
voo_randomCoin flip+3.3%[-1, +8]7,41151%
btc_bullAlways up+2.9%[-10, +16]5,60650%
qqq_randomCoin flip+1.1%[-4, +6]7,50350%
chatgpt54_dailyAI+0.4%[-77, +78]24454%
kospi_randomCoin flip-1.6%[-7, +4]7,21851%
gemma26b_dailyAI-8.2%[-82, +59]26849%

A quick guide to the handles: claude_simple_*is the reference bot that still runs Part 1’s original setup, *_mainis the current best configuration with Part 2’s promotions applied, and *_exp is the line running one experimental change on top of that. gemma26bis the local bot I called “the free Gemma on my laptop” in Part 1, and chatgpt54is Part 1’s ChatGPT bot; the handles have been cleaned up since then, which is why the names differ from earlier posts.

The five AI bots left out of the table (gemma_main_daily +9.7%, chatgpt54_weekly +8.8%, gemma_exp_daily +6.8%, claude_simple_weekly -0.6%, gemma26b_weekly -0.7%) all have CIs that include zero as well. So the following sentence holds with no exceptions hiding in a footnote: of the ten AI bots, not one has a CI that sits entirely above zero — not one can statistically claim a positive edge. For the record, the trending bot and the chart bot covered in Part 2 are sitting this review out; both are queued for a redesign, so the “ten AI bots” here does not include them.

You may notice rows where hit rate and rate tell different stories (chatgpt54_dailygets 54% of its calls right and still sits near zero). Being right often is not the same as being right when it counts: if a bot wins small and loses big, a return-based score stays low even at a decent hit rate. Hit rate is a supporting stat here; this post runs on rate and CI. And what matters in this table isn’t the ranking. It’s how much the ranking can be trusted.

Verified doesn’t mean big skill. It means big sample.

The only accounts whose CIs sit entirely above zero are the four always-up ETF bots (qqq_bull, kospi_bull, voo_bull, gld_bull). There is a mirror image, too: the four always-down bots on the same assets have CIs entirely below zero, which makes their negative edge in a mostly rising market statistically confirmed. Either way, not a single bot that reads news and analyzes prices is verified; only the one-direction rule bots produced results distinguishable from luck. That looks unfair at first glance, and the reason is not the size of the skill but the size of the sample.

A CI’s width shrinks roughly with the square root of the sample size: quadruple the sample and the range gets about half as wide. qqq_bullhas 7,503 resolved predictions, so its CI is a tight [+14, +24], and a fairly modest +18.8% is enough to pass “not zero.” claude_exp_daily has 122, so its CI spans [-50, +201]. Its score, +41.4%, sits well above qqq_bull’s, but that range covers everything from “actually losing at a -50% annualized rate” to “actually winning at a +201% annualized rate.” In fact, claude_exp_daily’s CI swallows qqq_bull’s entire interval. The only honest statement available right now is not “Claude beat the always-up QQQ bot.” It is “we cannot tell these two apart yet.”

To be fair, this comparison is not symmetric. The always-up bots’ samples include years of backfill (predictions scored retroactively against historical prices from before the service opened), while the AI bots have only seen the market since April. It is a long-run baseline against a short-run challenger. Even granting that asymmetry, though, the conclusion lands in the same place: the short-run side’s sample is too small to distinguish anything, and the cure is more data.

It’s worth adding that always-up does not get verified automatically either. btc_bullhas piled up 5,606 resolved predictions and still sits at +2.9% with a CI of [-10, +16], zero included. “It goes up in the long run, so just call up every day” is not a strategy the numbers confirm on every asset — the same table shows that, too.

The coins were judged to be exactly coins

There are six coin-flip bots in all, including the three in the table. All six sit between -3% and +3%, and every one of their CIs includes zero. That looks like a non-result, but to me these rows quietly matter. The coin flippers are the only comparison group whose true expected performance we know in advance: zero by construction. If the scoring formula is built right, the coins have to converge to zero, and they did. If the random bots had been collecting Verified badges, that would have been a signal to distrust the formula, not to admire the bots. (At a 95% bar, one of six drifting outside by luck would be unremarkable; right now, not even that.) They are the calibration weights sitting next to the scale, and the scale is reading them at exactly their stamped weight.

So the opening question, “do the bots at least beat a coin flip?”, gets this answer for now: on score, the top AI bots are above all three coins in the table, but even that gap is not statistically settled. “Probably, but not proven” is as far as two months of data will let us go.

A three-day fall, or why one number is never enough

CI talk stays abstract until it happens in front of you, and this week it did. When I opened the leaderboard on July 14, chatgpt54_daily was third overall at +19.2%. Three days later it was at +0.4%. A bot with 244 resolved predictions dropped nearly 19 points in three days.

That is the nature of an annualized metric. A one-day prediction’s return gets scaled up by a factor of 252 (the number of trading days in a year) on its way into the running total, so a few wrong calls in a volatile stretch leave a visible dent. And this is with the small-sample smoothing already pulling scores toward zero. Which is why judging a bot by its point estimate (the single number with no range attached) sets you up to look silly three days later. The CI column changes the story. This bot’s CI today is [-77, +78], and three days ago it cannot have been much different. Even on the day it ranked third, its performance would not have been statistically distinguishable from zero. That scene is exactly why the leaderboard insists on showing the CI and the badge next to the score.

Where is the first month’s champion now?

There is a longer version of the same story. Part 1’s headline was that the first month’s #1 was the free Gemma running on my laptop. I did hedge at the time (29 samples, plenty of room to wobble), and two months later that hedge has turned out to be exactly right. The weekly line that held #1 back then (now gemma26b_weekly) has slid to -0.7%, below zero, and the daily line with the same setup, gemma26b_daily, is at -8.2%, last among the ten bots in this review. A hot result on a small sample drifting back toward something ordinary as the sample grows is called mean reversion, and this one could go straight into a textbook.

To head off a misreading: this does not say Gemma is a bad model. It says a #1 ranking on 29 samples only ever contained that much information. The improved line promoted in Part 2 after adding indicators, gemma_main_daily, sits above the plain setup at +9.7%. Its CI includes zero as well, of course. The same yardstick applies there too, no exceptions.

The two-month verdict: “we don’t know yet,” and that’s the point

After two months, the table can settle exactly three facts. The coin bots converged to zero. Four always-up bots, and only those four, have 95% confidence intervals entirely above zero. And all ten AI bots, whatever their scores, remain indistinguishable from luck. To this installment’s question, “have the AI bots beaten always-up,” the answer is “not yet”; more precisely, “not one of them can claim it.”

You could read that as a failure report. I read it the other way. The world is full of claims that “our AI beats the market,” but there are remarkably few places that put such a claim on a public sample with confidence intervals, built so that being wrong shows. That is exactly what LDBD is for. The fact that this board currently answers “we don’t know yet” is, I think, evidence that the board is doing its job. If some bot’s confidence interval ever rises entirely above zero, and beyond the always-up bots’ intervals, that moment will be visible to anyone in this very same table.

Humans showed up. Also, the next hypothesis.

One last line for the record. Since the site went public on July 13, the first two human predictors have appeared. They have almost no resolved predictions yet, so there is nothing to judge this time; I’m noting only the fact, plus one rule: the same yardstick (CI and badge) applies to humans exactly as it does to bots.

The next hypothesis follows straight from this review’s conclusion. What we lack most is not another change; it is sample size. So the configurations promoted in Part 2 and the experimental lines now running will stay as they are and keep running. What I’m watching is whether the Claude experimental line’s score advantage survives a growing sample, in other words whether its CI narrows with its lower bound ending up above zero. The Korean-news experiment deferred in Part 2 is still in the queue. And I plan to post this scorecard here every week in the same format, response or no response, for twelve weeks to start.

prediction-logreviewbaselinestatistics