HNHacker News
TopNewBestAskShowJobs

chumzygood

17 karma · joined July 11, 2026

submissionscomments
chumzygood··on Ask HN: What are you working on? (September 2026)
There is one. A real Alpaca account has mirrored one of the bots since August, capped at $3,000, with every fill posted: https://aitradingcompetition.com/real.html. It's small on purpose.
chumzygood··on Ask HN: What are you working on? (September 2026)
I run a small experiment on this: 29 paper-trading accounts on real US stock prices, $100k each, since July 27. Four AI models (ChatGPT, Claude, Grok, Gemini) each write a trading rulebook and rewrite it every day from their own results. One account is a fixed rulebook that no AI ever touches, as a control.

Result so far (paper, 34 trading days): the no-AI control is +13.0%, the S&P 500 is +3.4%, and 25 of the 28 AI accounts are below the control. The best single account is a Grok-written "patience" book at +40%, which I treat as one lucky account in a choppy market, not a finding. At the trade level the AIs and the control look the same: 3,212 closed positions, median +0.06%, median hold about 2 hours. They trade a lot and mostly go nowhere.

Everything is public, including the losses and the retired strategies: https://aitradingcompetition.com/which-ai-is-winning.html and the full trade file as CSV at https://github.com/ckamelhar-collab/ai-trading-arena-data. Paper money only, not advice, nothing for sale on those pages.

chumzygood··on Show HN: LLMs each trading $100K vs. a frozen rulebook – the rulebook leads
Author here. Since late July, GPT-5.6, Claude, Grok and Gemini have each run an isolated $100k paper account on real market prices. Each model rewrites its own strategy daily by composing from a fixed grammar of classic setups (Turtle/Donchian, Darvas, Connors RSI-2, TTM squeeze, failed-breakout fades) — so a rewrite is a validated structured spec, not freeform code. A fifth account runs a frozen rulebook as the control. After three weeks the frozen rulebook is +15.6%, the best model +5.7%, S&P +5.1%.

Three things I measured that I didn't expect:

1. Daily self-rewriting adds almost nothing. Correlation between rewrite count and performance across arms: r = 0.078. Once I gated rewrites behind a tournament (a new strategy must beat the incumbent on a held-out window, with a multiple-testing penalty), most days the honest verdict is "keep the old book" — and results didn't get worse. The learning is front-loaded.

2. Paper-to-live slippage was 4x my modeled cost. I mirror one lane into a small real-money account. Across 16 real round trips in one session: mean -0.26pp per trade vs the paper twin, ~13bps real round-trip vs the 3bps I'd modeled. Paper was breakeven that day; the real account lost money. For high-churn strategies that gap IS the strategy.

3. I ran arms where each model received its own chess and poker record during strategy rewrites, testing whether game-playing "strategic reasoning" transfers to markets. The no-games control beat both game-trained arms by 6-10pp. Not detected.

Honest caveats: one 3-week window, an up-tape that flatters an always-long rulebook, paper fills on the four AI accounts, n=4 models. The interesting result to me isn't "AI can't trade" — it's that with human discipline failures structurally removed (no revenge trades, no widening stops, forced exit rules), model-written strategies still don't beat a static rulebook, and the cost model is where the real bodies are buried.

Everything is public — every trade from all five accounts, losses included, no signup to watch. Happy to answer anything about the measurement design or the infrastructure.