HNHacker News
TopNewBestAskShowJobs

bturtel

159 karma · joined August 31, 2019

Founder & CEO @ Lightning Rod Labs - www.lightningrod.ai Writing - bturtel.substack.com
submissionscomments
bturtel··on Show HN: Trained an LLM to predict "What will Trump do?"
Great question! It's probabilistic so not really "right vs wrong" on any single question, but who better estimated the likelihood. One big difference shows up when there's no useful context - we ran the same eval WITHOUT including any useful up-to-date context with questions. In this case, GPT-5 stays overconfident and its BSS drops to -11.3% (vs -4.3% ours) - worse than just guessing the base rate. So one advantage of the RL training is just learning to know what you don't know, and identify when there's real signal.
bturtel··on LLMs can teach themselves to better predict the future
Great question!

The key advantage of self-play is that we don't actually have labels for the "right" probability to assign any given question, only binary outcomes - each event either happened (1.0) or did not happen (0.0).

Our thinking was that by generating multiple predictions and ranking them by proximity to the ground truth, self-play incentivizes each agent to produce more finely calibrated probabilities - or else the other agent might come just slightly closer to the actual outcome.

bturtel··on LLMs can teach themselves to better predict the future
We're working on a follow up paper now to show similar results with larger models!
bturtel··on LLMs can teach themselves to better predict the future
Great read! Thanks for sharing.
bturtel··on Show HN: Log each meal with a photo to track your diet and monitor your weight
This could be huge. IIUC this is basically an AI-enabled version of MyFitnessPal, which has like 200M+ users, but with a massively streamlined user experience. Great idea.
bturtel··on Show HN: I made an app to challenge anxious thoughts with AI
This looks really cool - the UI in particular feels really approachable and polished. I like how you detect and call out "What you're doing wrong" to help build awareness of unhelpful thought patterns. Upvoted!

I recently launched something pretty similar (www.pensiveapp.com) but we took a very different approach (and I think the space is huge). Interesting to see how differently you approached the problem.

bturtel··on Show HN: Argument Encyclopedia: User-Posted Claims and Rebuttals
This is really cool - I think its really helpful in difficult conversations when you can encourage people to choose a single branch / claim of the argument and stick to resolving that before confounding by mixing in other claims. Personally, I'd love to be able to see a visual branching of different arguments / counterarguments, with some visual indicator for how much support each has.
bturtel··on Show HN: Demo of my web game about social persuasion
Yea, I think so - I could imagine this being really streamlined by just dropped me immediately into a conversation, with maybe the goal just written on a screen somewhere - no setup, no storyline, etc. I guess it just depends if most of your users are there for a gameplay experience vs a "practice" experience.
bturtel··on Show HN: Demo of my web game about social persuasion
I think this has a TON of potential. Situations like these are very non-obvious and anxiety-inducing for lots of people, so if you can make this a way for people to gain proficiency and confidence at navigating tricky social interactions, it could be a very powerful value prop. My only feedback would be that it took too long to get into the first challenge - lots of instructions / introduction / scene setting. Well done!
bturtel··on Show HN: Pensive – AI mental health coaching backed by science
Thanks! Great questions. We just launched - anecdotally we've had really great feedback from early users, but we're working with PhDs in the field to design an external validation study while tracking user-reported outcomes. We've also heard from a few providers that have started recommending Pensive to their clients for on-demand support between sessions. We see Pensive as complementary to therapy for some users and a standalone tool for others. Many people who don't need therapy can still benefit significantly from consistent, evidence-based practice. The key is making these proven techniques more accessible. Like physical health, most people don’t need a medical intervention - they need to exercise, but sorting through the research and implementing effective practice is a big barrier. Our focus is on delivering established practices in a convenient format with just a few minutes of guided conversation a day.
bturtel··on Show HN: I combined spaced repetition with emails so you can remember anything
This is very cool. Reminds me of Quantum Country (https://news.ycombinator.com/item?id=30467585) but for everything else in life.
bturtel··on Sunlight is more effective than censorship
I'm not sure what you're criticism here is - the article never uses the term "fake news", and it gives specific examples of factual inaccuracies promoted by the government. I fully agree that the cryptocurrency space if plagued with scams and theft, but I don't think that should prevent us from seeing the benefits of open & verifiable data that it has the potential to deliver.
bturtel··on Show HN: Model2vec – Lightning-fast Static Embeddings for RAG/Semantic Search
This seems awesome for enabling RAG queries for on-device LLMs.