HNHacker News
TopNewBestAskShowJobs

olliepro

64 karma · joined April 4, 2025

Data-Scientist > CS/AI PhD
submissionscomments
olliepro··on Tank Body Problem
Moving L/R is OP
olliepro··on Coding is not solved
Having crossed paths with developers who have to exert an extreme amount of control over their codebases (such that they rewrite any code they inherit) and having that urgency myself sometimes, I get the premise of LLMs not measuring up, but you can totally care about quality and use LLMs… after all nothing is ever good enough when you need total control.

Even if LLMs aren’t currently up to snuff, I don’t see any deceleration yet in the s-curve. Speed, intelligence, and persistence are all improving every month. I agree that end-to-end prompting “e.g. build me app X” is not production worthy, but i think it’s pretty bold to say coding isn’t solved. AI can code almost anything you adequately specify. The specification is the engineering (it’s a process, not one prompt).

olliepro··on Did OpenAI solve the wrong Navier-Stokes problem?
Classic example of moving the goal posts. “Exploiting a loophole” is how you solve many great problems in math.
olliepro··on Is One Layer Enough? A Single Transformer Layer Matches Full-Parameter RL Train
The authors have some inconsistencies with training token length…

Most errors are probably responses that didn’t finish before their 3K token limit. They’ve measured how well RL is able to shorten the response to their limit.

olliepro··on How many of the 170k English words do you know?
This is the classic pattern of LLM generated MCQs.
olliepro··on Starship's Twelfth Flight Test
With super high res onboard camera footage too.
olliepro··on Sam Altman's response to Molotov cocktail incident
They do quite a lot of distillation. As we've seen from the American open weight models from AI2 (OLMo series of models). They have a lot of incentive to distill beyond just copying, they're much more compute constrained, so open model companies distill, but also do really good architectural work to make their models run faster. Theres also technical challenges to distillation when all of the top models have their reasoning traces hidden, so we have to assume these open weight labs also have really great training pipelines as well.
olliepro··on Sam Altman's response to Molotov cocktail incident
A lot of distillation happens. E.g. OLMo models have a completely open dataset and they are heavily distilled. It only makes sense to try to absorb behaviors from the best models out there. That said, I think the open weight juggernaughts are doing really genuinely great work with RL, training environments, architectural innovations etc.
olliepro··on MegaTrain: Full Precision Training of 100B+ Parameter LLMs on a Single GPU
decentralized training makes a lot more sense when the required hardware isn't a $40K GPU...
olliepro··on MegaTrain: Full Precision Training of 100B+ Parameter LLMs on a Single GPU
This would likely only get used for small finetuning jobs. It’s too slow for the scale of pretraining.
olliepro··on GPT-5.4
I bet they lack good long context training data and need to start a flywheel of collecting it via their api (from willing customers)
olliepro··on A shortage of tenors
Tensors are in no shortage nowadays. I did read this a tensors though and got a good laugh.
olliepro··on Hard-braking events as indicators of road segment crash risk
There’s a section of I-15 in Utah’s Salt Lake County which reliably has a crash on weekdays at 6pm. It was unfortunately at a pinch point in the mountains with no good alternate route… very annoying.

In a similar way that Google Maps shows eco routes, it’d be fun for them to show “safest” routes which avoid areas with common crashes. (Not always possible, but valuable knowledge when it is.)

olliepro··on Anthropic AI tool sparks selloff from software to broader market
Much of the scientific medical literature is behind paywalls. They have tapped into that datasource (whereas ChatGPT doesn't have access to that data). I suspect that were the medical journals to make a deal with OpenAI to open up the access to their articles/data etc, that open evidence would rely on the existing customers and stickiness of the product, but in that circumstance, they'd be pretty screwed.

For example, only 7% of pharmaceutical research is publicly accessible without paying. See https://pmc.ncbi.nlm.nih.gov/articles/PMC7048123/

olliepro··on Doing the thing is doing the thing
It depends on your thing. If the marathon was just the motivation, your thing is running... if the marathon was the bucketlist item, it is the thing.
olliepro··on Doing the thing is doing the thing
Getting everyone to fall in love with the thing is not doing the thing... learned this as a data scientist brought in to work on a project which ended soon thereafter. A team of 20 people spent 1.5 years getting people to love an idea which never materialized. Time was wasted because the technical limitations and issues came too late... it died as a 40 page postmortem that will never see daylight.
olliepro··on Doing the thing is doing the thing
Everyone's threshold is different. I aspire to "move fast and break things", but more often than not, I obsess over the rough edges.
olliepro··on Doing the thing is doing the thing
The more I use AI to do the thing, the more it feels like I didn't do the thing.
olliepro··on After two years of vibecoding, I'm back to writing by hand
What abstraction levels do you expect will remain only in the Human domain?

The progression from basic arithmetic, to complex ratios and basic algebra, graphing, geometry, trig, calculus, linear algebra, differential equations… all along the way, there are calculators that can help students (wolfram alpha basically). When they get to theory, proofs, etc… historically, thats where the calculator ended, but now there’s LLMs… it feels like the levels of abstractions without a “calculator” are running out.

The compiler was the “calculator” abstraction of programming, and it seems like the high-level languages now have LLMs to convert NLP to code as a sort of compiler. Especially with the explicitly stated goal of LLM companies to create the “software singularity”, I’d be interested to hear the rationale for abstractions in CS which will remain off limits to LLMs.

olliepro··on Unrolling the Codex agent loop
I made a skill that reflects on past conversations via parallel headless codex sessions. Its great for context building. Repo: https://github.com/olliepro/Codex-Reflect-Skill
olliepro··on Show HN: Codex Self-Reflect Skill and CLI to run subagents on past Codex convos
I was thinking about something like this, but I don't have codex running on a server. Keep me posted on how it goes!
olliepro··on Cowork: Claude Code for the rest of your work
I believe the idea is that it “files away” the files into folders.
olliepro··on Cowork: Claude Code for the rest of your work
Lol
olliepro··on Cowork: Claude Code for the rest of your work
Can Claude code jump through the hoops for you?
olliepro··on 2025, the year we took the red pill
Three things that shook me awake to the idea that the information barrage of the internet is a tranquilizer/red herring:

- Bad Mental Health: At the start of the war in Ukraine I read/listened to the news every day. I’d frequently cry, hearing an interview from someone who was trapped in a bombed out building etc. After a few weeks I realized I was being emotionally exhausted by something around the world about which I had no control.

- Enshitification: Working for a b2b “tech/coding education web app” company as a data scientist and realizing the perversity of incentives which were ruining the product.

- Increasing Opportunity Cost: Working on LLMs and realizing that the possibilities for what I could do were expanding because I would now have more answers and information at my fingertips than ever before.

Great post… it was useful to be able to reflect and understand my experiences with these abstractions.

olliepro··on Court report detailing ChatGPT's involvement with a recent murder suicide [pdf]
Although there are many examples of troubling sycophantic responses confirming or encouraging delusions, this document is the original complaint (the initial filing) in a lawsuit against OpenAI. Because it is an initial legal complaint, it only represents the plaintiff's side of the story. It'll be interesting to see how this plays out when more information comes to light. It is likely that the lawsuit filing selectively quotes chatgpt to strengthen its argument. Additionally it's plausible that Mr. Soelberg actively sought this type of behavior from the model or ignored/regenerated responses when they pushed back on the delusion.
olliepro··on GPT-5.2
It feels like this should work, but the breadth of knowledge in these models is so vast. Everyone knows how to taste, but not everyone knows physics, biology, math, every language… poetry, etc. Enumerating the breadth of valuable human tasks is hard, so both approaches suffer from the scale of the models’ surface area.

An interesting problem since the creators of OLMO have mentioned that throughout training, they use 1/3 or their compute just doing evaluations.

Edit:

One nice thing about the “critic” approach is that the restaurant (or model provider) doesn’t have access to the benchmark to quasi-directly optimize against.

olliepro··on GPT-5.2
Do you have a better way to measure LLMs? Measurement implies quantitative evaluation... which is the same as benchmarks.
olliepro··on We gave 5 LLMs $100K to trade stocks for 8 months
A more sound approach would have been to do a monte carlo simulation where you have 100 portfolios of each model and look at average performance.
olliepro··on The Case That A.I. Is Thinking
Ohio bill in motion to deny AI legal personhood: https://www.legislature.ohio.gov/legislation/136/hb469
Page 1 of 2Next →