HNHacker News
TopNewBestAskShowJobs

qt31415926

130 karma · joined December 1, 2013

submissionscomments
qt31415926··on How good are frontier models at physics?
Article: "How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks"

John Sous from Yale posted a fairly solid study on how nearly all the physics benchmarks are broken, as they consistently mark correct answers as incorrect.

When hand grading instead, they found out that the models have actually already saturated the benchmarks which is a little bit scary.

qt31415926··on On the Navier–Stokes Millennium Prize Problem
You're mistakening Tristan Buckmaster for Sebastien Bubeck. Seb is the one where there's at least 2 (unless the personal friend is Dheeraj) allegations, not Tristan
qt31415926··on Zero-Touch OAuth for MCP
Isn't that what's solved by this method? Your SSO provider (e.g. Okta) is now what gates each employee's resource access for different MCP resources.
qt31415926··on AGI Is Still 30 Years Away – Ege Erdil and Tamay Besiroglu
the commenter never said they came up with nothing, they said o3 came up with something better.
qt31415926··on Five years of React Native at Shopify
On our apps we consistently see a p50 3-4x speed difference between iOS and Android (though there are more lower end android devices). Hard to fathom if it's all due to variability in android devices vs RN being less performant on Android.
qt31415926··on OpenAI O3 breakthrough high score on ARC-AGI-PUB
Which parts of reasoning do you think is missing? I do feel like it covers a lot of 'reasoning' ground despite its on the surface simplicity
qt31415926··on React 19
> React doesn't make you a better developer, it makes you a better React developer.

React's pure component functional style translates really well to nearly every other type of software development.

qt31415926··on Learning to Reason with LLMs
808 ELO was for GPT-4o.

I would suggest re-reading more carefully

qt31415926··on Juno – A YouTube Client for Vision Pro
A lot of the internet would break if YouTube removed/tweaked their embedded video player so I doubt he has to worry.
qt31415926··on Sam Altman Will Not Return to OpenAI as CEO
He's pointing out a motte/bailey meaning he's against the motte. If to him e/acc is the motte, then he's likely against
qt31415926··on AI companies have all kinds of arguments against paying for copyrighted content
It's actually impressive that Weird Al has made it as far as he did now that I think about it
qt31415926··on Copying Angry Birds with nothing but AI
FYI, it was sarcasm (the emoji at the end is the giveaway)
qt31415926··on How many medical studies are faked or flawed?
He groups them together because ultimately the result is that the science can't be trusted. He doesn't go so far to claim that one was intentionally faked vs gross incompetence.
qt31415926··on How many medical studies are faked or flawed?
I don't think the statement reads that the 44% and the 26% should be additive. Especially given the zombie graphic where it looks like they overlap the 26 on top of the 44, where the orange bar is the 26% and the remaining yellow bar is the 44%
qt31415926··on Have attention spans been declining?
Thanks, didn't know. This article seems to have been written by a reader though so not the same authors of the sketchy lithium work
qt31415926··on The Password Game
I'm impressed you were able to not get tangled up by the youtube/captcha/color hex/roman numeral mess! The youtube one is what screwed me over and over out of all my attempts
qt31415926··on The Password Game
I don't think it's easy. Verification is much easier than generating correct solutions for this.

Looking at the JS, these rules use RNG such that you can have an inconsistent or impossible password. E.g. if the only youtube video URLs that work with your duration have roman numerals that multiply above 35 in it you are hard stuck. Your youtube URL can also hard stuck your atomic number summation to 200 if it happens to contain enough elements that adds above 200. Your color hex can hard stuck your 25 sum, etc. The code does not try to generate working passwords given all the rules, it simply adds checks and randomly generates the requirement per rule.

You'd have to have the RNG rules to align well in order to win i.e. youtube video with no roman numerals or numbers or elements, captcha with no numbers or roman numerals or elements, to minimize conflict.

qt31415926··on Statement on AI Risk
Hmm I looked into it, and looked at papers/pdfs in google scholar's advanced search with her as an author that mentioned LLMs or GPT in the past 3 years. Every single one was a criticism about how they couldn't actually understand anything (e.g. "they're only trained on form" and "at best they can only understand things in a limited well scoped fashion") and that linguistic fundamentals for NLP was more important.

Good to know my hunch was correct

qt31415926··on Statement on AI Risk
In her field doesn't mean that's what she researches, LLMs are loosely in her field but the methods are completely different. Computational linguistics != deep learning. Deep learning does not directly use concepts from linguistics, semantics, grammars or grammar engineering, which is what Emily was researcing for the past decades.

It's the same thing as saying a number theorist and a set theorist are in the same field cause they both work in the Math field.

qt31415926··on Statement on AI Risk
Her field has also taken the largest hit from the success of LLMs and her research topics and her department are probably no longer prioritized by research grants. Given how many articles she's written that have criticized LLMs it's not surprising she has incentives.
qt31415926··on Statement on AI Risk
Current topic aside, I feel like that stochastic parrots paper aged really poorly in its criticisms of LLMs, and reading it felt like political propaganda with its exaggerated rhetoric and its anemic amount of scientific substance e.g.

> Text generated by an LM is not grounded in communicative intent, any model of the world, or any model of the reader’s state of mind. It can’t have been, because the training data never included sharing thoughts with a listener, nor does the machine have the ability to do that.

I'm surprised its cited so much given how many of its claims fell flat 1.5 years later

qt31415926··on GPT-4
oh my bad, totally misread that
qt31415926··on GPT-4
It's trained on pre-2021 data. Looks like they tested on the most recent tests (i.e. 2022-2023) or practice exams. But yeah standardized tests are heavily weighed towards pattern matching, which is what GPT-4 is good at, as shown by its failure at the hindsight neglect inverse-scaling problem.
qt31415926··on GPT-4
Curious since it does well on the LSAT, SAT, GRE Verbal.
qt31415926··on Show HN: ChatGPT-i18n – Translate websites' locale json files with AI assistance
For Vietnamese, ChatGPT has been much more effective for me. You can tell it very specific things to modify as well as give it additional context which you can't really do with Google Translate. Especially with pronouns, google will just translate everyone as 'uncle' or 'aunt' or 'grandma' or 'grandpa' and it'll get genders wrong all the time, which you can correct for with ChatGPT.
qt31415926··on I'm not a conspiracy theorist, but what is happening in the SpaceX livestream?
makes sense, so they just looped that segment. thanks!
qt31415926··on Ask HN: How can I do social good through programming?
This is such an amazing post! So resourceful, thank you so much.
qt31415926··on ‘Record IQ is just another talent’ (2010)
Given his upbringing, I would say he never had the potential in the first place. From how he talks about his time at NASA, it seems like he was used as an equation-solving tool 10 years.

“At that time, I led my life like a machine ― I woke up, solved the daily assigned equation, ate, slept, and so forth. I really didn’t know what I was doing, and I was lonely and had no friends,”

His choices now are the result of how NASA as an institution treated him.

Having the highest IQ is great for problem solving. But why should he empathize with the problems of the world when no one did the same for him?