HNHacker News
TopNewBestAskShowJobs

astrobiased

630 karma · joined November 15, 2012

2x11.xyz
submissionscomments
astrobiased··on GPT-6 Astra
Appreciate the input! Responses below to your first two items:

1) I'm leaning more into a broader sense, which is that given the priors a system already possesses, how efficient can it acquire competence on a novel task? If I'm reading the point you make, you're focussing on continual learning right? If so, I'm not necessarily restricting my statement above to that.

Here's another reframing: How much of the benchmark improvements come from overwhelmingly large training distributions vs improving the models for adapting to things genuinely outside of it?

2) Excellent points about crystalized and fluid intelligence. Wouldn't the LLM scaling gains be a representation of crystallized capabilities? In regard to Gf, that is exactly what I am asking about. That is what seems to be lacking, Gf like adaptation under genuine novelty.

My concern is that it's increasingly difficult to tell of what looks like Gf like behavior is really coming from better adaption vs. having broad priors from the model's large learned distributions.

astrobiased··on GPT-6 Astra
I think the moat that China has is energy costs. It's taking learnings from the Bitter Lesson. If you role up scale and compute to the next level, it's energy resources. China has it and sharing open weight models is an effective means of removing the tech moat. This idea has been floating around for a bit now (I'm not taking credit for it).
astrobiased··on GPT-6 Astra
I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547

Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area.

It seems more about coverage-driven competence. Somewhat analogous to overfitting at scale.

The harder question, in Chollet’s framing, is: how efficiently can a system learn to do something genuinely new?

With our current AI architectures and training in place, I think we will only continue on skill acquisition optimization vs. truly novel intelligence.

astrobiased··on Nvidia Nemotron 3.5 Lightning and NeMo Switchyard
Depends on the type of agentic task though. For simple operations, a small model can be quite beneficial.
astrobiased··on Prime Agent: A self-improving RLM agent
Not RL. SFT.
astrobiased··on Prime Agent: A self-improving RLM agent
Yes, used bert model with decent results.
astrobiased··on Pi's Minimalism Is Its Advantage
Pi does one thing that I love, developing a tool that has minimalism where it's easily configurable with good documentation. The leads to new use cases that the the author(s) would have never dreamed of. The organic growth process of the Pi ecosystem has been fascinating to observe. It's one of the reasons why Pi has become one of my favorite coding agents to this day, flexible beyond personal uses and extensible to larger environments.

IMO, I view it more than a coding agent, it's a coding agent platform with powerful extensibility.

astrobiased··on Show HN: Cactus Hybrid: We taught Gemma 4 to know when it's wrong
Is this in any way similar to Goodfire's work? https://www.goodfire.ai/research/rlfr#
astrobiased··on Who's afraid of Chinese models?
The post hits it spot on with unequal access to the models in terms of security. I'm developing OSS where security is important for the user ... but the frontier models like GPT 5.6 and Fable flake out and state that I cannot get the info/access.

This is extremely lopsided I'll have to resort to GLM 5.2/K3 to ensure that those security issues (hopefully) are resolved properly.

For OSS, this is one of the most counterintuitive experiences I have ever had. More than ever I'm convinced that open weight and open pipelines models are 100% critical for progress on the AI and societal fronts.

astrobiased··on Real-Time LuaTeX: Recompiling Large Documents in 1ms [pdf]
Same, so much in fact that I have used it for my website now and it works beautifully for graphics and formulas.
astrobiased··on Agents need control flow, not more prompts
It's the right direction, but control flow introduces limitations within a system that is quite adaptable to dynamic situations. The more control flow you try to do, the more buggy edge cases that pop up if done poorly.

Still have yet to see a universal treatment that tackles this well.

astrobiased··on What is an elliptic curve? (2019)
John is not only smart and knowledgeable, but an incredibly great person to know in general. I worked with him on a project briefly back in 2012 and he stood out as a champion for science, coding, and education. His posts clearly reflect him well.
astrobiased··on Toucan Wireless Split Keyboard with Touchpad
Absolutely love these type of keyboards. But ... with how much security I work with for logins, etc, the fingerprint button on my Mac keyboards are amazing time savers that I don't want to live without. Has anyone found a workaround?
astrobiased··on 19% of California houses are owned by investors
What I find odd is that definition of "investor" is not that clear. When you click through the links you get blocked at the data provider with no context. There's also a link to another post by the same news provider. When clicking through reference to the data source, the link doesn't work.
astrobiased··on Define policy forbidding use of AI code generators
It would need to be more than that. A prompt for one model can have different results vs another. Even when the model has different treatment for inference, eg quantization, the same prompt for the unquantized and quantized model could differ.
astrobiased··on GravityLight: Generate light with gravity
Regarding the energy of lifting a 10KG weight 2 KM high is not a 1:1 comparison with the tech here. It's using an LED which much more efficient than a Kerosene powered flame, since most of the energy in the flame is spent in infrared.
astrobiased··on Pyxley: Python Powered Dashboards
Fixed! Thanks for the catch.
astrobiased··on Seaborn: a high-level Python interface for drawing statistical graphics
Seaborn is my favorite statistical plotting package in Python. I wrote an astro plotting package that digs deep into the Matplotlib internals and it was not easy. Big props to the developer behind Seaborn and the great aesthetics he imbued it with.
astrobiased··on Partially Derivative: The One About Sarah Connor
The podcast is as entertaining as it is informative. It's really a goldmine of good resources/ideas and it's my favorite podcast to date.
astrobiased··on Hubble captures the sharpest ever view of Andromeda Galaxy
You're correct, the likelihood of stars colliding is near zero, but the gas that forms stars in both galaxies will collide and that will create quite a spectacular view. It will likely resemble something like this: http://hugepic.io/bfc195a2b/4.00/2.02/-77.61
astrobiased··on Show HN: Discover Awesome New Repositories on Github
Great finds! This is the kind of stuff that Inspector Git aims to do. I'm honestly surprised how well it did with the Hadoop recommendations.
astrobiased··on Egyptian statue spins all by itself at Manchester Museum
Your question rang with mine. Why is this on HN? Is there a policy about the content that posts should contain on HN? In the past 6 months it seems like much more non-code/hack/start up material is posted on HN and that's a bit concerning.
astrobiased··on In Head-Hunting, Big Data May Not Be Such a Big Deal
My dad was part of a company in Boston that only had employees with an IQ of 140 and up. I asked him how it went and he laughed and said it naturally fell into ruins. Key point: there's a lot more to employees than just quantifiable numbers like GPA, IQ and such.
astrobiased··on Forget Your GPA
This depends on which field of study you're in. If you want to go to grad school for astrophysics, getting that high GPA score is important. But, I do agree that halving your time to get a B+ versus an A is good. This type of decision making shows that you're good at managing your time and setting your priorities right.
astrobiased··on Seattle drinking den bans Google Glass geeks
After reading Amped from Daniel H. Wilson, this post seems to strike the same chord as the book. Almost freaky.
astrobiased··on Authorea: Write research papers inside your browser
[Authorea consultant here] This feature is quite exciting. Using JS based plotting capabilities like D3.js will allow users to do two things at once.

(1) Provide dynamic figures to represent data that is normally ineffective in static form and (2) provide the code and data that was used to create it. This helps others reproduce the results from the data to figure form, which is a feature that is definitely missing in PDF publications.

astrobiased··on 100,000 stars
Awesome! The only issue is that the visualization makes it look like we are in a cluster of stars. That is not correct. We're part of the diffuse field star population in the Milky Way.