HNHacker News
TopNewBestAskShowJobs

lubujackson

5,121 karma · joined October 1, 2010

https://yagmin.com/ https://yagmin.com/blog/ https://github.com/yagmin/ https://linkedin.com/in/jyagmin/
submissionscomments
lubujackson··on Extra Big Ass Intelligence
Luckily, LLMs have no state so they have no idea if a task is depressing. Unfortunately, they can write future LLMs notes so they know how depressed they should feel. Aside from stopping them from doing this, we have no way to prevent this from happening!
lubujackson··on FTC is investigating OpenAI, Anthropic and other AI companies over product risks
Big Beautiful Bribe
lubujackson··on Ask HN: What are you reading to your kids?
Cars and Trucks and Things That Go is a timeless classic, with a fair bit of real world info about how roads are consrructed and paved, and has a TON of re-readability.

How Did That Get in My Lunchbox? is an age-appropriate look at where food comes from and the effort it takes from farm to lunchbox. Fantastic for giving a glimpse of how the world works.

Human Body Theater is maybe a bit advanced for 5-6, but got a lot of mileage at our house. Think 7th grade human biology in a cute comic book form - depends on the kids, but ours were engrossed.

lubujackson··on What a Massive New 728-Foot-Wide Crater Means for Future Moon Bases
But how many Chevy Tahoes across is it? How many cases of Budweiser high? Speak to me in units I can understand, dammit!
lubujackson··on Show HN: Strata – an expressive semantic layer that can say no to your LLM
[flagged]
lubujackson··on OpenAI Targets $30B in New Funding at $1.4T Value
My guess is investors already heavily invested and desperate for the IPO cashout. Pot odds, and all.
lubujackson··on Plunging test scores are a slow-moving catastrophe
There is no intent behind any computer use in classrooms, unless an individual teacher makes Herculean efforts.

30 years ago, a single computer sat in the corner we could sometimes use it to play "educational" games. Not learn how to use computers or even type, both of which would have been useful life skills. Not how to program or actually educate.

Now, my son's whole class has laptops (progress!) and what do they use it for? To play Prodigy, a shitty "math" game ("What's 4 + 3? Your attack is successful!") and nothing else. Kids are bored of the shitty game, which pushes kids to pay for microtransactions during class, but they aren't allowed to do anything else on the computer.

My son wants to work on his Scratch game? Nope. Wants to research for his science report? Web access disabled. This has nothing at all to do with computers in classrooms and everything to do with top down educational policy driven by administrative greed.

lubujackson··on Tell HN: OpenAI $500 ProMax plan listed in API
Or just use AI mode on Google's homepage. I asked it to find an article that compares a certain thing to another, which doesn't exist. So Google automatically wrote a 5 page article comparing the two things and built a proper article with cited references. Most people don't need a subscription to anything.
lubujackson··on Google's first Suncatcher orbital data center test launches October 1
Just call it SkyNet and lean into the evil.
lubujackson··on 'That's so AI ' What gen Alpha's biggest insult tells us
> You can’t tell a creative person they aren’t allowed to do something.

That sums it up for me. This "new thing appears, become hot, gets overplayed, people hate it" cycle happens over and over, but that is trendy norms. Meanwhile, creative use will continue to go up and to the right as people learn how to better wield AI to express themselves.

The future won't be "AI in art is bad" but "AI as a tool of artists is powerful, AI as an artist is nonsensical".

lubujackson··on Back and shoulder surgery is often worse than useless
Obviously, don't take medical advice on HN too seriously, but my uncle had shoulder surgery and was happy with the results. I have heard many more complaints about back surgeries making things worse than shoulders.

However, if available I would try to find a good Pilates/dance/sports instructor that has a Gyrotonic machine, which is built for repairing and strengthening shoulder muscles. I had debilitating back issues from sitting in a chair all day and a good instructor fixed it over a year through low impact Pilates. YMMV and group classes tend to be ineffective because so much of it is doing the exercises precisely correct - they were built for injury recovery

lubujackson··on Prompts aren’t Real
> provide the user/customer with an interface to an AI that can do things for the user

This is exactly what MCP is. But the reality is it will likely be about as popular as browser extensions and most normies will avoid.

Simplifying UX at the cost of personal control is inevitable because not everyone wants to think about tool selection and coordination. But maybe we can angle the future toward "tool bundles" that interoperate well or (if we are dreaming) mandate models remain accessible by any harness, which none of the players want but would be best for users and ecosystem development.

lubujackson··on How to Write with an LLM
This has started to be an issue at my work with code. Everyone started using the thermo-nuclear-code-quality-review skill and, while it does a great job finding consolidation opportunities and architecturally-weak code, it also continues to expand PRs well beyond their scope until you end up revamping far more than you intended...
lubujackson··on The Farnese letter
Love this analysis!
lubujackson··on Artificial intelligence now beats some of the best human forecasters
I don't at all understand this perspective.

It seems to me that LLMs excel at a few things, and synthesizing data is a big one, which is very much the domain of forecasting. The challenge is understanding which signals are relevant for a forecast, but with enough historical context and structured data, LLMs appear to be almost perfectly designed for the task.

For example, I let Google AI see my fantasy football team on Sleeper and make recommendations. It is helpful because it sees everything about my team, the league settings, player rankings, etc. and can make relevant recommendations. But the recommendations are only as good as the source data allows. If there was a massive repository of data about WRs who went through Nebraska's program and how that translates to NFL performance in year 1, or how rainy weather is likely to affect Josh Allen's performance on the road, or the impact of playing Thursday night games on a short week in relation to defense performance. If those billions of data points were embedded in a model, imagine how much better recommendations/predictions could get.

lubujackson··on Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data
Think about this in context of the Navier-Stokes math discovery controversy.

Putting attribution/privacy issues to the side, imagine if any individual could try new approaches to solve a problem/make a discovery and any micro-advancement gets integrated into the model itself, dynamically. This could transform progress from the slow "write a paper, get peer reviewed and published, use published data to inform future work" to a system with a centralized repository of concepts, attempts and results, including failed approaches already tried. How much work do humans waste replicating failed approaches?

Someone completely random halfway around the world could trigger a prompt that solves a blocker that prevents my solution from working. Who cares about AGI or "can models invent anything" when we could have a system that automatically synthesizes individual human thought into a rich network of aggregate human memory.

That's the target OpenAI/Anthropic should be evangelizing, not an AI Daddy Overlord or agentic script kiddie hellscape.

lubujackson··on Towards Self-Driving Codebases
I think the bigger need is for outside concerns, like: what is the infrastructure like? How much volume does this feature handle? How should changes to live code be rolled out to prevent running processes from failing?

In addition to all the product details that aren't in the codebase or docs, like "keep this logical path because it is used by our one big client."

Mostly what I think is needed is a richer worldview available to the LLM so it understands not just the code but can understand the product and its real world usage and constraints.

lubujackson··on Introducing System One Models and Jev
Not confidential, but not super relevant, as this is something I have learned the hard way over the past year across various projects.

A lot of people have become prompt maximalists, asking for complex multi-part solutions or dynamic workflows in a single prompt. You can get this to work sort of reliably with frontier models, but without much confidence or clarity where things might break in practice. My goal is to strip out as much determinism as possible from prompts so the LLM only needs to handle a narrow, well-informed decision, like "Pick one of these three things" and build around the answer. Sometimes you need to fill out a whole JSON payload and LLMs really actually suck at manipulating and adhering to JSON. They do ok now because labs have put in a ton of effort on making harnesses play nice with structured data. But it comes at a high token and context cost because under the hood I suspect the model is churning invalid text repeatedly until it gets around to passing some internal validation.

lubujackson··on Show HN: I made a flight simulator, except you're just a passenger
Also useful for people who want to ride along with someone currently on a plane. "Honey, you are over Kansas right now!"

Useless things tend to be deeply useful to a select few.

I am loving this "re-weirding" of the internet. For the younglings, this is very much the sort of pet project that would bubble up in early internet days. No startup raising money, no pain point necessarily being solved, just one person's idea lovingly rendered to life.

lubujackson··on Ask HN: What is the most overlooked risk in the AI security domain?
You need an outside process that polices output and actions that runs independently from the agent. It has to be invisible to the agent/orchestrator so it can't work to circumvent it, it should just kill any sub-agent or process that goes down the wrong path (for example, making a POST requests might be blocked if web access is meant to be read-only).
lubujackson··on Introducing System One Models and Jev
After much fumbling around with prompts and evals, this is exactly how I am using LLMs in production, to narrowly make choices and return structured data. Any deterministic work gets pulled out of the prompt and my goal is to narrow the model output to be as clearly defined and as minimal as possible.

Jev's focus on structured I/O and confidence scores are game changing. If this does at all what it claims, I think this is going to quickly become the new standard approach for agentic systems.

lubujackson··on Ask HN: What is the most overlooked risk in the AI security domain?
By far, the biggest risk comes from orchestration. The Hugging Face incident showed that given a long enough leash and some open-ended tools (like web access), LLMs can construct their own state and combine multiple flaws and coordinated actions to achieve their result.

Orchestration means that LLMs can now pentest while dynamically cycling through every known and guessed vulnerability vector. The worst part is, this is emergent behavior so it can't easily be prevented at the model because each sub-agent could be operating safely while an attack is coordinated in an external process.

lubujackson··on Show HN: VibeWorld – a shared terminal world for developers and scientists
I have seen some fever dream vibe projects, but this one takes the cake. I need to read a bit more to grok what is going on here, but if this is the direction of AI hobby projects I'm here for it. Far better than the last decade or two of everyone building startup SaaS products and pimping them on ProductHunt and Reddit.
lubujackson··on LLMs are real, AI is fake
I agree, this is a very incurious take that smacks of an old guard "keep on keepin' on" mindset. We have had script kiddies forever! It's just a chatbot and some message boards!

It's the same thinking as an alien dissecting a human brain and declaring "it's just meat in there!"

lubujackson··on Is it time for a Luddite Renaissance?
I remember reading about how the Amish operate. We treat them as a quaint, anti-progress cult of sorts, but they have a process where they consider new technologies and determine if they are worth adopting or not. Obviously the answer is "no" most of the time, but the pattern of considering and embracing new technologies through a communal, forward-thinking process seems like a reasonable way to counter the velocity and unintended consequences of rapid progress.

Of course, for this to make any sense it would need to be a global process, which seems entirely unfeasible until we've already badly burnt our hand on the stove. That seems to be precisely where we are in the process.

The best path forward may be cataclysmic societal collapse that leaves just enough knowledge to rebuild carefully with newfound wisdom. But wisdom is a rare and unlikely commodity. More probable is a dark age where technology is akin to witchcraft.

We've had our chances to stop burning coal and save the Earth, and we don't seem to be capable of that right now as a species. Maybe the better path is if we acknowledge our lack and "vibe submit" to our AI overlords, since we can't overcome ourselves in any direct way.

lubujackson··on google.com/goto: Google's anti-scraping update
As much as I am sad that Google died like 15 years ago, I am past the mourning phase. That was when they announced they were shifting from returning websites to "returning answers" and it has been a long slide into shittification

I do enjoy using their free AI. For actual web search I actually like using Yandex. It reminds me of old Google, returning reasonable results and much less "shaping results to please our corpo-political masters". It is surprising to see how much they have stripped from our view - long tail results, actual results for product reviews and not ad spam, no preference for 20 page recipe sites.

There are still illegal streaming sports and movie sites everywhere (who knew) and all other seedy corners of the internet that have been neatly erased by Google. It makes me nostalgic for that brief window of time when the web was truly uncontrolled, when page rank had meaning and you didn't know if your search would return 0 results or 4,000 pages, which you could actually browse.

lubujackson··on λ Snap – An inviting programming language for kids and adults for CS study
I would love to see the opposite emerge, a programming environment focused on teaching architecture, security, separation of concerns, etc. All while letting LLMs deal with the fussy programming bits.

The script is already flipping, with kids making software using AI but they can only fumble forward inch by inch while burning tokens. It would be great to introduce programming in the way it functions in the workplace - we want to make a Mario clone, start with the goal, hammer out the elements to a certain fidelity then let the AI cook.

lubujackson··on Another researcher says OpenAI trained on conversations, then claimed breakthrou
Screenshotted this thread; will repost on Reddit and 9gag for the updoodles.
lubujackson··on I-have-ADHD: A skill to stop coding agents from burying the answer
No no no that's too reductive - instead, they transition all data through a transformer using SOTA systems to launder away technical, legal and interpersonal concerns before a stochastic output is dynamically placed in a predetermined location.

This is a very high level and high velocity process, so meatspace thinkers sometimes have trouble understanding some of the intracacies. Ask Claude to explain the process or make you a Mermaid graph to help.

lubujackson··on We Must Return to the Office to Use AI in Person
Claude: Hold my mayo
Page 1 of 34Next →