HNHacker News
TopNewBestAskShowJobs

shay_ker

712 karma · joined September 21, 2018

you can dm me on twitter, @__smiz
submissionscomments
shay_ker··on Solaris, Our Interface World Models
if you want to try something like this, play MIRA: https://mira-wm.com

notably, rocket league is a closed loop, deterministic game, which is why kyutai labs chose it as a "base case" for world modeling.

it looks cool at first but really quickly you'll start to see the gaps

shay_ker··on Highlighting My Code Based on How Much I Care
i'm actually curious if anyone has explored syntax highlighting with ML. doesn't have to be an LLM, but something that actually highlights what's interesting, or even strange, about a code snippet.
shay_ker··on Migrating to HTTPX2
i assume this is largely about the cybersecurity blitz, to ensure hot-path code is secure?
shay_ker··on Previewing the Model Hardware Standard
How is this different from tool calls? Or rather, MCP?
shay_ker··on WebMCP: Teaching Your Website to Talk to AI Agents
I think pinning WebMCP to the browser was the wrong choice. To list or call tools, you have to have access to the browser environment. "Declarative" WebMCP allows you to read the HTML, but to call tools you still need the JS execution env.

This limits the utility of WebMCP to in-browser agents. In practice, what's the use case? Computer Use can work, but at that point you can circumvent WebMCP entirely.

IMO the better option would be to unify MCP with a standard protocol, like HTTP. Right now, you have to rely on the MCP registry for discovering MCP servers, but wouldn't it make more sense to have this natively in HTTP?

shay_ker··on The turbulent AI era is here
> I doubt I would have put in the same work if I had had an AI companion back then. They talk to you in ways you’re already comfortable with. They don’t push you outside your comfort zone. They are always available and never get mad at you. This gives them the potential to become highly addictive and to rob us of the lessons we learn from connecting with other people.

I've been waiting for a long, long time for us to figure out how AI makes us more social, not less. Outside of memes of course.

shay_ker··on The turbulent AI era is here
> I’m also worried about AI’s impact on education. Ironically, the same tool that will allow people to learn more than ever could also lead to many people learning less.

For teachers/admins who are on HN: what has worked in a past for a "new way to educate" students to actually be implemented by educators? Does it happen through conferences? Do some states take the lead and others follow? Does it follow from private schools into public schools?

I know AI labs are trying to include teachers here and there... but IME teachers are very "anti-AI", "anti-datacenter", etc. etc. In coastal cities, the unions will be entirely opposed. This seems less of a "can AI be useful for education" question and more around politics and messaging. But I'd love to learn more.

shay_ker··on The turbulent AI era is here
> In the United States, I’ve met families who, understandably, were overwhelmed by the process of applying for health insurance, student aid, or food assistance. Faced with a huge stack of complicated bureaucratic forms, many felt like giving up. AI can streamline things dramatically

I'd love to see anyone building Grokbot or ChatGPT Work talk about how they're thinking about this use case, specifically. It seems greatly important and tractable given how little government websites & forms change over time.

On the other end - yes, it will create massive vectors for abuse & fraud, since govt services are designed to have roadblocks to reduce usage. But fraud has always been possible, even without AI (so many recent examples to pull from).

IMO it's best to maximize social benefits to see where the tradeoffs really are. We won't know until we push on it!

shay_ker··on RAG Is Simpler Than You Think
How long have "large scale RAG systems" really existed in the first place? I'm always surprised at this, given how new all this really is, relatively speaking.
shay_ker··on Feature Request: Support AGENTS.md
how does agents.md work for subagents and swarms? is it actually that useful?
shay_ker··on Apple Silicon and macOS VMs: 11–16× Faster LLM Inference with Llama.cpp
I recall there was another YC startup that was working on Mac-specific ML optimizations for local inference (and perhaps fine-tuning).

I wonder if their work is related?

shay_ker··on Rust SIMD on the GPU
Hm is the intent to one day replace the CPU?
shay_ker··on How I use LLMs to learn complex topics
i've asked llms to write a presentation for me on a given topic. for whatever reason, that finds the hidden layer responsible for explaining things well.
shay_ker··on Shopify replaced Redis with MySQL for inventory reservations–and it scaled
outside the slop, i liked this post that was linked on innodb locking: https://jahfer.com/posts/innodb-locks/
shay_ker··on Databricks drove down AI coding spend 70%
how do any of these routing approaches handle kv cache misses? Devin Fusion is the only one that explicitly addresses this, though it does so by switching models during compaction (not sure this isn't still a cache miss though)
shay_ker··on Responding to the next frontier of critical cyber capabilities
i'm surprised there isn't more commentary on the vulnerabilities themselves (mostly in apps developed on the jvm, e.g. artifactory)
shay_ker··on Devtools must be open source
Separately, JetBrains's revenue grew 25%: https://www.jetbrains.com/lp/annualreport-2026/

But back to OP, for this prompt:

> Set up a nightly cron job that executes the prompt: fetch upstream changes to the <software> and rebase all local changes on top of upstream. Check that the software works as intended and replace the current version.

Seems like nice syntax sugar to add a `/maintain-fork` command.

shay_ker··on Agent swarms and the new model economics
How do we know if these models weren’t trained on Turso’s rewrite of SQLite in Rust?

It seems both likely that they were and impossible to remove that code from pretraining. Doesn’t that make this just about LLM memorization of the training set? What am I missing?

shay_ker··on Towards a harness that can do anything
How much do the labs post-train on the harness inputs & outputs? That's a critical piece to understand if a "generic" harness is at all possible
shay_ker··on Profiling the "Abundance" housing bottleneck with real data
There's no free lunch. Housing costs are effectively carried by immigrants and transplants: https://www.aei.org/wp-content/uploads/2023/09/Setting-the-r...

In general, the Vienna model is very difficult to copy without being in extreme conditions: https://www.threads.com/@__smiz/post/Cxpz28CIkT6

shay_ker··on Separating signal from noise in coding evaluations
Didn't we all know from the start that all of SWE-Bench was flawed? Even the authors concede the limitations and have long since moved on.
shay_ker··on Show HN: MIRA – Multiplayer World Model Trained on Rocket League
Great work! My question is about what you mentioned at the end - how well do world models operate when out of distribution? In some sense we hope these models learn something "deeper" about how the world works and can apply that knowledge to different tasks.

I saw lots of awesome ablations in the paper (loved it!), but I'm curious if you analyzed the latents to get an intuition for what the model actually learned. Or, is it just that it learned the training data distribution really, really well?

shay_ker··on Astro 7.0
I saw the integration with Hono - hadn't heard of it before, do many people use it?
shay_ker··on Reducing Doom Loops with Final Token Preference Optimization
I'm curious if this approach can be generalized beyond "doom loops".

For instance, another way of thinking about a "doom loop" is wasted tokens, which happens all the time with larger models that are inefficient at test time. Can "bad-ish" tokens be identified and penalized?

Maybe this is already SOTA but would love to learn more!

shay_ker··on Price per 1M tokens is meaningless
cost per benchmark task is definitely interesting!

i've always wanted cost per prompt, but even that has too much variation.

shay_ker··on GeoLibre 1.0
random thought, but i wonder if html-in-canvas can really improve development for visualization tools specifically
shay_ker··on Looking Forward to Postgres 19: Query Hints
I have a naive question: why did this take 15 years? I understand that good APIs need time and thoughtful design, but I struggle to understand why we couldn’t get to the same (or better) solution faster.
shay_ker··on Angular v22
I do wonder - why not add this to jsx?
shay_ker··on Angular v22
Seems like Angular has gotten better since v2 (my last experience).

Has anyone done a modern Angular vs. React comparison that's not an AI slop article?

I'm also curious if it's "simple made easy" for performant applications. React is arguably "simple made hard", but there are notable, highly performant applications written with it (Linear comes to mind).

shay_ker··on OpenAI frontier models and Codex are now available on AWS
It's fascinating that cloud providers like AWS/GCP/Azure are now immovable "enterprise" technologies, in the way that IBM, Oracle, SAP, etc. were 15 years ago (and still are!).

Fond memories when only startups used S3 and EC2....

It's both an incredible triumph and tremendously sad that cloud providers are now the dinosaurs. So many companies are locked in, just as they were before. It's only going to get worse.

I wish the "cloud" was more fungible.

Page 1 of 9Next →