HNHacker News
TopNewBestAskShowJobs

jasonjmcghee

4,817 karma · joined March 25, 2015

Website: https://jason.today

GitHub: https://github.com/jasonjmcghee

submissionscomments
jasonjmcghee··on Gemini 4 Argon
Gemini 1.5 Pro claimed 10M input tokens before release.

And was 2M tokens IIRC after release.

There were also many rumors that Gemini 4 was going back to 2M. Just seems odd not to say what it is.

jasonjmcghee··on Gemini 4 Argon
> 1M output token limit

what about input?

(Maybe I missed it)

jasonjmcghee··on Everybody’s home. No one’s coming over
Lots of focus on the negatives, but it's rather uplifting to see that fathers are spending 4x as much time with their kids.
jasonjmcghee··on Ollaya – Ollama for open-source, Jev-style decision models
In my experience it's not close and the benchmarks I've seen don't reflect my experience at all.

But I'm guessing people will find the right training regime and data mix soon to close the gap.

But big things I see are instability and inaccuracy - like pick a random problem.

jasonjmcghee··on OpenAI is well positioned to fast-follow Jev
They are not equal or better. Try them out on real things. I did - it's not close. The benchmarks are a misrepresentation.

That being said- hard to believe OpenAI etc couldn't build a frontier classifier too.

jasonjmcghee··on Transformers Explained Visually
For the uninitiated, I can't recommend enough, The Illustrated Transformer:

https://jalammar.github.io/illustrated-transformer/

jasonjmcghee··on Grok 4.7
For what it's worth - over the last few years or whatever, it seems like Anthropic benchmaxxes the least.

That being said, I currently prefer Sol / Astra to Opus / Fable as I find both to be a better cost payoff to me.

jasonjmcghee··on Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
I think the reason is in the general ballpark of people throwing LLMs at a huge variety of problems and being too slow disqualifies them from a bunch of things.

Now there's a new training-free thing that is fast enough to be useful on a new class of problem.

If you have a little data and can ask a frontier LLM to train a model, you can probably beat it on average for a specific task.

But... This is the case with LLMs too.

jasonjmcghee··on Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
TypeSafe is still the provider. It's just proxied.

So it's a new subprocessor. Which can often be painful to onboard, especially if not compliant according to your needs.

jasonjmcghee··on Kev: Tiny Jev-like family of decision models built on top of Qwen3.5
The open source ones- I downloaded a number and tried them and compared to Jev.

Anything that required knowledge / familiarity mmBERT and ModernBERT post-trains performed much worse.

So it seems like they did some kind of useful expansive pre-training.

Things that were Qwen or Gemma Diffusion did better at those kinds of tasks but were generally pretty inconsistent in terms of whether they could succeed repeatedly (and be stable + reliable) on the many types of tasks that are in the cookbook part of the Jev docs.

If you ask Jev similar input + questions, it's pretty stable. And does a reasonable job on a lot of questions.

This one public benchmark (the only I've seen) seems to give the open versions way too much credit. It wasn't my experience at all.

It gave my a false wrong sense of what might be required to get it working for something at work to avoid needing a new subprocessor as - at least on Cloudflare / OpenRouter Jev is third-party not hosted.

jasonjmcghee··on Typesafe-computer-use drives a Mac toward a goal for 1/50th of a cent per step
Did you do the follow up questions? I was invited within a few hours of joining today.

It's also now on openrouter and cloudflare

jasonjmcghee··on When the fractional part of a float fixes your shader
There absolutely is a "main" function in GLSL. This is the entry point.

And there are for loops- and if statements, and function declarations, etc. It's very c-like without dynamic memory allocation and executed in a highly parallel context.

Your shader is also not necessarily called every frame, you control this. It could be called multiple times per frame or only at certain times.

Fragment shaders (one type of shader) is called for each fragment of the rendered geometry - in 2D or in screen space this is just a quad so effectively a panel of pixels.

You get to know your own location with something like UV or a pixel coordinate if you pass it from a vertex shader, so that you can sample input textures as needed to build whatever it is you're building.

If you're not doing multiple render targets, your only responsibility is setting to the fragment color variable which will be the final color of that pixel for the output texture.

And that could in turn be the input to another shader etc.

I usually recommend https://thebookofshaders.com/

It's very sad this was never completed. It's so good.

jasonjmcghee··on Show HN: How Stale Is Your AI? Release age and training cutoff for 20 models
ollama is just an inference engine - it just runs models.

it must ship with some default old model if you didn't need to explicitly download one

jasonjmcghee··on Show HN: How Stale Is Your AI? Release age and training cutoff for 20 models
Or other kind of search / knowledge acquisition / computer use etc to get the information needed
jasonjmcghee··on Show HN: How Stale Is Your AI? Release age and training cutoff for 20 models
It still matters, but in the age of good reasoning, tool use, and web search, this is much less of a problem than it used to be.
jasonjmcghee··on I'm not addicted to the internet or my smartphone. I'm addicted to information
Just a small note and I'm guessing this is an intentional aesthetic choice but the chromatic aberration makes the text very uncomfortable to read.
jasonjmcghee··on Show HN: Kinesis – Control your Mac with the Meta Neural Band
Feels like such a misnomer, but I guess that's marketing.

I remember when OCZ NIA came out and it claimed BCI and it was effectively doing muscle activity too.

At least meta is properly representing what it does!

jasonjmcghee··on Spaceships (Reverse Asteroid)
Fun twist on the game!

I didn't play for too long, but I just clicked between the ship and asteroids and near-instantly destroyed it, quickly passing the waves I played.

Is there difficulty ramp later on?

jasonjmcghee··on google.com/goto: Google's anti-scraping update
I've never heard of this particular SERP provider, but some marketer is very excited they wrote this blog post right now (500+ upvotes on an seo blog post).

And their fix here - they just resolve all the urls- which I suppose could make the service slightly more expensive? But otherwise isn't that what every similar provider/ crawler etc will do and this change will only hurt users?

jasonjmcghee··on Show HN: Toast, a by default in-terminal IDE
First file I look at (concerning):

"github.com/yourusername/toast/internal/components/breadcrumbs" "github.com/yourusername/toast/internal/components/closedialog" "github.com/yourusername/toast/internal/components/commandpalette" "github.com/yourusername/toast/internal/components/editor"

jasonjmcghee··on RTK reports token savings, but our cost benchmarks disagree
I see the same thing and have effectively the same philosophy. If I'm using something like figma or glean or playwright/chrome dev tools, plugin/skill/mcp - likely very useful.

But so many of the weird collections of skills that people on YouTube get viral followings for - I just don't get it.

People excitedly ask me what skills I use and I feel bad just saying only things we've directly authored for some express purpose. None of the "hot" ones.

I've written a large handful of skills, but they aren't like vim plugins. I don't just leave them "on".

This has been my experience at least- curious if I'm just behind the times.

I also effectively didn't leave the IDE+ChatGPT copy/paste workflow until the first release of Claude code. So maybe I'm slow to adopt.

jasonjmcghee··on Herdr Studio
Using their name, logo, and css styles - community project.

Feels wildly misleading.

jasonjmcghee··on DeepSeek v4.1 Flash
Actually... https://x.com/antirez/status/2098121665771110540
jasonjmcghee··on Tell HN: OpenAI keeps re-enabling the 'allow training' setting
I've done the formal opt-out process - like "Make a privacy request" where you fill out a form.

Not sure if that's region specific or something though.

jasonjmcghee··on No Man's Sky Cosmos
The original massive controversy was the founder of the studio claiming it had multiplayer when it very clearly didn't which was such an odd thing to claim and easy to prove false.

That and the trailer was very different than the released game.

They apologized IIRC and then continued to work on it for like a decade - adding multiplayer and many large changes to the game

jasonjmcghee··on No Man's Sky Cosmos
A big name in proc gen is Kate Compton who coined the term "10,000 bowls of oatmeal" to describe this phenomenon of technically unique but not different enough to matter - I googled "10000 bowls of oatmeal" to see if there was a good article and the second link was about it and nms

https://www.challies.com/articles/no-mans-sky-and-10000-bowl...

jasonjmcghee··on Extracting Steering Vectors from J space
Steering can degrade and bias output

Even with basic experiments I've done, it frequently introduces much more hallucination etc and allowing arbitrary steering...

Not to mention you just can't trust a model's judgement if the highest bidder chooses what it thinks

jasonjmcghee··on Extracting Steering Vectors from J space
Fwiw you can just Google this for a model and often someone has done it

https://huggingface.co/eyes-ml/Qwen3.8-27B_jacobian-lens

jasonjmcghee··on Three sites made 215,128 “best software” pages for AI. Perplexity cites them
If by root urls you mean domains, openai at least supports this.

https://developers.openai.com/api/docs/guides/tools-web-sear...

jasonjmcghee··on Building Autonomous Goal Loops That Deliver
Sorry for the non-substantive comment, but I found the animated line at the top very distracting.
Page 1 of 34Next →