HNHacker News
TopNewBestAskShowJobs

thegeomaster

2,573 karma · joined May 1, 2014

Programmer and tinkerer; based in Belgrade, Serbia. I like building products, systems programming, distributed, and gamedev.

Building Carthagine (AI-powered Figma to React): https://carthagine.ai

submissionscomments
thegeomaster··on Itadakimasu: A word you say to the food, not the cook
Probably has something to do with the fact that the blog post is not human-written.
thegeomaster··on The session you cannot take with you
Pangram is a very reliable tool and it does clock the text as AI generated.
thegeomaster··on The session you cannot take with you
If even Earendil is publishing AI generated blog posts...
thegeomaster··on Launch HN: Superset (YC P26) – IDE for the agents era
Looks great, even has Linux support!
thegeomaster··on A few words on DS4
> `from opentele while import trace`

FYI, this to me points to an inference bug, bad sampling, or a non-native quant. OpenRouter is known to route requests to absolutely terrible, borked implementations. A model like DeepSeek V4 Flash shouldn't be making syntax errors like this.

thegeomaster··on Muse Spark: Scaling towards personal superintelligence
Alexandr Wang on Twitter [0] mentioned open source plans:

"this is step one. bigger models are already in development with infrastructure scaling to match. private api preview open to select partners today, with plans to open-source future versions. incredibly proud of the MSL team. excited for what’s to come!"

https://x.com/alexandr_wang/status/2041909388852748717

thegeomaster··on Project Glasswing: Securing critical software for the AI era
What's the "attention window"? Are you alleging these frontier models use something like SWA? Seems highly unlikely.
thegeomaster··on Qwen3.6-Plus: Towards real world agents
And it seems they've decided to go closed-source for their largest, best models.
thegeomaster··on Launch HN: IonRouter (YC W26) – High-throughput, low-cost inference
Tried on a few of our production prompts and got comparable speeds to what we normally get with Fireworks Serverless (Kimi K2.5), but at a better price. Rooting for you!
thegeomaster··on Show HN: Mowgli – Figma for the agent era, with Claude Code and design export
Thank you so much for the kind words and for the feedback!

1. Duly noted on USDC and other payment options - I have to see how easy this is do to as we're using stripe.

2. Teams and orgs are very high on the priority list and we hope to have something on this front very soon.

To keep up with our development, Discord is probably the best place: https://discord.gg/ptDRKRJpPV

thegeomaster··on Show HN: Mowgli – Figma for the agent era, with Claude Code and design export
Thanks for the feedback! I'm trying to fix that. The trouble is actually that changing the src of an iframe on a page pushes an entry into the history implicitly. Since we use iframes to display the contents of designs, and they can update, this results in a lot of state history pollution. Will prioritize!
thegeomaster··on GLM-5: Targeting complex systems engineering and long-horizon agentic tasks
Not the parent commenter, but in my testing, all recent Claudes (4.5 onward) and the Gemini 3 series have been pretty much flawless in custom tool call formats.
thegeomaster··on Cursor's latest “browser experiment” implied success without evidence
I actually ran this one. It measures some 700k lines of code, and seems to contain things like a full VBA implementation, complex currency and date parsing, etc. But the UI is extremely basic, doesn't seem to expose any of this advanced functionality, and and is buggy to the point of being unusable. Focus will jump around as you type, cells will reset to old values, it will stop responding to keyboard events, etc.
thegeomaster··on The inefficiency of RL, and implications for RLVR progress
Article talks about all of this and references DeepSeek R1 paper[0], section 4.2 (first bullet point on PRM) on why this is much trickier to do than it appears.

[0]: https://arxiv.org/abs/2501.12948

thegeomaster··on The inefficiency of RL, and implications for RLVR progress
You could think of supervised learning as learning against a known ground truth, which pretraining certainly is.
thegeomaster··on Cloudflare Sandbox SDK
It's interesting to also compare this to getting a bare metal instance and provisioning microVMs on it using Firecracker. (Obviously something you shouldn't roll yourself in most cases.)

You can get a bare metal AX162 from Hetzner for 200 EUR/mo, with 48 cores and 128GB of RAM. For 4:1 virtual:physical oversubscription, you could run 192 guests on such a machine, yielding a cost of 200/192 = 1.04 EUR/mo, and giving each guest a bit over 1GiB of RAM. Interestingly, that's not groundbreakingly cheaper than just getting one of Hetzner's virtual machines!

thegeomaster··on Sora 2
You didn't include the amortized cost of a Blackwell GPU, which is an order of magnitude larger expense than electricity.
thegeomaster··on Shai-Hulud malware attack: Tinycolor and over 40 NPM packages compromised
Warning: LLM-generated article, terribly difficult to follow and full of irrelevant details.
thegeomaster··on Irrlicht Engine – a cross-platform realtime 3D engine
Well this was a trip down the memory lane. I built a small game on Irrlicht at the time and I remember these discussions also.

Irrlicht had its editor (irrEdit), a sound system (irrKlang), and some basic collision detection and FPS controller was built right into the engine. This was enough to get you a considerable way through a fully featured tech demo, at the very least. (I even remember Irrlicht including a beautiful first-person tech demo of traversing a large BSP-partitioned castle level.)

However, for those not afraid to stitch these additional parts from other promising libraries (or derive them from first principles, as was fashionable), OGRE offered more raw rendering prowess: a working deferred shading system (this was the heyday of deferred shading), a pop-less terrain implementation with texture splatting, and more impressive shader and rendering pipeline support, with the Cg multi-platform shading language. I remember a fairly impressive ocean surface and Fresnel refraction/reflection demos from OGRE at the time.

thegeomaster··on SkiftOS: A hobby OS built from scratch using C/C++ for ARM, x86, and RISC-V
What an astounding achievement. In 6 years, this person has written not only a very well-designed microkernel, but a build system, UEFI bootloader, graphical shell, UI framework, and a browser engine.

The story of 10x developers among us is not a myth... if anything, it's understated.

thegeomaster··on Are OpenAI and Anthropic losing money on inference?
Common sense:

- The compute requirements would be massive compared to the rest of the industry

- Not a single large open source lab has trained anything over 32B dense in the recent past

- There is considerable crosstalk between researchers at large labs; notice how all of them seem to be going in similar directions all the time. If dense models of this size actually provided benefit compared to MoE, the info would've spread like wildfire.

thegeomaster··on Are OpenAI and Anthropic losing money on inference?
tok/s cannot in any way be used to estimate parameters. It's a tradeoff made at inference time. You can adjust your batch size to serve 1 user at a huge tok/s or many users at a slow tok/s.
thegeomaster··on Are OpenAI and Anthropic losing money on inference?
There's no way Sonnet 4 or Opus 4 are dense models.
thegeomaster··on Are OpenAI and Anthropic losing money on inference?
Are you saying that you think Sonnet 4 has 100B-200B _active_ params? And that Opus has 2T active? What data are you basing these outlandish assumptions on?
thegeomaster··on Show HN: Strix - Open-source AI hackers for your apps
Seems heavily vibe coded, down to the Claude-generated README and a lot of the LLM prompts themselves (which I have found works very poorly compared to human-written prompts). While none of this is necessarily bad, it requires a higher burden of proof that it actually works beyond toy problems [0]. I think everyone would appreciate some examples of vulnerabilities it can find. The missing JWT check showcased in the screenshot would've probably been caught with ordinary AI code review, so to my eye that by itself is not persuasive.

Good luck!

[0]: Why I say this --- a 10kLOC piece of software that was mostly human-written would require a large amount of testing, even manual, to ensure that it works, reliably, at all. All this testing and experimentation would naturally force a certain depth of exploration for the approach, the LLM prompts, etc across a variety of usecases. A mostly AI-written codebase of this size would've required much less testing to get it to "doesn't crash and runs reliably", and so this depth is not a given anymore.

thegeomaster··on Why LLMs can't really build software
Thanks for sharing this! It's difficult to find good examples of useful codebases where coding agents have done most of the work. I'm always actively looking at how I can push these agents to do more for me and it's very instructive to hear from somebody who has had success on this level. (Would be nice to read a writeup, too)
thegeomaster··on Gemma 3 270M: Compact model for hyper-efficient AI
Well, Gemini Flash Lite is at least one, or likely two orders of magnitude larger than this model.
thegeomaster··on Benchmarking GPT-5 on 400 real-world code reviews
Gemini 2.5 Pro is severely kneecapped in this evaluation. Limit of 4096 thinking tokens is way too low; I bet o3 is generating significantly more.
thegeomaster··on GPT-5
It just matches the 90% discount that Claude models have had for quite a while. I don't see anything groundbreaking...
thegeomaster··on GPT-5
o3 pricing: $8/Mtok out

GPT-5 pricing: $10/Mtok out

What am I missing?

Page 1 of 20Next →