HNHacker News
TopNewBestAskShowJobs

ponyous

1,417 karma · joined November 30, 2011

Full-stack Developer. Based in London.

vmeznaric@outlook.com

Working on AI 3D modeling software - https://grandpacad.com

X.com/otivdev

submissionscomments
ponyous··on Ask HN: What are people doing with their OpenClaw set ups now?
How did you connect it to search console and do you like the integration? I would like to do this too, and eventually to Google Ads too.
ponyous··on Cartesian – AI 3D Modeling for Design
It really depends on the complexity of the model you are trying to generate and the budget you have.
ponyous··on Steam Frame starts at $1059
You mean people blind in one eye? I think it still works, but obviously depth perception will not be there
ponyous··on Ask HN: What are you working on? (September 2026)
Surprisingly Gemini models. Spatial understanding seems to be the best for the price and speed.
ponyous··on Ask HN: What are you working on? (September 2026)
https://grandpacad.com - AI modeling for 3D printing Been at it for about year and a half.

Really exciting stuff is happening literally every month, because underlying models are getting better and better. When I started it was pretty basic: “make a cube with a hole through it”. Now it’s at the point “make a raspberry pi 4 case” and the agent searches, builds, verifies…

What surprised me in this process is how little meaning AI benchmarks have. Pareto frontier for my use case looks completely different than any other benchmark portrays.

ponyous··on DeepSeek v4.1 Flash
> I have a feeling that the price of 1 million tokens transmitted over the internet is more expensive than cache hit.

And this kinda makes sense. What is cheaper few KB of disk space or internet bandwidth?

ponyous··on Claude Fable 5.1 and Claude Mythos 5.1
We went for 16% intelligence bump according to artificial analysis for +82% of the cost. Interesting.

Comparing 4.8 Opus with Fable 5.1

ponyous··on Gemini 3.7 Flash
On my benchmark where AIs generate ~20 different 3D models about 1/2 the time of Opus and 1/3 of the time of Kimi K3 and 2/3 of time of sonnet.
ponyous··on Ask HN: What are you working on? (August 2026)
Yeah you can absolutely do multiple parts, although the mating features are still a bit rough you can do a lot already.
ponyous··on Ask HN: What are you working on? (August 2026)
GrandpaCAD - AI 3D modeling software focused on simplicity. Made it so even my grandpa could model. He’s been asking me for years when will I teach him how to 3D model. I tried, we failed and then I seen him use ChatGPT so I knew there was a better way than traditional CAD tools.

Recently we also got European funding and the project got some traction. Very exciting times ahead.

https://grandpacad.com

ponyous··on DeepSeek V4 Flash 0731
You are right, relatively to other llm providers this is not slow. But if you think what is possible when you have 1000t/s a sec you might find it slow.
ponyous··on Modern email can be built from borrowed parts
Spark email client supports this. It’s great
ponyous··on Are AI labs pelicanmaxxing?
I agree with the conclusion and am happy to see this blog post, but this killed a bit of credibility for me:

> Using a single LLM judge for scoring. Every score here comes from one model, GPT-5.6 Luna, looking at one image at a time. I didn’t do much alignment and didn’t check how often it agrees with itself on a re-run.

Having used a similar setup (with previous gen LLMs) to evaluate the 3D models that my product[0] generates, it turned out there was no correlation at all. LLM judgments were very much random and I assume judging SVGs is not that far from judging 3D models. I guess I have to re-test this with current gen.

[0]: https://grandpacad.com

ponyous··on Hy3
Interesting way to show off a model last on every benchmark. Not sure any other lab is doing this
ponyous··on Ask HN: Has anyone had success with SBIR grants and what is the process like?
As someone who just had success with equivalent system in EU my recommendation is to get someone who's done it before to do it for you. I hired an agency. They took 8% fee, which is pretty low, usually it's between 10% and 15%.
ponyous··on GLM-5.2 is the new leading open weights model on Artificial Analysis
Absolutely. Running it now, will update this comment in about 30 mins.

Edit: Surprisingly very good results with 3.0 flash with high thinking.

Cost: $0.06

Duration: 3.22 min

Code Errors: 1.3 per attempts (meaning on average it had to retry 1.3 times)

Adherence was on par with 3.5 flash Low thinking

ponyous··on GLM-5.2 is the new leading open weights model on Artificial Analysis
I don't have the eval results live yet, so I cannot share them yet.

I was benchmarking using a soon to be released new version of my AI CAD modeling software[0]. It's basically an agent that has access to tools that can execute build123d scripts, get sculpted models, blender to combine sculpts + parametric models, tools to inspect the model (visually and with code), search datasheets, ...

I tried what you recommend a while ago (asking an AI to evaluate using different angles) and the AI evaluations were extremely bad - barely any correlation to what I scored. Things have gotten better, but I don't trust it enough yet.

Here is how I score adherence (and how AI did as well, but I tried methods where it would just give back a boolean "pass" or not):

    <0.2 → Poor – Misses core intent; largely irrelevant or incorrect.
    <0.4 → Weak – Partially relevant; significant omissions or errors.
    <0.6 → Fair – Covers main points but lacks completeness or precision.
    <0.8 → Good – Mostly accurate; minor gaps or deviations.
    <=1.0 → Excellent – Fully aligned; precise, comprehensive, and faithful to intent.
Here is the scenario list (prompts are much more detailed):

    dragon-bottle-stopper
    editing-param-mid-conv
    editing-parametric-enclosure
    editing-swap-material-param
    editing-text-edit-cube
    multi-turn-bird-house
    multi-turn-dice-tower
    multi-turn-modular-planter
    multi-turn-phone-stand
    multi-turn-shelf
    one-shot-bookend
    one-shot-cable-clip
    one-shot-chess-queen
    one-shot-coaster
    one-shot-coffee-cup
    one-shot-dog-tag
    one-shot-dragon-figurine
    one-shot-hex-bracket
    one-shot-keychain-fob
    one-shot-low-poly-tree
    one-shot-pegboard-hook
    one-shot-pi4-case
    one-shot-threaded-jar


[0]: https://grandpacad.com
ponyous··on GLM-5.2 is the new leading open weights model on Artificial Analysis
Just ran and scored 63 3d model generations (via code) across high and no reasoning. 3D Modeling benchmark quickly shows spatial, logic and code performance of the model so I think it's a very good indicator of the quality.

Here are the results compared to Gemini 3.5 Flash:

    Model + config          CodeErr/gen   Cost/gen   Median time   Quality
    gemini-3.5-flash, low      0.71        $0.18        68s       baseline
    GLM 5.2, reasoning high    0.61        $0.18       289s         -6.0%
    GLM 5.2, reasoning off     1.52        $0.10       126s        -13.6%

Although it is cheaper, it is significantly slower, and results are worse overall. Surprisingly - high reasoning produces less code errors than gemini 3.5 flash, but when I actually look at the models they are worse.

Edit: I recently ran evals with Kimi 2.7 and MiniMax-M3 and this is clearly open source SOTA model, by far.

ponyous··on Claude Fable 5
Basically double from Opus 4.8 IIRC
ponyous··on Spanish traders set the standard for GnuCash database design
> Surprisingly written by a human :)

Article ends with this

ponyous··on Bun's unreleased Rust port has 13,365 unsafe blocks
Bun is(was?) a lot about performance. How does it compare to zig?
ponyous··on Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark
I've run a tons of benchmarks for OpenSCAD for all kinds of models and setups, and what I realised is:

- Models are very jagged (might excel in one type of 3d model, but not another)

- Gemini models are the least jagged in my experience and have the best image understanding

- Gemini models are also the most creative (which may be undesirable if you want precise CAD part)

- Overall this benchmark doesn't prove much because one 3d model (and one attempt) is just not enough. I am usually testing on at least a dozen models each generated 3 times, but should really do much more, but it's too pricey for a solo dev.

Still, thanks for publishing this. Will be definitely run flash 3.5 soon to see how it performs.

ponyous··on GenCAD
I've seen this and other attempts like this[0] while exploring improvements for my CAD AI[1]. And I think these are potentially powerful solutions, but none of the current projects/weights have enough training (data or time) to make it work on arbitrary models. MeshCoder pretty much works only on models based of their training data. I haven't tried GenCAD but other commenters have confirmed my suspicion.

[0]: https://daibingquan.github.io/MeshCoder/

[1]: https://grandpacad.com

ponyous··on Arena AI Model ELO History
Seems like Chinese labs are the only ones that are trustworthy (at least when it gets to this specific issue). This feels so ironic haha
ponyous··on Google Chrome silently installs a 4 GB AI model on your device without consent
The site is currently unavailable 503 so I can't read it. But I wonder, what should you consent to? Every dependency? Every dependency above 1GB?
ponyous··on Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge
Kimi is nowhere near GPT or Opus unfortunately. I really wish it was. I’m running evals where models have to generate code that produces 3D models and it’s obvious that it lacks spatial understanding and makes many more code errors before it succeeds.

Maybe it’s better in one particular case here and there and I think this blog post is example of that.

ponyous··on Show HN: AI CAD Harness
Gregor! Not sure I want to say more on here.
ponyous··on Show HN: AI CAD Harness
Been following you guys a while, seems like you've been gaining some traction recently, lets goo and congrats!

I have been working on GrandpaCAD[0] for a while, a very similar product. I thought of you as my biggest competitors but noticed recently you are focusing more and more on professionals while I am focusing on total noobs in modeling who just want to whip out a quick model. So I guess we are not competitors anymore?

My evals[1] show that Opus 4.7 and GPT 5.5 are very comparable in terms of generation quality, but GPT 5.5 is slower and costs sooo much more in my harness. And the original breakthrough model was Gemini 3.1. I'm curious do you have more written about your benchmarks setup?

If you want to chat email is in my profile. Btw, just met "your"(?) neighbour on a plane a couple of days ago. World is small.

[0]: https://grandpacad.com

[1]: https://grandpacad.com/en/blog/public-benchmarks-misled-me-o...

ponyous··on CadQuery is an open-source Python library for building 3D CAD models
Another library I have to integrate and benchmark against OpenSCAD for my AI SaaS[0]. I am really curious how constructive solid geometry compares to sketching and extruding that CadQuery is build on.

Anyone curious in the writeup? I have a pretty good harness for evaluating 3d generation performance.

[0]: https://grandpacad.com

ponyous··on Darkbloom – Private inference on idle Macs
Why does M1 Max project significantly higher revenue than M3 Max with double the ram?
Page 1 of 17Next →