HNHacker News
TopNewBestAskShowJobs

tezza

1,853 karma · joined April 7, 2009

PM: hackernews (at) terrylurie DOT com

In spare time I make:

- https://epicwin.team - browser video games. Remote Teams and Solo

- https://generative-ai.review - Reviewing AI tools from a professional angle

- LLM London organiser

Based in London. Light inventor on the side of a regular day job

Any opinions expressed are my own and not those of my employer.

submissionscomments
tezza··on ChatGPT Images 2.5
I've done a side by side comparison of the two new models across all quality levels. Versus all the previous models

https://generative-ai.review/2026/09/rush-openai-image-gen-2...

right, off to bed now

tezza··on Clamiga: Common Lisp for the Amiga
Clamiga should be treated immediately with a large dose of Clatari S.T.
tezza··on Kill The Cookie Banner
Even well meaning bodies like TFL (Transport for London) have cookie warnings that impede the actual access of the website.

Need to look up a bus time? Full screen cookie consent with accept buttons drawn OFF THE SCREEN.

tezza··on It's getting harder to focus every day
I would add that Brazil (1985) foretold this 41 years ago.

The civil servants are addicted to their personal TV and only pause watching it compulsively when the boss comes around. As soon as the boss is out of sight they resume their fixation.

From memory it is a loop of a cowboy galloping on a horse.

tezza··on Kimi K3, and what we can still learn from the pelican benchmark
Yes, I see your point.

Your pelican output is thus both in the training set and yet still outside the capability of the model architecture.

And so you are tracking both the capability of the training and also the capability of the querying!

When you receive your first outstanding pelican it will track a gain of capability.

(btw I first mentioned simonw-pelican-into-training-set in May 2025 on twitter.)

My 3D-egyptology-explainer showed a massive uplift for Kimi K3 and this tracks a much improved 3D capability.

tezza··on Kimi K3: Open Frontier Intelligence
Nice qualitative test!

Just like you I am super impressed by Kimi K3.

I do a qualitative benchmark series making 3D explainers and so here's Kimi K3 vs Claude Fable:

https://generative-ai.review/2026/07/kimi-k3-rush-test-vs-cl...

I've put links to the posts on GLM5.2, Opus 4.8, Chat GPT 5.5. I grab video screencaps so you can compare in detail. The full interactive Kimi output is at the bottom of the post if you want a comprehensive 3D play around

tezza··on After 7 years in production, Scarf has reluctantly moved away from Haskell
well i’ve come across it loads, especially 1/2) during REPL style build outs and 2/2) calling libraries and frameworks you are not yet familiar with.

perl another offender… is it a hash? is it an arrayref? over time you get it right, but by trial and error and looping. json suffers this too, arrays different from strings, different from numbers etc, but opaque until checked and liable to change

tezza··on GPT-5.6
Absolutely, the change in quality over time is a great yardstick.

Nice to see you did the quality level comparisons and did three passes.

I've been using that technique myself on my image gen reviews[1] and it also works well in presentations and for personal study.

[1] https://generative-ai.review/2025/12/beast-mode-activated-op...

tezza··on GLM 5.2 vs. Opus
I've just put GLM 5.2 through my __qualitative__ benchmark. I was quick enough to capture Claude Fable, so now you can compare GLM5.2 vs Claude Fable vs Opus 4.8 vs Chat GPT 5.5

https://generative-ai.review/2026/06/glm5-2-from-z-ai-vs-cla...

I've structured it side-by-side. You can clearly see where the private models excel, and where GLM 5.2 is still really good.

tezza··on Claude Fable 5
There are benchmarks if you want quantitative results. Mine is qualitative, and clearly billed as such. Comparison and contrast still possible.
tezza··on Claude Fable 5
check the backlinks[1][2] in the article before you start throwing around accusations. I am not (yet) a person that has advanced notice and access to models.

Fable just got announced and I did a rush out article because people are curious. I released the post mere hours afterwards and it takes time to create the output, slice into videos, make a wordpress article on top of taking my son to basketball training and eating dinner. I’m in London and this was all happening at 1am.

If you check the links my previous articles have all the juicy stuff you are criticising me for not having with little preparation.

How is a side by side direct comparison NOT precise?

[1] first in series from 2025: https://generative-ai.review/2025/05/vibe-coding-my-way-to-e... . This has all the background you are talking about in the Appendix

.

[2] https://generative-ai.review/2026/05/vibe-coding-my-way-to-e... . Second in series 2026 has a side by side table of what changed. This is what is possible with more than a few hours advanced warning.

tezza··on Claude Fable 5
I did a qualitative side-by-side of Claude Fable vs Opus 4.8 vs ChatGPT 5.5

https://generative-ai.review/2026/06/claude-fable-rush-test-...

I get them to make a 3D explainer animation. You can clearly see Fable is much improved on both Opus 4.8 and ChatGPT 5.5.

Better Textures . A nifty camera follow . Humans rendered better . ... see for yourselves

tezza··on Side by side videos of Claude Fable vs. Opus 4.8 vs. ChatGPT 5.5
I did a rush review of Claude Fable in my benchmark test and compared it to Opus 4.8 and ChatGPT 5.5.

Side by side videos to compare and contrast.

The 3D viewers are also available run right in your browser.

tezza··on Ask HN: What was your "oh shit" moment with GenAI?
MidJourney public discord channel.

The amount of masterpiece level art flowing per hour was astounding.

For every one doing a ninja waifu, there were ten doing art from davinci and leonardo crossed with hockney.

it almost gave you art sickness

tezza··on London's Smallest Public Sculptures
The John Snow Pump… where they sealed it up to stop Cholera is very small.

Also Novelty Automation (WC1R 4AX) is a tiny interactive museum/wharf end arcade which has some very intricate mechanical entertainments. Some are tiny.

tezza··on ChatGPT Images 2.0
I've rushed out my standardised quality check images for gpt-image-2:

https://generative-ai.review/2026/04/rush-openai-gpt-image-2...

I've done a series over all the OpenAI models.

gpt-image-2 has a lot more action, especially in the Apple Cart images.

tezza··on Britannica11.org – a structured edition of the 1911 Encyclopædia Britannica
2004: https://archive.org/details/britannica-2004

2009: https://archive.org/details/britannica-multimedia-dvd-2009-d...

2012: https://archive.org/details/britannica-dvd_20230709

2013: https://archive.org/details/encyclopedia-britannica-dvd-2013

tezza··on Traders placed over $1B in perfectly timed bets on the Iran war
Wait until they find out how many people place perfectly timed horse racing bets on the winning horse, JUST before the start of the race. 100% of them knew to back the winning horse
tezza··on JVM Options Explorer
How is this different to system tuning parameters in Linux /proc, FreeBsd, Windows Registry, Firefox about:config, sockopt, ioctl, postgres?

Zillions of options. Some important, some not

tezza··on John Bradley, author of xv, has died
> take part of the color space and map it uniformly to a different part of the color space

fyi Affinity Photo has recolor and hue filters that will do just that.

I used it for my video game art.

tezza··on Data centers are transitioning from AC to DC
They were Thunderstruck?
tezza··on Ask ChatGPT to pick a number from 1-10000, it generally selects from 7200-7500
when you make a program that has a random seed, many LLMs choose

   42
as the seed value rather than zero. A nice nod to Hitchhikers’
tezza··on Excommunicated devs making games with AI
My latest game BossBattle[1] (html5, have a play) uses AI for graphics, some strobe effects, a C64 loading screen shader.

I have decided to lean in to it and I will document all the places I use AI in the game on my blog[2]. Not everything works, notably 3D assets[3] and sound effects.

There is a lot of human content… i paid for a lot out of my own pocket and have limited budget. It started in 2021 before chatgpt. LLMs cannot do everything and that’s not the purpose.

Generative AI makes me as an solo indie dev able to make the game. Without the AI the game wouldn’t exist

[1] http://epicwin.team/play/solo/BossBattle/ - (public beta) .

[2] https://generative-ai.review .

[3] https://generative-ai.review/2025/08/3d-assets-made-by-genai...

tezza··on If AI writes code, should the session be part of the commit?
I put a link to the LLM session at the end of the commit, and prefix with POH: if I wrote it by hand.

POH = Plain Old Human

Easy to achieve.

Why NOT include a link back? Why deprive yourself of information?

tezza··on Terminals should generate the 256-color palette
Terminals are text. Text adds features missing from gui namely:

* Ad Hoc

requirements change and terminal gives ultimate empty workbench flexibility. awesome for tasks you never new you had until that moment.

* Precision

run precisely what you want, when you want it. you are not constrained by gui UX limits.

* Pipeline

cat file.txt | perl/awk/sed/jq | tee output.result

* Equal Status

everything is text so you can combine clipboard, files, netcat output, curl output and then you can transform (above) and save. whatever you like in whatever form you like, named whatever you like.

tezza··on I want to wash my car. The car wash is 50 meters away. Should I walk or drive?
Thank you all! We needed further data points.

comparing one shot results is a foolish way to evaluate a statistical process like LLM answers. we need multiple samples.

for https://generative-ai.review I do at least three samples of output. this often yields very differnt results even from the same query.

e.g: https://generative-ai.review/2025/11/gpt-image-1-mini-vs-gpt...

tezza··on Software factories and the agentic moment
Not sure “Digital Twin Universe” is required here. They seem rather to have rediscovered Simulators in Integration Tests from first principles? The DTU comes off as XML Databases or Information Superhighway.

Still… a really good application of agent hands-off replication.

Seems like creating a quality negative mould and then that single negative mould makes multiple positive objects en-masse.

tezza··on Start all of your commands with a comma (2009)
This is a really good practical step if you worry about name collisions

quick, easy and consistent. entirely voluntary.

Bravo

tezza··on Iconify: Library of Open Source Icons
I’ve used FlatIcon extensively. My use case is video games rather than web design.

https://www.flaticon.com/

tezza··on CLI agents make self-hosting on a home server easier and fun
Wait… tailscale connection to your own network, and unsupervised sysadmin from an oracle that hallucinates and bases its decisions on blog post aggregates?

p0wnland. this will have script kiddies rubbing their hands

Page 1 of 22Next →