HNHacker News
TopNewBestAskShowJobs

vunderba

9,036 karma · joined July 20, 2023

Velkommen to the profile of Shaun Pedicini!

Fun fact about me: I enjoy crawling through really tight awkward spaces in the remote possibility that it's actually a portal to Narnia.... it never is.

**** My blog ****

https://mordenstar.com/blog

**** My projects ****

https://mordenstar.com/projects

Loopweave - Opensource app that lets you drop a audio file in and automatically extract seamless loops for ambient soundscapes

https://loopweave.specr.net

Hackerman Recorder - Endless animated screencast of random source code pulled from Github in an old Borland IDE. Useful for ASMR in the background while you work.

https://hackerman.specr.net

Redhook's Revenge II - Official sequel to the original DOS game from 1993 which patches binary to inject new trivia questions.

https://redhook.specr.net

Clockwork Chamber - A bunch of different ways to visualize clocks because I have nothing but time on my hands.

https://clocks.specr.net

Shah Kur - Invisible Chess - A 2d/3d blindfold chess trainer with full voice control to play on your phone as you walk.

https://shahkur.specr.net

Lend Me Your Ears - A Simon-toy inspired educational game to help teach how to play piano by ear.

https://lend-me-your-ears.specr.net

Glyphshift - A browser extension which swaps out words with braille/morse/kana while you browse.

https://mordenstar.com/projects/glyphshift

GenAI Showdown - A comprehensive comparison of state-of-the-art generative image models.

https://genai-showdown.specr.net

Aladdin's Mathemagical Flying Carpet - Teaches the times tables as you fly through the cave of wonders

https://mordenstar.com/projects/mathemagic

submissionscomments
vunderba··on Strange Parodies of Atari 2600 Video Game Box Cover Art (2008)
These are great.

Salvador Dalí’s Pinball Thrills could also have been the title of the pinball segment on Sesame Street back in the day.

Every sport ever in pong form was completely true. Magnavox Odyssey Pong + a crappy blue cellophane transparency overlay == ice hockey!.

Anyone remember Slalom for the NES?

https://imgpb.com/hvszd

vunderba··on Strange Parodies of Atari 2600 Video Game Box Cover Art (2008)
Yeah I think the first-party early titles (Excitebike, etc) were definitely more representative of the game play.

But third parties definitely didn't respect this. I'm sure everyone who grew up in this era remembers the infamous Megaman cart by Capcom [1].

One of the things that Nintendo learned from Atari was that massive amounts of shovelware did not help their cause.

As a third party company, you were actually limited to a certain number of games you were allowed to release per year in North America (five I believe). There were ways around that, like Konami forming the subsidiary Ultra.

Additionally they locked the system behind the 10NES chip that made unauthorized game production difficult (unless you pulled the classic Taiwan voltage snap trick or just blatantly copied the source from the US Copyright Office like Tengen).

[1] - https://imgur.com/a/5hUUDnM

vunderba··on Show HN: Pac-Bench – How well can models one-shot a Pac-Man game?
I was just coming here to say this. I looked at all of them, and Opus 5.5 at high effort seems significantly better than every other clone.

Buffered controls, different pathfinding AI for each ghost, level transitions, etc.

Gpt-5.6 sol high technically completed the assignment as well, but the gaps within the pipes used to construct the maze, the lack of pause when eating a ghost, etc made it feel far less polished.

vunderba··on Show HN: Lofi Cities – Pixel-art city nights with browser-generated lofi
I haven't had a chance to play with Opus 5.5 yet, but I guess if there were a way to generate each of the pieces (aka the layers) then it might be possible:

- parallax, which might be some clouds moving in the background.

- weather effects which is mostly just an overlay

- animated sprites that interact with, or at least move along, the different sections of the static scene.

It'd be neat to see how well it could recreate some of the larger set pieces like Westminster palace.

I've definitely seen some pretty impressive automated AI voxel stuff (and that was 6 months ago) and that's having to deal an extra dimension. [1]

Personally I still don’t see a situation in which you could do it completely hands-off though. I think you’d end up with some pretty good scenes and some really terrible scenes requiring post-adjustments. It would be an interesting experiment at any rate!

[1] - https://github.com/Ammaar-Alam/minebench

vunderba··on Show HN: Lofi Cities – Pixel-art city nights with browser-generated lofi
Maybe to a degree but this isn’t just somebody using Seeddance to generate a first-frame-and-last-frame video loop and calling it a day.

There's definitely some manual curation. They even have assets that are clearly being placed into the scene, like the boats moving in the water.

You could theoretically generate all possible animated sprites (birds, people, vehicles, etc.) for use in any given cityscape, but automatically segmenting/detecting the different valid "paths/waypoints" where those assets should be placed would probably end up looking pretty janky.

vunderba··on Show HN: Lofi Cities – Pixel-art city nights with browser-generated lofi
Yes according to the author - the assets were created with AI tooling.

https://safaelmali.gumroad.com/l/lofi-cities-complete-collec...

vunderba··on 2DWillNeverDie
Yeah, I’d agree, I definitely wouldn’t classify FF7 as graphically timeless.

I think most people would agree that early 3D games (PSX and N64 era) - particularly ones attempting to achieve texture-mapped photorealism, like FF7, GoldenEye, Tomb Raider etc. did not age gracefully, at least where graphics are concerned.

Contrast that with something like Mario 64, which, while still chonky, looks way better because it’s so highly stylized.

vunderba··on What even is an OS now?
> As for DOS, I have had the privilege of hacking together a bitbanging VGA adapter, so I know for a fact that drawing to the screen is really as simple as setting values in an array, then, scanning it out to the display.

Screen mode 13h was definitely fantastic. Being able to light up a pixel by mapping directly to the 64k contiguous bytes in the address space was very intuitive.

And yeah dealing with Win32 stuff definitely came with its own set of complexities, I practically had the Charles Petzold Programming Windows book memorized back in the day.

But let’s also not forget Visual Basic, which outside of maybe Delphi, I'd consider one of the greatest RAD environments ever created. Being able to just grab a button onto a form and then double-click it which would instantly create a handler that you could add your logic to was very approachable even for younger audiences.

The only problem of course is that Visual Basic wasn’t bundled with Windows, so it was a lot less accessible than QBasic which came bundled for free with MS DOS 5.0.

vunderba··on What even is an OS now?
Absolutely~ I remember many summer days as a kid trying to figure out how to send commands to our serial-port driven 2400 baud modem using the AT command set:

  OPEN "COM1:2400,N,8,1,BIN" FOR OUTPUT AS #1 
Followed by my parents yelling at me to get off the phone since we only had the one landline. :)
vunderba··on What even is an OS now?
Yes! To my knowledge, I think almost every single keyword in the built-in help had a small example program at the bottom that you could type in.

For example GET/PUT had this one:

  SCREEN 1
  DIM Box%(1 TO 200)
  x1% = 0: x2% = 10: y1% = 0: y2% = 10
  LINE (x1%, y1%)-(x2%, y2%), 2, BF
  GET (x1%, y1%)-(x2%, y2%), Box%
  DO
      PUT (x1%, y1%), Box%, XOR
      x1% = RND * 300
      y1% = RND * 180
      PUT (x1%, y1%), Box%
  LOOP WHILE INKEY$ = ""
vunderba··on What even is an OS now?
Agreed! What’s this weird `SCREEN` command? Punch F1 and you get a complete breakdown of all the screen modes, color palettes, and their resolutions.

As I mentioned in another comment, I think if you didn’t grow up with it, you just underestimate how amazing QBASIC was as an IDE, a programming language, and perhaps most importantly as an educational tool. I mean, it even came with actual BASIC programs (GORILLA, NIBBLES, etc) right there that you could use to explore and tweak the values as you learned.

There’s no better way to learn how something works than to physically play with the actual values, change constants, change variables, and observe how that affects the behavior of the game.

vunderba··on Show HN: 12 Puzzle – a sliding puzzle on the hyperbolic plane
As somebody who has also built sort of "different" slide puzzle (trivia meets slide puzzles), I'm a huge fan of new ones.

Thanks for adding the little touches like letting me click a tile and have it displace the entire row to the left/right, and also the quotes underneath the tiles is pretty great too!

vunderba··on What even is an OS now?
Reading through all the comments, I think there’s also a bit of a generational gap here.

There’s a huge difference between a kid turning on a Commodore 64 or an Apple II and being dropped into the relatively no-frills versions the BASIC interpreter versus a kid in the early ’90s who had a DOS machine with MS-DOS 5 bundled with QBasic which IMHO is a more representative example of a “batteries included” BASIC environment that was significantly more user-friendly.

I also encountered QBASIC as an elementary school kid - it felt like the closest thing to sorcery. People who grew up in the age of the internet really underestimate how awesome it was to be able to just be able to place the cursor on a more obscure command (like PEEK/POKE or POINT) and see something like this:

  Returns the current graphics cursor coordinates or the color attribute of a specified pixel.

  'This example requires a color graphics adapter.
  SCREEN 1
  LINE (0, 0)-(100, 100), 2
  LOCATE 14, 1
  FOR y% = 1 TO 10
      FOR x% = 1 TO 10
          PRINT POINT(x%, y%);
      NEXT x%
      PRINT
  NEXT y%
vunderba··on Show HN: I made a WWI dogfight game that runs in the browser
Obligatory Calvin and Hobbes comic:

https://imgur.com/a/dNS5r5X

vunderba··on Show HN: QBasic 1.1 in the Browser
Nice job~

For the past 6 months I've also been working on a period-accurate QBASIC 1.1 simulator (and corresponding virtual hardware layer) that runs entirely in the browser.

I've got a stack of books on my desk from when I was a kid in the 90s that have literally been my religious texts (Undocumented DOS, The Programmer's PC Sourcebook, etc), so I know how hard it is to get this right.

I've actually posted about it a couple times here in HN.

Anyway small bit of feedback:

Technically if you don't clear the screen (using CLS, or set a SCREEN mode other than zero), the cursor retains its position even between program executions.

So basically if you have a program like this:

  PRINT "WHY IS A RAVEN LIKE A WRITING DESK"
  END
And then you run it again - the output screen should display:

  WHY IS A RAVEN LIKE A WRITING DESK
  WHY IS A RAVEN LIKE A WRITING DESK

There's a thousand other subtleties like this that make the proverbial "last mile" sometimes feel insurmountable. Fun though!

* Your definition of "fun" may vary.

vunderba··on Qwen Image 2.1 beats Google Nano Banana 2.0 with minuscule 7B parameter model
I haven't done enough tests on the editing capabilities but as far as the strict text to image goes, Alibaba is reaching to claim it is comparable to NB 2.

It's definitely a nice upgrade from the last open-weight version (Qwen-Image 1.0) at least in terms of strict coherence and prompt understanding... but a lot of the outputs seem distilled for lack of a better word - likely trained on poor synthetic data. There's also some elements of tinging that very much reminds me of early gpt-image outputs.

You can mitigate it a bit using better samplers (like res_2m paired with the beta_57 scheduler, etc). A lot of people have also seen better results using a higher CFG than what is recommended by the Qwen team.

On my GenAI Image Showdown bench which emphasizes adherence to prompts, Qwen-Image 2.1 clocked in at 7 out of 15 which is SOTA for an open-weight model (Ideogram 4 is the only other open-weight model that outscored it), but the quality is frustratingly inconsistent.

https://genai-showdown.specr.net/?models=local

vunderba··on Truman World
As avaer already mentioned it seems to be a thin wrapper around H3 Max Director, which generates real-time video and takes “natural language directions” that are then relatively seamlessly incorporated into the continuous video stream.

It's like a Choose-Your-Own-Adventure dialed up to 11.

I played around with it briefly on Fal.ai generating a cartoon, but it’s more of a technical curiosity for the time being since I don’t have a Scrooge McDuck vault of gold coins.

vunderba··on I'm addicted to buying domains for my vibe-coded apps that get 0 visitors
This is the way - 99% of my stuff is Caddy-proxied through subdomains to a single Debian VPS and works great even when deluged with the occasional HN-levels of traffic.
vunderba··on Qwen Image 2.1
I agree. In fact the entire reason I initially built GenAI Showdown was because many of the comparative tests on places like Image Arena Leaderboard [1] are not designed to challenge models on prompt adherence. Even when they are, a considerable number of amateur judges tend to prioritize aesthetics over adherence or instruction-following.

I'll likely be redoing that particular bench with added minimum passing criteria of an anvil.

[1] - https://huggingface.co/spaces/ArtificialAnalysis/Text-to-Ima...

vunderba··on The Effect of CRTs on Pixel Art (2024)
If you want to look at a company that does pixel art right for the modern era, Yacht Club Games is probably the gold standard with games like Shovel Knight and Mina the Hollower.

https://www.yachtclubgames.com/blog/the-art-of-the-game

vunderba··on Qwen Image 2.1
Well, the results are in, at least for text-to-image (the editing bench will come later).

Qwen-Image 2.1 is definitely a pretty big leap over the last open-weight version, Qwen-Image 1.0, released back in August of last year and managed to score 7 out of 15 as opposed to its predecessor which scored 4 out of 15.

Even though it's significantly smaller, 7b vs 20b, it's multimodal (so you don't need a separate image-to-image model like you did with Qwen-Edit), more coherent, and significantly faster even when outputting at higher 2K resolutions. However, in my testing, I found that I had to play with dialing up the CFG depending on the complexity of the prompt.

I've also added a progress dropdown under Model Performance so you can see how cloud vs. local models have been trending since 2024. Spoiler: June of this year released some of the biggest bangers (Krea 2, Ideogram 4, and the kind of slept-on Boogu-Image 0.1).

Downsides:

- It was clearly trained on at least some level of synthetic training data, and it shows in some of the subpar outputs in terms of fidelity. Some of this you might be able to iron out with a refiner model downstream or a custom LoRA but time will tell.

- They've moved away from the permissive Apache license. Commercial usage is only allowed by request.

Comparisons:

https://genai-showdown.specr.net

If you just want to compare local models only:

http://genai-showdown.specr.net/?models=local

vunderba··on Qwen Image 2.1
Agreed. There's also a lot of bad tinging/yellow saturation that very much reminds me of early gpt-image outputs on a lot of the non-cherry picked stuff I've been seeing on Twitter/Reddit.

A lot of people were putting ZiT as a refiner downstream in early Qwen-Image 1.0 workflows, so I'm wondering if we're going to see something similar with 2.1.

vunderba··on Qwen Image 2.1
Well this is HN - original home of the "ummm actually..." - so I appreciate when people pick all the nits. :)

Even though I prompted for a crucible in the prompt, I think the fact that the prompt also contained terms like “blacksmith” and “hammer,” caused it to lean towards anvils over crucibles in some of the pictures (which as you brought up makes more sense anyway).

vunderba··on Qwen Image 2.1
Wait... is that true? I don't think the original Queen Image 2.0 was ever released beyond an API. At least, I don't remember a public weights release.
vunderba··on Qwen Image 2.1
Can do! GPT-Image-2 already scored unsurprisingly very high: 12 out of 15 on text-to-image, and 10 out of 12 on image-to-image.

The three benchmarks it failed on (D20, Flat Earth, and Banded Snake) are pretty difficult, so I'd be surprised if 2.5 manages to pass them, but I’ll add it for completeness’ sake later this week.

vunderba··on Qwen Image 2.1
That’s a good catch. Yes, all scoring is done through manual review since relying on a VL model for these kinds of meta-metrics is a sort of loose equivalent of gödel's second incompleteness theorem.

I’ll have to think about this one. When I crafted the prompt, I wasn’t really thinking about the differences between a crucible and an anvil. It was more the visual of an archangel smelting halos for newly arrived heavenly beings.

vunderba··on Qwen-Image-2.1: Compact, efficient, and unified image creation
This is my experience as well. Ideogram4 (assuming you are willing to put in the work to use the proper structured JSON input) is very accurate when it comes to text rendering in an image.
vunderba··on Qwen Image 2.1
So thoughts

Positives

• It's a heck of a lot smaller than Qwen-Image 1 (20b parameters) at only 7b, making it one of the smaller open-weight models available (Z-Image Turbo is one of the few that is smaller at 6b) when compared to Ideogram, Krea2, Flux2, etc.

• It supports native transparency (Qwen's team, as far as I know, is the only one attempting to tackle this). Even though it's relatively trivial to set up background removal postprocessors, it's also neat to see it natively supported.

• It's fast using QwenImage2.1 convrot, a 1MP image took around ~5 seconds on an RTX4090.

Negatives

• The license (assuming you respect it) is far more restrictive. The original Qwen Image 1 was released under the standard Apache license; this one explicitly forbids commercial usage without obtaining a separate license. On the other hand, a lot of us didn't expect the Qwen team to ever release "weights-available" ever again.

Qwen-Image 1.0, released about a year ago, only scored 4/15 on my GenAI Showdown Benchmarks. Since that time, they've been upstaged by Krea 2 (6/15) and Ideogram4 (8/15). I'll post the new results once I have some more time to run them.

https://genai-showdown.specr.net

vunderba··on Qwen Image 2.1
Boogu-Image has the Apache 2.0 License [1] (good coherence, but outputs can look synthetic).

And Krea 2 has a community license [2] that is fairly permissive - I think commercial usage is allowed under $1 million.

Boogu-Image scored 6/15 and Krea 2 scored 7/15 on my GenAI Showdown benchmark [3] - only Ideogram4 eclipses them in terms of local models, but its got a far more restrictive license and the JSON structured inputs can be a pain to work with.

[1] - https://github.com/Boogu-Project/Boogu-Image

[2] - https://www.krea.ai/krea-2-licensing

[3] - https://genai-showdown.specr.net/?models=fd,hd,kd,qi,f2d,zt,...

vunderba··on Tabasa – A blindfold chess tactics trainer
Nice. Some Feedback

• Wish this was a web app though, I have a general policy of not isntalling unknown software on my computer.

• Website needs a demonstration video (or at the very least some screenshots)

If you're interested in improving your blindfold chess techniques, I actually put together a pretty in-depth blindfold chess trainer website about a year ago called Shah Kur. It lets you play Stockfish AI with different blindfold variations (Last N hidden, pieces reduced to color only, etc.), and it uses voice activation/TTS so you can play on BT headphones while you're out walking without having to physically look at your phone.

It also has a positional recall game constructed from historical games, so you’re building an actual practical “memory chunking” perspective instead of just random pieces on a board that would never correspond to an actual game.

https://shahkur.specr.net

Page 1 of 34Next →