HNHacker News
TopNewBestAskShowJobs

Macuyiko

1,124 karma · joined July 30, 2010

opinionated researcher • data scientist • programmer • hacker • book reader • gamer • fast walker

see seppe.net and blog.macuyiko.com

submissionscomments
Macuyiko··on VNC Resolver
Wow, this brought back memories. I could swear I wrote a blog post about this years ago but couldn't find it.

A quick search on the local file system revealed `vnccrawl/crawler.py` from 2016 [1] using what looks like a Shodan data dump and calling out to `vncviewer.exe`. I remember randomly logging into some instances and also seeing a lot of cool random systems, including a lot of them controlling industrial systems. Guess I never ended up writing that post.

One would think that on today's Internet it would take only a couple of seconds for those to get compromised, but obfuscation as security, perhaps?

[1]: A random tip from that file: Using a password of 12345678 gives access to way more 'weakly secure' instances.

Macuyiko··on Long-term nuclear waste warning messages
This reminds me of a short story by Ken Liu, The Message, which details a xeno-archaeologist digging into a place full of radiation. The main character doesn't get the warning message until it is too late and almost loses his daughter.

Googling it now it seems at one point is was going to get adapted to film [1], but seems like that went nowhere.

[1]: https://reactormag.com/ken-lius-the-message-to-get-big-scree...

Macuyiko··on Teachable Machine (2017)
v1 used a very limited (albeit very easy and already quite impressive) form of transfer learning, e.g. take a pretrained network's 1000dim vector outputs given a bunch of images belonging to three sets (since the original was trained on Imagenet), and then just use K-NN to predict what a set "new" image falls into.

v2 does actually finetune weights of a pretrained network. At the time, it was a nice showcase how fast fast JS ML libraries were evolving.

Macuyiko··on Trinary Decision Trees for missing value handling
Came here to cite your work, I even mention "CloudForest" in my slides still as "an interesting implementation that is also capable of handling NANs in DTs in a slightly different way." Crazy this has already been 10 years.
Macuyiko··on Gaussian splatting is pretty cool
Very interesting, indeed, they seem to be driven by better fluid simulations... remarkable that they find their way into games. I was always under the impression that Navier Stokes was hard in 3d, but it does seem like there are performant solutions now that are easily offloaded to the GPU, e.g. https://github.com/chrismile/cfd3d (and NVIDIA also has some blog posts about it).

Edit: I also just found this: https://www.youtube.com/live/569oSOSoKDc?si=8V5buRMoI3IKqLQp... -- which is very close to what you describe and fully matches the kind of particle systems I was hinting at, thanks!

Macuyiko··on Gaussian splatting is pretty cool
Very cool work.

As a bit of tangent (but wondering whether someone can answer) - the article also makes mention of point based rendering and indeed the fact is has been a staple of particle systems for a long time. However, especially with recent games, I have noticed (purely subjectively) a very subtle shift to a new style of particle systems which are on the one hand fully point oriented (compared to (textured) fragments) but on the other behave more like a physics systems.

Examples:

- Hogwarts (heavily): https://www.gamespot.com/a/uploads/original/1816/18167535/40... - Forspoken (heavily): https://oyster.ignimgs.com/mediawiki/apis.ign.com/project-at... - Starfield (though more rarely): https://dotesports.com/wp-content/uploads/2023/08/temple-loc... - AC6 - FF16 (heavily)

It's more obvious when you see it 'in motion'. The common denominator seems to be particles as colored transparent points with physics. Especially on console systems it seems that developers are using this for very cheap (CPU-wise, all on GPU) effects.

Anyone in gamedev who has some insight in this?

Macuyiko··on Valve is not willing to publish games with AI generated content anymore?
Well... Valve is a very interesting study in that regard. They have voted for violence, topics bordering to meme hate speech, pornography, but have voted against crypto shit and now AI.

Not making a verdict either way but I find it interesting. I'd like to know more about the internal discussion(s) that took place to establish their frameworks. Especially given the company is private.

Macuyiko··on Ask HN: Is GitHub down?
Very good points. Meanwhile I have clients asking me why they can't have a status page to which I reply: you can, but ultimately to be completely fail proof it will be a human updating it slowly. To which they reply: but GitHub or X does it...

Very infuriating, that.

Macuyiko··on Ask HN: Is GitHub down?
Update

We have identified the root cause of the outage and are working toward mitigation Posted 4 minutes ago. Jun 29, 2023 - 18:02 UTC

Macuyiko··on Ask HN: Is GitHub down?
EU here. Actions are failing to run. Rest is kinda ok.
Macuyiko··on AI Is Catapulting Nvidia Toward the $1 Trillion Club
You're right, but I don't think so. From the moment that 80/99% realizes they're out of work, it's over. That's why you see idiot anti-AI spokespeople showing up, why Altman is invited to Bilderberg, why EU is making AI-laws. They're not against AI, as such, but please do not "awaken" the working class. Keep it for "trusted parties" or the military only. What I am curious about is how NVIDIA will position itself against that background.

Personally, I did really wish this would have been a new-era moment where society would take a step back and evaluate how we are organizing ourselves (and living, even), but I fear that AI comes too late for us, in the sense that we're so rusted and backwards now that we cannot accept it. Or any important change, in fact. It's pretty depressing.

Macuyiko··on Modern SAT solvers: fast, neat and underused (2018)
Curious to hear about your preferred alternative. Poetry?
Macuyiko··on John Carmack’s ‘Different Path’ to Artificial General Intelligence
Fun interview but barely any meat for those in the field. Just very general questions and answers, saying that the road is murky, but nothing about e.g. if Transformers / attention are the way to go forward, multi-modal models, reinforcement learning + self-supervised learning.
Macuyiko··on Summer Afternoon – A WebGL Experiment
Just played through Lil Gator Game with my 3.5y old and she loved it.
Macuyiko··on Ask HN: Do others with young kids prefer work to holidays?
Lots of recognizable stories in this answer - thanks for sharing.
Macuyiko··on YouTube confirms that it has removed the “sort by oldest/newest” option

    yt-dlp --flat-playlist --extractor-args youtubetab:approximate_date "https://www.youtube.com/c/AndreasKling" --print "%(upload_date)s %(id)s" 1> out.txt && sort out.txt
Takes longer if you want precise dates:

    yt-dlp --simulate "https://www.youtube.com/c/AndreasKling" --print "%(upload_date)s %(id)s" 1> out.txt && sort out.txt
Macuyiko··on Lost something? Search through 91.7M files from the 80s, 90s, and 2000s
This by any chance? https://www.youtube.com/watch?v=ccEpeuqdHGM
Macuyiko··on Users reporting artifacts appearing in old images stored in Google Photos
Each passing month I am more and more convinced that I should urgently get a) my mails out of my personal (non paying) Gmail account and b) get my associated photos out of Google Photos, mainly because of baby pictures which I cannot fathom to lose.

But then I think about the convenience of my 2TB Google Drive and I can't do it.

And I think, too, of the Polaroid pictures my parents took and gave me in a box once I got married and are now fainted. Mayhaps algorithms deciding that it's time to let go of your own data is just the modern form of yellowing.

Macuyiko··on This X Does Not Exist
I was planning to make a snarky comment here as follows: Ah yes, thisxdoesnotexist.com, let's repeat the fad from a few years ago when StyleGAN 1/2 was all the rage but has been widely outdated since. Still, it can come in useful when giving an AI talk to completely unsuspecting managers.

Though I was pleasantly surprised https://www.derekau.net/this-vessel-does-not-exist at least had been updating their approach, up to DALL-E 2 even. Not yet Stable Diffusion but I do applaud the effort.

(In all seriousness: just goes to show how quick things evolve in this space.)

Macuyiko··on Flip the Switch for 5.5 Seconds
First try:

> You got 5.523 seconds. Aaand just like that, you went over.

I think that is good enough.

Macuyiko··on On Device Learning
> Sorry if I got a little bit off topic there.

Not at all. Though forgive me for not agreeing with some key points you raise.

> Curiosity is not a necessity. While curiosity is integral part of what makes a human well, human, it doesn't have to be hardwired. > The keyword here is cheap. Any sufficiently powerful maximizer with an infinite horizon __has__ to develop curiosity otherwise they will not be able to maximise their reward function.

With that I do in fact agree. I think curiosity was a quick solution to fix some immediate problems I was seeing from the pain-slash-survival angle, perhaps from a belief there should be more to humanity. I also feel it is rather emergent and should (must!) emerge pretty fast, even, in order to survive.

Actually typing out the last sentence made me realize another meta-meta-level of intelligence. Whereas the basic reward function is level 0, the chemistry surrounding and interactions with out body it might be level 1 (might be, because they are probably emergent as well), evolution is definitely level 2 (or 1) - the multi-simulations of agents. On top of that, there's the fact that initialisation is cheap: meaning that even if some emergent properties are highly necessary on top of the basic reward function, and might lead to very complex aspects later on, a designer (and this is a very badly chosen word mayhaps) would be prepared to deal with those given the fact that there are many chances. Many one-cell organism striving to do better in the "soup". The more I think about this, the more I start becoming convinced that computational biology should have been a serious field (and many are saying this).

> In-fact, I'd argue that this is true for most mammalian functions such as taking care of our pack, exhibiting pro-social behaviour and so on, but there is a caveat. For this to happen, there needs to be an actual benefit in the behaviour.

See, this is where I respectfully disagree. The benefit can be emerging from a longer-term simulation rather than immediately. You might say: sure but what are the chances of this happening? Well how many intelligent species like ours have we encountered so far? On this planet, in this universe?

> Due to evolution operating mostly linearly with small changes through genetic and epigenetic information passing, there seems to be relatively little variation between generations, which then implies that it is difficult for some candidates to overwhelm everyone else in a winner takes all fashion, hence maximization of replication will eventually result in cooperation simply because that allows genes in support of it to continue replicating, effectively self-selecting for itself.

Yes and no. From a gut feeling I agree though I also think small changes tend to take over very rapidly in a population pool once they show up. The waiting is mainly for the showing up part.

> We saw this in the OpenAI video where agents eventually learned to cooperate in what was effectively a prisoners dilemma. In the video, there were two teams, the hiders and the seekers, in an environment that could be manipulated.

I saw that as well, and this is why although I believe at some point this will be possible, the main reason why I agree with hotz is because our simulations suck and always allow for exploitation. Unless it doesn't, of course. The on-device part, hence, for me, is not that necessary. But it means we should have a very robust simulation (which so far we don't have in any area of RL and associated topics; digital twins are a joke; people making them care more about dataviz; and so on).

> Regardless, what causes us to learn is neurotransmitters getting released because certain circuits activate, the neurotransmitters charge a neuron which causes it to fire. Connections between neurons get reinforced if they are frequently used, which reinforces them, makes them cheaper, and that inevitably reinforces certain patterns of behaviour.

Too tired to go into this but you touch upon some key differences between gradient descent and biological neuron learning, though we are getting closer to that: spiking neurons, memory cells, even real-neuron cell chips. I am not sure electronics wouldn't be able to emulate it correctly in the end. If after all, P <> NP, then what does the "real" difference of a computational time step make?

> What i propose is that we should instead look into analogies as a means of learning. Humans seem to be great at using analogies. Mathematically speaking, an analogy is a functor between categories.

Agree, and surprising that this is still such an open question in AI even given the one-shot and zero-shot learning research. Though it seems like this has been put on the backburner yet again. It amazes me even today how young humans are so good at that. Like someone said: show a cartoon tomato and a real tomato to a toddler. Next time show a cartoon of an elephant and wait until they see a real elephant. They will shout: elephant. Though on the other hand, the solution for this might be very close to use. A small architectural or multi-modal change. I was more pessimistic about this a few years ago, but less sure today.

I think the main piece missing of the puzzle is stepping away from supervised learning, self-supervised learning, and going for a continuous self-supervised reinforcement learning, where predictions for t+1 are continuously matched with reality, like a human brain does. The only problem is that you need to have a continuous reality. But we have that.

Macuyiko··on On Device Learning
Exactly. This is the crux of the question: what is necessary as a starting point. What is not. And given what is: what emerges over time (like nationalism).
Macuyiko··on On Device Learning
Interesting thoughts. I was having a very similar chat with a friend about this recently, including exactly the question of "what should the reward function be", or at least the most minimal one.

Some aspects we thought of:

- pain (minimal reward): this should probably be hardwired straight to the artificial brain, though it can't be enough as otherwise no activity would take place. The agent would learn that the best course of action is to sit still

- so we also came up with curiosity being a necessity. Encountering an unseen or hard to predict state leads to positive reward

- although I am not sure whether pain is not sufficient by itself. E.g. in nature, actions are still necessary. Otherwise, other pain signals (hunger and thirst) start showing up.

- what is tricky to figure out is how this works for more complicated intelligence such as humans. Let's be honest, most babies are fed whenever necessary by their caretakers. What causes them to learn? What causes grown ups to learn?

So something important that we'll have to figure out is what needs to be hardwired versus what emerges from 2nd order things such as chemistry, hormons, gut bacteria, upbringing, role of parents etc., and whether there's a difference between non-conscious or simple intelligence versus complex ones w.r.t. the necessity of these aspects. E.g. you might talk about aspects such as "love" (towards your partner, children, born or not, future generations) but it is much more unclear how necessary this is, and how to quantify it.

Perhaps indeed this emerges from the basic reward function but only after a meta-simulation:

> I don’t think there’s a way to learn it aside from millennia of multi agent survival competition.

Macuyiko··on Remembering the best shareware-era DOS games that time forgot (2019)
See https://news.ycombinator.com/item?id=21245827

Especially the commenter on my post: Rocks and Diamonds is great.

Macuyiko··on Dillo web browser domain is for sale
You're in luck, Andreas has been hacking on that since a couple of months. They're calling the Linux version of the browser Ladybird: https://github.com/awesomekling/ladybird
Macuyiko··on Simon Tatham's Portable Puzzle Collection
I was just thinking that dungeons and diagrams would make a great addition to this collection. Maybe a fun weekend project to implement that.
Macuyiko··on Tunneling Wikipedia through WhatsApp to (maybe?) get around WiFi restrictions
I also used this a lot while travelling to access the internet through captive wifi portals. Especially in asia this worked very well, given the huge amount of telco wifi providers in cities.
Macuyiko··on Making a falling sand simulator
Actually part 4 should be most interesting if you're interested to know how different materials can be mixed. It also looks at Noita:

https://blog.macuyiko.com/post/2020/an-exploration-of-cellul...

Kudos to the OP article. Although it doesn't go very far the presentation style is very clean and appealing.

Macuyiko··on Ali, what’s up
I don't know, I opened nl.aliexpress.com in a private browser tab and immediately got the NSFW suggestion on first try.

The class of the div where the suggestions are placed is class="hot-words", so I would say this is the result of someone spamming the search box with nasty terms and AliExpress just showing them based on 'popularity'.

Macuyiko··on Hexells – Self-Organising System of Cells Drawing Textures [WebGL]
https://distill.pub/2020/growing-ca/ - you can either train them to generate expected outcomes

https://distill.pub/selforg/2021/textures/ - this is the paper the authors refer to. Here, the network is trained to maximize activation of another filter of a CNN pretrained network (VGG, for example), which typically correspond to certain textures and shapes. The losses are typical style/gram losses.

← PreviousPage 2 of 9Next →