HNHacker News
TopNewBestAskShowJobs

tbalsam

1,036 karma · joined January 14, 2021

hi. ;P
submissionscomments
tbalsam··on Anti-AI fonts are useless and harmful
There was a story once about a boy with a wheelchair who needed a ramp to get into school, and the school made him use the loading dock ramp used for garbage and other things at the back. The school argued that it was an appropriate accommodation.

Accessibility is not accessible if you need to go through extra steps to get it.

Cerebrally, this as a solution makes sense. But if you know anyone with a vision or other impairment, gating it behind a request is not only cruel but gets within dangerous striking distance of an ADA lawsuit, for general applications.

Maybe in the legal field or specific niche cases it's possible. But this would represent a major step backwards in the work we've done lowering barriers for a population whose only difficulty in accessing common resources is because they were born, or got sick, differently than anyone else.

tbalsam··on NanoChat – The best ChatGPT that $100 can buy
This is the common belief but not quite correct! The Muon update was proposed by Bernstein as the result of a theoretical paper suggesting concrete realizations of the theory, and Keller implemented it and added practical things to get it to work well (input/output AdamW, aggressive coefficients, post-Nesterov, etc).

Both share equal credit I feel (also, the paper's co-authors!), both put in a lot of hard work for it, though I tend to bring up Bernstein since he tends to be pretty quiet about it himself.

(Source: am experienced speedrunner who's been in these circles for a decent amount of time)

tbalsam··on Fp8 runs ~100 tflops faster when the kernel name has "cutlass" in it
shocked quack
tbalsam··on Problem solving using Markov chains (2007) [pdf]
I'm not entirely sure, to be honest. If you look at the linked video, they state that it's oftentimes not in the best interest of the private equity group's moneymaking capabilities to announce that a channel has been sold out to them.

How that is in practice, I'm not sure, and I'm sure with some sleuthing it would be possible to find out at least some of it. But on the whole, I'm honestly not sure beyond that.

tbalsam··on Problem solving using Markov chains (2007) [pdf]
They unfortunately recently (last few years) sold out to private equity (which tends to glaze over fundamentals and tries to pump out massive content using previous brand quality to give it credence), so beware of quality in more recent vids:

https://youtu.be/hJ-rRXWhElI?si=Zdsj9i_raNLnajzi

tbalsam··on Large language models are improving exponentially?
There are versions of this kind of benchmark with a higher threshold, however, it only seems to adjust the timetables by a linear amount, so you're only buying 1-2 years or so depending on what you want that % success rate to be.
tbalsam··on Large language models are improving exponentially
For those curious: https://en.m.wikipedia.org/wiki/Zombo.com
tbalsam··on Large language models are improving exponentially?
The only limit is yourself

Source: One of the most classic internet websites, zombo.com (sound on)

tbalsam··on Show HN: TokenDagger – A tokenizer faster than OpenAI's Tiktoken
No! This is not good.

Iteration speed trumps all in research, most of what Python does is launch GPU operations, if you're having slowdowns from Pythonland then you're doing something terribly wrong.

Python is an excellent (and yes, fast!) language for orchestrating and calling ML stuff. If C++ code is needed, call it as a module.

tbalsam··on Look Ma, No Bubbles: Designing a Low-Latency Megakernel for Llama-1B
This is (and was) the dream of Cerebras and I am very glad to see it embraced if even in small part on a GPU. Wild to see how much performance is left on the table for these things, it's crazy to think how much can be done by a few bold individuals when it comes to pushing the SOTA of these kinds of things (not just in kernels either -- in other areas as well!)

My experience has been that getting over the daunting factor of feeling afraid of a big wide world with a lot of noise and marketing and simply committing to a problem, learning it, and slowly bootstrapping it over time, tends to yield phenomenal results in the long run for most applications. And, if not, then there's often an applicable one/side field that can be pivoted to for still making immense/incredible progress.

The big players may have the advantage of scale, but there is so, so much that can be done still if you look around and keep a good feel for it. <3 :)

tbalsam··on The Speed of VITs and CNNs
Not bad frustrations at all. That said -- IoU is how the final box scores are calculated, that doesn't change how you do feature aggregation, this will happen in basically any technique you use.

Modern SSD/YOLO-style detectors use efficient feature pyramids, you need that to know where to propose where things are in the image.

This sounds a lot like going back to the old school object detection techniques which end up being more inefficient in general, generally very compute inefficient.

tbalsam··on The Speed of VITs and CNNs
> The MSE here is not intended to be a training loss, but as a means to demonstrate that both approaches lead to almost the same result except for some rounding error.

Ah, gotcha

> I don't think that max pooling the last feature maps would be a good idea here, because it would cut off about 98 % of the gradients and training would take much longer. (The shape of the input feature layer is (1, 768, 7, 7), pooled to (1, 768, 1, 1).)

MaxPooling is generally only useful if you're training your network for it, but in most cases it ends up performing better. That sparsity actually ends up being a good thing -- you generally need to suppress all of those unused activations! It ends up being quite a wide gap in practice (and, if you have convolutions beforehand -- using avgpooling2d is a bit of extra wasted extra computation blurring the input)

> Could you elaborate on that?

Variable-sized inputs don't batch easily as the input dims need to match, you can go down the padding route but that has its own particularly hellacious costs with it that end up taking away from compute that you could be using for other useful things.

tbalsam··on The Speed of VITs and CNNs
As someone who's done a fair bit of architecture work -- both are important! Making it either or is a very silly thing, both are the limiting factor for the other and there are no two ways about it.

Also, for classification, MaxPooling is often far superior, you can learn an average smoothing filter in your convolutions beforehand in a data-dependent manner so that Nyquist sampling stuff is properly preserved.

Also, please do smoothed crossentropy for image class stuff (generally speaking, unless maybe data is hilariously large), MSE won't nearly cut it!

But that being said, adaptive stuff certainly is great when doing classification. Something to note is that batching does become an issue at a certain point -- as well as certain other fine-grained details if you're simply going to average it all down to one single vector (IIUC).

tbalsam··on The Speed of VITs and CNNs
As someone who has worked in computer vision ML for nearly a decade, this sounds like a terrible idea.

You don't need RL remotely for this usecase. Image resolution pyramids are pretty normal tho and handling them well/efficiently is the big thing. Using RL for this would be like trying to use graphene to make a computer screen because it's new and flashy and everyone's talking about it. RL is inherently very sample inefficient, and is there to approximate when you don't have certain defined informative components, which we do have in computer vision in spades. Crossentropy losses (and the like) are (generally, IME/IMO) what RL losses try to approximate, only on a much larger (and more poorly-defined) scale.

Please mark speculation as such -- I've seen people see confident statements like this and spend a lot of time/manhours on it (because it seems plausible). It is not a bad idea from a creativity standpoint, but practically is most certainly not the way to go about it.

(That being said, you can try for dynamic sparsity stuff, it has some painful tradeoffs that generally don't scale but no way in Illinois do you need RL for that)

tbalsam··on Show HN: I built a synthesizer based on 3D physics
McCormick is a popular brand of seasonings hahaha

https://i5.walmartimages.com/seo/McCormick-Pure-Ground-Black...

tbalsam··on Show HN: I built a synthesizer based on 3D physics
Yes, it's a synthesizer -- you may know it inside and out, but having demo videos showing what it can do will help people with no context get that quick "ahhh, that makes sense" moment from things. :)
tbalsam··on The Impossible Contradictions of Mark Twain
If you would like an original link non-heisted through the gwern domain, I'd encourage you to read it from the original UPenn link (University the professor who wrote this works at): https://web.english.upenn.edu/~cavitch/pdf-library/Cavitch_T...
tbalsam··on Hacker News Hug of Deaf
I did my part and manually reloaded the page about once a second for 5 minutes so that Andrew could get their dev validation beep quota in for the day (unless it's not naive hits, and unique user based, in which case this has a been a fantastically hilarious waste of time).
tbalsam··on Japanese scientists create new plastic that dissolves in saltwater overnight
I don't think they really did? A single scratch to cause it to break down doesn't seem like it would really be a scalable solution for any kind of mass produced material like this. Would cause chaos if any individual container went bad in a shipment, so it's not really addressed I feel. OP's concerns still stand.
tbalsam··on Speedrunners are vulnerability researchers, they just don't know it yet
I'm a speedrunner, and I'm pretty sure this is well known -- and accepted as standard in some categories! It's a pretty well accepted standard (to the point of the headline being almost a mild offense!).

In the gaming world, undefined software behavior is critical to this sort of thing, we see this especially in some games like the legendary exploits found in the Ocarina of Time speedruns for example.

I mean, in Super Mario World, SethBling did code injection to manually run a version of Flappy Bird (how ironic given the origin of the pipes!) in the game. By hand. No savestates. It took forever and the run through is really and truly fascinating: https://youtu.be/hB6eY73sLV0?si=nIP07o_fa6O9rauW

I speedrun things other than games as well -- and so the generalization is not just that we are security researchers, we are people who fundamentally learn the "shape" of a thing very, very well, and ways that this shape can be used to get from one state on that shape to another.

In conclusion -- yes, it can be something as simple as security research! But the joy and the beauty of speedrunning is something so much bigger and beautiful than that -- though it certainly is one outcome that can be had!

tbalsam··on Generating an infinite world with the Wave Function Collapse algorithm
Yes, this is the principle that violates the WFC algorithm and makes it no longer WFC

It is now just a procedural algorithm, which is faster than but loses some of the magic of what makes WFC _so good_.

You can tell by looking at the renders too, the before-and-after of both methods. The difference is incomparable.

That being said, it is cool as a runtime-optimized non-WFC WFC-approximating algorithm.

tbalsam··on Generating an infinite world with the Wave Function Collapse algorithm
This is neat! However, it is now longer the WFC algorithm. The idea of the WFC algorithm is that it is limited to one-at-a-time iteration, this mathematically means that every element in the input communicates its full "state" to the output.

The parallel tiling is neat, and the resulting cross-boundary replacements are neat as well, but they are in no way WFC, and as a result will lack a lot of the positive properties of the original algorithm! This is a kind of procedural generation now that does approach the properties of WFC in the limit as the number of chunks and comparisons goes to infinity.

It is a great way to scale an approximate/approximating WFC algorithm (and I'm a huge fan of the first post on this -- sent a video of results from it to someone just a month ago or so!), so it may end up being much more practical in use day-to-day due to the speed of chunk loading.

I do maintain that WFC looked a bit nicer and more natural due to the iterative collapse bit and the constraints out on it, it felt, very "cohesive". I feel like some of that has been lost in the speedier chunk-based algorithm (i.e. the entropy of the generations is much lower), but, this doesn't mean that it's not as useful or anything like that. Just a tradeoff for parallel computation (and I'd call it something like "approximate WFC" or "pre-baked approx WFC" or the like just for clarity's sake).

In any case, good work and I appreciate the problem solving skills and especially a lot of the hard problem solves like the heightmap differences and such. Very cool stuff and I'd love to see how this (and/or adjacent techniques) get incorporated into games at some point in the future! :)

tbalsam··on What is entropy? A measure of just how little we know
>> The classification of several microstates into the same macrostate, is this not a distinctly observer-centred function?

> It seems that way if we consider only our neat models, but it fails to explain why experimental measurements of the entropy of a given materials are consistent and independent of whatever model the people doing the experiment were operating on. Fundamentally, entropy depends on the probability distribution, not the observer.

I am not sure that I agree with this -- it feels a little too "neat and tidy" to me. One could argue, for example, that these seemingly-emergent agglomerations of states into these cohesive "macro" units are an emergent property limitations of modelling based of the physical properties of the universe -- but there's no way to necessarily easily tell if this set of behaviors comes from an underlying limitation of _dynamics_ of the underlying state of the system(s) based on the rules or this universe or the limitations of our _capacity to model_ the underlying system based on constraints imposed by the rules of this universe.

Entropy by definition involves a relationship (generally at least) between two quantities -- even if implicitly, and oftentimes this is some amount of data and a model used to describe this data. In some senses, being unable to model what we don't know (the unknown unknowns) about this particular kind of emergent state (agglomeration into apparent macrostates) is in some form a necessary and complete requirement for modelling the whole system of possible systems as a whole.

As a general rule, I tend to consider all discretizations of things that can be described as apparently-continuous processes inherently "wrong", but still useful. This goes for any kind of definition -- the explicit definitions we use for determining the relationship of entropy between quantities, how we define integers, words we use when relating concepts with seemingly different characteristics (different kinds of uncertainty, for example).

We induce a form of loss over the original quantity when doing so -- entropy w.r.t. the underlying model, but this loss is the very thing that also allows us to reason over seemingly previously-unreasonable-about things (for example -- mathematical axioms, etc). These forms of "informational straightjackets" offer tradeoffs in how much we can do with them, vs how much we comprehend them. So, even in this light, the very idea of modelling a thing will always induce some form of loss over what we are working with, meaning that said system can never be used to reason about the properties of itself in a larger form -- never verifiably, ever.

Using this induction, we can extend it to attempt to reason then about this meta-level of analysis, showing that because it is indeed a form of model sub-selected from the larger possible space of models, that there is some form of inherent measurable loss, and it cannot be trusted to reason even about itself. And therein lies a contradiction!

However, one could postulate that this form of loss results in any model necessarily has some form of "collision" or inherent contradiction in it -- theories like Borsuk-Ulam come to mind, and so we must eventually come to the naked depravity of picking some flawed model to analyze our understanding of the world, and hope to realize along the way that we find a sense of comfort and security in the knowledge that it is built on sand and strings, and its validity may unwind and slip away at any minute.

A very curious ideal, indeed.

tbalsam··on Arc Prize 2024 Winners and Technical Report
That is an understandable statement, and probably fair as well I feel.

Much of this comes in reference to statements from fchollet w.r.t. replacing deep learning -- around the time of the initial prize, with a lot of the much more hype marketing, this was essentially the thru-line that was used, and it left a bitter taste in a number of peoples' mouths. W.r.t. misquoting, they did say that we needed something "beyond" deep learning, not "other than" here, and that is on me.

The utility is certainly still present, if I feel diminished, and it probably is a case of my own frustrations due to previous similar issues leading up to the ARC prize.

That being said, I do agree in retrospect that my response skewed from being objective -- it is a benchmark with a mixed history, but that doesn't mean that I should get personally caught up in it.

tbalsam··on Arc Prize 2024 Winners and Technical Report
> Now that the AI research field is coming around to the idea that something beyond deep learning is needed,

I have not heard this from anyone that I work with! It would be a curious violation of info theory were this to be the case.

Certainly, some things cannot efficiently be learned from data. This is a case where some other kind of inductive bias or prior is needed (again, from info theory) -- but replacing deep learning entirely would be rather silly.

Part of the reason that a number of researchers don't take the benchmark more seriously is because it's meant to cripple the results. For example, in the name of reducing brute force search, the compute was severely limited! This turned many off to begin with. The general contention as I understand was to let compute be a reasonable amount, but this would not play well with the numbers game. Because if you restrict compute beyond a reasonable point, it makes the numbers artificially low for people who don't know what's going on behind the scenes. And this ends up biasing the results unreasonably to favor the original messaging, (i.e., "We need something other than deep learning.")

If it was structured with a reasonable amount of compute, and instead, time-accuracy gates were used for prizes, it would be much more open. But people do not use it because the game is rigged to begin with!

Unfortunately due to that, plus the consistent goal-post moving of the benchmark is why it's generally not really held with staying power in the research community -- the messaging changes based upon what is convenient for publicity, and there's unfortunately been a history of similar things in the past in the pedigree leading up to the ARC prize itself.

It is not entirely unsalvageable, but there really needs to be a turnaround of how the competition and prize is managed in order to win back people's trust. Placing a thumb on the scales to confirm a prior bias/previous messaging may work for a little while, but over time it robs the metric of its usability over time as the greater research community loses trust.

tbalsam··on Arc Prize 2024 Winners and Technical Report
I feel rather consternated that this response effectively boils down to "yes, we know we overhyped this to get people's attention, and now that we have it we can be more honest about it". Fighting for place in the attention economy is understandable, being deceptive about it is not.

This is part of the ethical morass of why some more serious researchers aren't touching the benchmark. People are not going to take it seriously if it continues like this!

tbalsam··on Arc Prize 2024 Winners and Technical Report
As a rather experienced ML researcher, ARC is a great benchmark on its own, but is punching below its weight in terms of claiming that it is a gate (or in terms of this post -- a "steward") towards AGI, and in my perspective and the perspective of several researchers near me this has watered down the value of the ARC benchmark as a test.

It is a great unit test for reasoning -- that's fantastic! And maybe it is indeed the best way to test for this -- who knows exactly. But the claim is a little grandiose for what it is, this is somewhat similar to saying that testing on string parity is the One True Test for testing an optimizer's efficiency.

I'd heartily recommend maybe taking down the marketing vibrance down a notch and keep things a bit more measured, it's not entirely a meme, though some of the more-serious researchers don't take it as seriously as a result. And that's the kind of people that you want to attract to this sort of thing!

I think there is a potentially good future for ARC! But it might struggle to attract some of the kind of talent that you want to work on this problem as a result.

tbalsam··on Long Fatigue: The exhaustion that lingers after an infection
I have chronic fatigue issues and after trying thousands of dollars' worth of supplements over the years what worked best for me were blood glutamate scavengers (N-acetyl cysteine for me, I take an unnaturally high dose of up to 8g a day), which seem to help other chronic fatigue people, as well as berberine & niacinamide.

Other things help me that are specific to me but those definitely had a rather dramatic impact on me (after an initial period of feeling like I'd been hit by a train -- it tended make me experience my fatigue up front instead of as PEMS and also reduce the symptoms of it somewhat)

Taurine is great too, there's also TUDCA which is somewhat legendary in some parts of the fatigue community, very good supplement.

Hope this helps! <3 Doing the personal work of letting go of emotional stressors and ceasing all stimulant use when possible (I know it sounds horrible) was also really hard and made a big impact over the years for me.

Still not perfect but so much more functional than I used to be. <3

tbalsam··on Bluesky is ushering in a pick-your-own algorithm era of social media
Do you have a link to said starter pack? Have been having trouble finding a good one of those specifically in this vein. <3
tbalsam··on Bluesky is currently gaining more than 1M users a day
Do...do I open the modem and read the link off a piece of paper inside, kinda like a fortune cookie?
Page 1 of 12Next →