Luckily, there are more axiomatic reasons for why softmax is the preferred way to map inputs to a probability distribution.
469 karma · joined February 27, 2012
Luckily, there are more axiomatic reasons for why softmax is the preferred way to map inputs to a probability distribution.
The sketch of the justification is something like this. We first need a function that maps from (-inf, inf) to a unique positive value, and then we need to normalize the resulting values. Setting aside the normalizing step, we imagine a f(x) that needs to fit the following properties:
1. It should be strictly positive, so that we can normalize it into a (0, 1) probability.
2. It should preserve the relative ordering of the logits to allow them to be interpreted as scores. Thus $f(x)$ should be monotonically increasing.
3. It should be continuous and differentiable everywhere, since we are interested in learning through this function via backpropagation.
4. It should have shift-invariance with respect to the input, as we don't want the model to have to learn some preferred logit-space where there is a stronger learning signal. For example, applying softmax on the values `(-1, 1, 3, 5)` would yield the same result as applying it to `(9, 11, 13, 15)`. This property can also be restated as a "scale invariance of probability ratios", where the ratio between $f(x)$ and $f(x+c)$ for a given $c$ is a constant. One useful interpretation of this property is that the learning domain or "gradient-learning surface" is stable, and high-magnitude initializations won't impede the learning process.
Taken at face value, these properties uniquely define e^x. The last property is actually pretty debatable, because in the context of machine learning, we actually do have a "preferred logit-space", namely closer to zero, for numerical stability. But there are other ways to enforce this in a post-hoc manner (e.g. weight initialization, normalization layers, etc.)
Another property that is uniquely justifies e^x and thus softmax is IIA (independence of irrelevant alternatives), which states that the odds for two classes, p_i / p_j, only depend on the logits/inputs for i and j, and an irrelevant class k has no impact. For example, for Softmax([5, 7, 1]) and Softmax([5, 7, 10]), the resulting odds for the first two values (p_i/p_j) should be the same from both distributions, regardless of the third value.
Finally, if the "desired properties" approach is not satisfying, a more theoretical route for justifying the form of the softmax uses the framework of maximum entropy (E. T. Jaynes published this in 1957 to justify the Boltzmann distribution).
TL;DR, softmax is not a the only solution to mapping function of unnormalized values to a probability distribution, but it can be justified through axiomatic properties.
(1) one could say that the exponential shows up from the Boltzmann distribution, but then the same question applies.
(also, log(p) is not the formal definition of a logit)
relative_entropy(p, q) = cross_entropy(p, q) - entropy(p)
Information theory is some of the most accessible and approachable math for ML practitioners, and it shows up everywhere. In my experience, it's worthwhile to dig into the foundations as opposed to just memorizing the formulas.
(bits assume base 2 here)
> One of Bill Atkinson’s amazing feats (which we are so accustomed to nowadays that we rarely marvel at it) was to allow the windows on a screen to overlap so that the “top” one clipped into the ones “below” it. Atkinson made it possible to move these windows around, just like shuffling papers on a desk, with those below becoming visible or hidden as you moved the top ones. Of course, on a computer screen there are no layers of pixels underneath the pixels that you see, so there are no windows actually lurking underneath the ones that appear to be on top. To create the illusion of overlapping windows requires complex coding that involves what are called “regions.” Atkinson pushed himself to make this trick work because he thought he had seen this capability during his visit to Xerox PARC. In fact the folks at PARC had never accomplished it, and they later told him they were amazed that he had done so. “I got a feeling for the empowering aspect of naïveté”, Atkinson said. “Because I didn’t know it couldn’t be done, I was enabled to do it.” He was working so hard that one morning, in a daze, he drove his Corvette into a parked truck and nearly killed himself. Jobs immediately drove to the hospital to see him. “We were pretty worried about you”, he said when Atkinson regained consciousness. Atkinson gave him a pained smile and replied, “Don’t worry, I still remember regions.”
Pinterest’s Advanced Technologies Group (ATG) is an ML applied research organization within the company, focusing on large-scale foundation models (e.g. multimodal encoders, graph representation models, content embeddings, generative models, computer vision signals, etc.) that are deployed throughout the company. ATG is composed primarily of ML engineers and researchers, backed by a strong infrastructure team, and a small product prototyping + design team for deploying new AI/ML features in Pinterest. The organization is highly collaborative, research-driven, and delivers deep impact. The team is hiring for several engineering position
- iOS engineer for generative AI products: we are looking for senior or staff iOS engineers who have a track record of building fast prototyping work in the AI space — no deep machine learning domain expertise is required, but the ideal candidate would be comfortable interfacing with our ATG’s ML teams daily. An engineer in this role would be building entirely new features for Pinterest leveraging emerging technologies across LLMs, visual models, recommendation systems, and more.
- Computer vision domain specialist: we are looking for researchers or applied engineers with industry experience in the computer vision / visual-language modeling field (e.g. multimodal representation learning, visual diffusion models, visual encoders/decoders, etc.) We encourage the team to regularly publish, and the team works in a highly collaborative, research-driven environment, with full access to the Pinterest image-board-style graph for large-scale pre-training.
Please reach out to me directly (dkislyuk@pinterest.com) if you’re interested in either of these roles.
Additionally, the team is currently hiring for fall 2025 ML research internships for Master’s / PhD students, with opportunities to publish or to work on frontier models in the visual understanding and multimodal representation learning space: https://grnh.se/dad7c60e1us
So, self-information is uniquely defined by (1) assuming that information is a function transform of probability, (2) that no information is transmitted for an event that certainly happens (i.e. f(1) = 0), and (3) independent information is additive. h(x) = -log p(x) is the only set of functions that satisfies all of these properties.
> One of the anomalies in the history of mathematics is the fact that logarithms were discovered before exponents were in use.
One can treat the discovery of logarithms as the search for a computation tool to turn multiplication (which was difficult in the 17th century) into addition. There were previous approaches for simplifying multiplication dating back to antiquity (quarter square multiplication, prosthaphaeresis), and A Brief History of Logarithms by R. C. Pierce covers this, where it’s framed as establishing correspondences between geometric and and arithmetic sequences. Playing around with functions that could possibly fit the functional equation f(ab) = f(a) + f(b) is a good, if manual, way to convince oneself that such functions do exist and that this is the defining characteristic of the logarithm (and not just a convenient property). For example, log probability is central to information theory and thus many ML topics, and the fundamental reason is because Claude Shannon wanted a transformation on top of probability (self-information) that would turn the probability of multiple events into an addition — the aforementioned "f" is the transformation that fits this additive property (and a few others), hence log() everywhere.
Interestingly, the logarithm “algorithm” was considered quite groundbreaking at the time; Johannes Kepler, a primary beneficiary of the breakthrough, dedicated one of his books to Napier. R. C. Pierce wrote:
> Indeed, it has been postulated that logarithms literally lengthened the life spans of astronomers, who had formerly been sorely bent and often broken early by the masses of calculations their art required.
Pinterest’s Advanced Technologies Group (ATG) is hiring for an engineering position on our visual modeling team for developing Pinterest Canvas. Canvas is a foundation text-to-image model developed internally for helping various visualization, inpainting, and outpainting products. In this role, you’ll get to work with Pinterest’s rich visual-text dataset to build large-scale generative models which are continuously being shipped to production. The core Canvas pod is a small group (~6 engineers) inside of ATG, which focuses on a broad variety of AI/ML initiatives, such as core computer vision, multimodal representation learning, heterogeneous graph neural networks, recommender systems, etc.
New-grads are welcome to apply (preferably with a masters or PhD). Candidates should have diffusion modeling experience (e.g. diffusion transformers, LoRA fine-tuning, complex {text, image} conditioning, style transfer, etc.) and some form of industry experience. Engineers within ATG have a lot of leeway in terms of product contribution, so both ML engineers and research scientists are welcome to apply. We encourage the team to regularly publish, and the role can be either in person (SF, NY) or hybrid is preferred.
Please reach out to me directly (dkislyuk@pinterest.com) if you’re interested.
We’re looking for strong engineers to help us build consumer AI products within Pinterest’s Advanced Technologies Group (ATG), our in-house ML research division. You’d be working with a full-stack team of ML researchers and product engineers on projects that bring LLMs, diffusion models, and other core models in the generative multimodal ML and computer vision space to life inside the Pinterest product. Projects include assistants, new ways to search, restyling of boards / pins / rooms, and many other new applications. Your work will directly impact how millions of users experience Pinterest.
Tracks:
*iOS engineer*: You’ll craft beautiful and intuitive user experiences for our new AI products. Strong command of iOS and UI/UX craftsmanship required. Bonus points if you’re an opinionated product thinker with 0-1 mentality or have experience working with ML models. Please apply here: https://www.pinterestcareers.com/jobs/5426324/staff-ios-soft...
*Applied ML*: If you think you’d be a better fit as an applied ML or research engineer with an interest in directly translating research into user-facing products, feel free to contact me directly (@dkislyuk everywhere).
The ML and product engineering teams on ATG work directly together, along with design. The team consists of long-tenured employees who care deeply about both the quality of the Pinterest experience, and taking full advantage of the new capabilities emerging in the ML space over the last two years. ATG more broadly has spent the past decade+ bringing various ML technologies into the Pinterest ecosystem, and values publishing our work, building long-term infrastructure, and a collaborative and remote-friendly culture (though we do expect everyone to join company onsites a few times a year).
Unfortunately, ImageNet is not a useful benchmark for a while now since pre-training is so important for production visual foundation models.
If a new algorithm appears from with a novel approach (analog compute, heterogeneous computation graphs from genetic algorithms, quantum, much more...), there will be a whole generation of R&D + tool + framework building, which gives the major players enough time to adapt.
Maybe this time it's different, and maybe it's not, but that's why most recent robotics predictions fail to convince the ML industry broadly.
Of course, optimized autograd / autodiff is more parallelized than node-based message passing, but it's a useful model to start with.
At least in road and trail running, this doesn't line up with my observation. There are plenty of elite athletes, at least in the US, who are posting their training on Strava (Jim Walmsley, CJ Albertson, Molly Seidel, etc.) At the fast amateur level (e.g. those chasing the Olympic Trials Qualifier time), Strava adoption is pretty ubiquitous.
Out of the the athletes who are active on social media in the first place, those who don't post on Strava are usually doing it either for privacy reasons, or because they do not want to share their training plans for their competitors to copy.
Since the postscript mentions that the article was co-written by GPT-3, perhaps this is one of the generated lines? :)
(compared to the prior SOTA for generative ML in images i.e. GANs, diffusion models use far more compute-heavy and are much slower!)
I use it for every aspect of knowledge management, building a personal wiki, personal logging and writing, task tracking, reading notes, academic paper notes+metadata, planning, and more. Other tools offer similar features, but they all seem to have tradeoffs on data ownership or offline support or lack of extensibility or non-standard text format (i.e. not markdown). I wrote last year in another HN post that it's remarkable that the Obsidian team has delivered a superior product in a _very_ crowded note-taking / PKM space, and 18 months later it remains the single tool that I couldn't imagine abandoning.
Perhaps PKM as a cottage industry is overdone, and the author's point that everybody's system is highly personal and should not be replicated is valid. That said, Obsidian has allowed me to organize all of my reading, learning, writing, and thinking into a format that allows me to recall knowledge and make connections that I wouldn't have otherwise found. Seeing in what other contexts you have read about a concept or a historical figure or an essay, etc. is so valuable for life-long learning that I can't imagine going back to disconnected or separated note-taking.
Doing anything on Twitter scale has significant technical challenges.