2,808 karma · joined June 13, 2012
About the Superposition paper - this is close to what I've been thinking about over the past week. I'm thinking that concepts or choices in a "superposition" are harder for a fully-differentiable neural net to reason about. For example, if there's a "green" vs "purple" choice to be made, it can't fully commit to either (especially if they're 50-50), and will have to reason about both simultaneously (difficult due to nonlinear manifold space). Discretizing to tokens (non-differentiable argmax) forces a choice, and that allows it to reason about a single concept separately and easier.
But with such scales, low sample rates and averaging are key.
Anything HPC will benefit from thinking about how things map onto hardware (or, in case of SQL, onto data structures).
I think way too few people use profilers. If your code is slow, profiling is the first tool you should reach for. Unfortunately, the state of profiling tools outside of NSight and Visual Studio (non-Code) is pretty disappointing.
* shadertoy - in-browser, the most popular and easiest to get started with
* Shadron - my personal preference due to ease of use and high capability, but a bit niche
* SHADERed - the UX can take a bit of getting used to, but it gets the job done
* KodeLife - heard of it, never tried it
I've mentioned it before, but I'd love for sparse operations to be more widespread in HPC hardware and software.
I don't have a good idea of what happened inside or what they could have done differently, but I do remember them going from a world-leading LLM AI lab to selling embeddings to enterprise.
I'm working on a few wide gamut art pieces, and so far the test prints have been less than stellar. Disclaimer - I'm an amateur in this field.
I'm all for a good Acrobat alternative though.