I don't know that's a safe assumption tbh. Try throwing them some chapter 2 exercises from any category theory textbook.
334 karma · joined September 20, 2016
Applied type theorist gone haywire for GenAI-ish things. Previously a game-dev. https://runeblaze.github.io/
Reach me at `me ~at~ baqiaoliu.com`.
I don't know that's a safe assumption tbh. Try throwing them some chapter 2 exercises from any category theory textbook.
> The NS equations are far,far,far less meaningful. Like I mentioned earlier, if you actually want accurate CFD, you dont even use them.
Sure. Consider this: in algorithm research often the most optimal algorithm in big-O is not the one used IRL; examples are numerous: matrix multiplication, LCA data structures, many variants of shortest paths.
An academic can work two years on faster-in-theory matrix multiplication that no one expects to be used in practice (in our currently imaginable univese). Do you consider that less meaningful than working on faster matmul kernels?
like say if god lets me find a single counter example to P=NP and thus disproving it — I think we can learn tons about complexity theory from this counter example by studying it. we should not have the hubris of assuming “oh a single counterexample is generally useless” — why, how. this is the same hubris imo that produced like “number theory is useless” until it is not
we all have worked on software. we all know we aren't exactly the best people to point to ToS or how decisions are made higher up from us
don’t do that, it is weird, use “bruh” or “dude”
i guess you do? claude code is the commercial closed sourced version provides by ant. reading a mini version of vllm or sglang and then read codex source code or grok build source code will teach you all things taught by this article, fully in the open
it is like saying that you have no insights into some $commercial_db_system which is kinda true but imagine if the article is to teach you indices, query normalization, etc..
1. make distillation much harder
2. safety: prevent modifications to the thinking leading to injection attacks.
3. also honestly sometimes the model raw thoughts can be deranged and is not a good user experience (consider the varied audience in the market, etc.)
also often the mass underestimate/the model makers over-estimate how people love distilling models
> So basically... openrouter
:skull:
i now really wonder how many people of the public understood my thesis defense lol
> SigLIP 2
Maybe visual-semantic similarity is more appropriate? Nonetheless the design is fantastic
Some profs have a team of PhDs and things go to shit all the time. I don’t know why we expect $FRONTIER_LLM to do better
2. people often use openrouter for the sole purpose of using a unified chat completions API
3. OpenAI invented chat completions; if you use openrouter for chat completions often you can just switch your endpoint URL to point to the OAI endpoint to avoid the openrouter surcharge!
4. Hence anyone with large enough volume will very likely not use openrouter for OpenAI; there is an active incentive to take the easy route of changing the endpoint URL to OAI’s
With that said, the model is pretty good at it.
Let’s be realistic in our portrayal here.
I literally invoke sglang and vllm in Python. You are supposed to (if not using them over-the-network) use the two fastest inference engines there is via Python.
Disclaimer: I used to work at Adobe GenAI. Opinions are of my own ofc.
disclaimer: not expert, on top of my head
I can’t even make stuff with fundamental groups.