105 karma · joined December 30, 2024
https://mlepath.com/
Basically, do not carelessly use any similarity metric.
The effectiveness of coding assistant is directly proportional to context length and the open models you can run on your computer are usually much smaller. Would love to see something more quantified around the usefulness on more complex codebases.
On production side of "AI" (I don't love the term being thrown around this loosely as true AI should include planning, etc not just inference) the only question is how well do you solve the one problem that's in front of you. In most business usecases today that problem is narrow.
LLMs drive a minuscule (but growing) amount of value today. Recommender systems drive huge amount of value. Recommender systems are very specialized.
For these models probably no. But for proprietary things that are mission critical and purpose-built (think Adobe Creative Suite) the calculus is very different.
MS, Google, Amazon all win from infra for open source models. I have no idea what game Meta is playing
I think thinking about it IS useful. Problems like value alignment are really difficult. No one alive knows what to do about value alignment and even if it takes us 500 years to reach real singularity it may take that long to solve value alignment problem
DEI always seemed like an activity they did for show. This changes nothing honestly.
Do you see any artifacts from having trained on synthetic data? Is there a natural benchmark dataset (real tables in the wild)?
In my experience synthetic data can only take you so far, it has all the quirk the dataset creator can think of but the real value is usually in patterns they cannot. Vision took a huge leap forward with ImageNet dataset release
Interested to know how debugging in a real application would work since WASM is pretty hard to debug and GPU code is pretty hard to debug. I assume WASM GPU is ... very difficult to debug.
I find chat for search is really helpful (as the article states)
I think we came to any metric that is used for rewards loses its value as a metric. That includes citations, sadly.
We were lucky enough to grow up with the industry and progressively learn more complexity. The kids out of school today are faced with decades worth of complexity on day one on the job.
Basically if something appeared online or was transmitted over the wire should no longer be eligible to evaluate on. D. Sculley had a great talk at NeurIPS 2024 (same conference this paper was in) titled Empirical Rigor at Scale – or, How Not to Fool Yourself
Basically no one knows how to properly evaluate LLMs.