453 karma · joined January 4, 2022
You should be excited because there's so many low-hanging classification problems (e.g. "Is this email spam?", "Is this yelp review happy or sad or neutral?") that a "cheap" LLM is overkill for. It really should be this cheap, and now Jev is the first to do it well and do it at scale.
If I was explaining it to my mom, I'd say "Classifying 1000 yelp reviews used to cost $50. Now it costs 5¢, at similar accuracy"
Its not that I don't like it as a game, but it's just so chock full of character.
The premise alone is so compelling; in the Land of the Dead, even dead men chase money. Even with the literal afterlife a train ticket away, you have conmen and good guys living (or dying?) like there's no end of the line.
Love it. One of those rare pieces of art that continues to live in my head rent-free.
I wish they had stayed this as the very first paragraph.
Maybe I'm too stupid to grok the original declaration. Does anyone have a better summation?
[0] https://github.com/mouseland/cellpose
This. I run small models (>50MB) for bio-imaging/biotech applications, it feels like every README implies that you need a discrete GPU to get started. While some do, many, especially the most useful ones, do not. Sure it matters if you're also going to do fine-tuning, but I believe your typical user just wants to detect some nuclei and get some cell-body ratios.
The laptop on your desk won't be running Meta's SAM, but it has more than enough compute to crunch 100's of your H&E slides overnight.
There's something wonderful about having my coworker walk up with an issue, and being able to bang out a Napari plugin that solves their exact problem before lunch.
My order of 'tool escalation' usually goes: - Can I solve their problem from napari's inline terminal? - Can I solve it with a one-off script? - Can I solve it with a one-off script that creates a one-off plugin interface? - Should I add the plugin to our company-wide repo since this problem seems to occur a lot?
There is an ever-growing bodycount in the NIO-GM graveyard [0], but I too hope that one day, it'll get figured out. My old roommate and good friend was T1 and monitoring one's glucose and remembering not to eat too much/little is half your life.
I'm not super familiar with JAX though.
I'd stick with the venv if it's heavy duty crunching and you do it often.
However, if this is a one off or doesn't need heavy compute, and you don't mind waiting a little longer, use the notebook.link.
Sure you could write Cython but then you have to have a build step and make wheels for every platform you and python version. Sure numpy has gotten faster over the years but you're still hampered by the GIL.
Numba is a little magic because you get to write stuff that feels like numpy, but get literal bytecode perf.
BUT there's a cost to this, which u learned the hard way when I imported a color map extension for matplotlib recently.
I thought i was going to be importing a couple megabytes at most. But Numba+llvmlite alone is almost 100MB!
This might be a drop in the bucket in some applications but for a color map library that has only two hot paths that need to be JITed, it's excessive.
Overall though, love this achievement, and i love what's being done for in-browser (aka local-first) scientific computing!
I just went back to dig up some old sources[0], and I can't believe this post is almost a decade old now. This guy's explainers and animations were leagues beyond any other resource I could find through Google searc at the time.
I think this is such a great reframing. It makes so much sense; I need an AI that acts more as a HUD and gives me superpowers, not just a copilot that can tell me when I've misspelled a word.
Not a huge deal since it's still cents per session, but my bigger issue was the weird change in tone. It became a lot more pretentious and over-explanatory.
Heavy prompt reworking helped but maybe that's just the cost of being better at coding and ARC-AGI?
Obviously a CTO is not going to walk away from the technology just because it's not good enough. That much more incentive for someone to create a powerful enough harness that can direct that power safely and productively. Like a nuclear core, we'll need to come up with the graphite rods and water tank. And if tokens are essentially free, why not, for every million tokens, spend 10x tokens on code review, testing, etc?
- Deepseek V4 Flash is impressively capable. Sonnet still beats it out by a thin margin, but the real kicker is that a typical session with Sonnet at current API costs is ~$2. The same session with Deepseek is 2 cents (ha). Its even allowed me to consider offering free-with-limits API usage on my own app. - Taalas (or competitors) have a lot going for them. If anything I feel like they need to join hands with these smaller model makers and converge in 2028
Maybe I'm thinking too far ahead but I'm hoping maybe I can cover the (future?) compute costs by charging for the heavy duty ML that would have to be run non-locally but IDK. This isn't critical yet but would love to hear your musings.