HNHacker News
TopNewBestAskShowJobs

npmipg

97 karma · joined October 24, 2023

submissionscomments
npmipg··on Anthropic, Meta, and Snap are paying up to 350k+ base for a DevRel
Devrel salaries appear to be growing very rapidly.

This was unheard of just a few years ago. It's possible it's because previously these roles would be called 'Head of Product Marketing,' and have just been rebranded as devrel.

npmipg··on Decoding an SF Craigslist "furniture" listing that appears to be a coded drug ad
lmao yes but the language is interesting
npmipg··on Fast Cheap Image Captioning Model Trained on Video Frames
We trained a frame captioning model that's 2x faster and 17x cheaper than Claude 4 Sonnet
npmipg··on OS 12B model that beats Claude 4 Sonnet at video captioning and is 17x cheaper
hf: https://huggingface.co/inference-net/ClipTagger-12b

blog post: https://inference.net/blog/cliptagger-12b

docs: https://docs.inference.net/use-cases/video-understanding

serverless API: https://inference.net/models/cliptagger-12b

npmipg··on GPU-rich labs have won: What's left for the rest of us is distillation
Note that distilling a general model is several orders of magnitude more expensive than distilling a task-specific model, which is what I'm trying to promote here. Smart general models make distilling great task specific models with no expert labelers way easier.
npmipg··on GPU-rich labs have won: What's left for the rest of us is distillation
Hey, I'm the author of the post.

The image has been fixed, and the point I'm making is that proprietary models are almost always ahead, and this gap is widening. OS models that are nearly at the same quality are usually distilled versions of proprietary models, or somehow get training data from them. Sometimes, after massive, expensive training runs models are open sourced anyway, and at some point that becomes unsustainable.

The difference between a top model and a model with a similar ELO might seem small, but the value of even a marginal increase in intelligence is extremely high--for example I only use the best coding model for coding, whatever the cost.

There's also lots of evidence that large labs are only getting started. In the past year, they have secured massive amounts of compute, which is still not utilized well. I expect lots of big training runs in the future, which will shift the gap further between OS and proprietary models.

The major problem for these companies is they spend hundreds of millions of dollars training a model, and then someone comes in the next day and distills something almost as good for far less money (still a VERY large sum of money.)

I don't know how this will be resolved long term.

npmipg··on Show HN: Spegel, a Terminal Browser That Uses LLMs to Rewrite Webpages
working on this as we speak!