2,072 karma · joined April 24, 2015
Some of my tests do point towards what this post says. I was not successful in getting it's score to align with an existing rubric I had. It 'worked' but it was off and compressed from where I wanted it to be. Not a bad starting point, but I couldn't get it to move to where I intended the rubric to be. It wasn't the most robust test and I didn't spend a lot of time trying, but it wasn't just instantly magical.
However, it does seem genuinely useful just by being fast and cheap and good enough, so I am still a bit hyped.
I ran it through MMLU a few days ago and it scored ~90% so seems to have a lot of general world knowledge trained in. Makes me think your speculation is right. I have some credits left, might try and think of an experiment. I saw a gist where someone was asking it which model it was and it was picking qwen a lot, but who knows...
Anyway, thank you for the interesting discussion!
In your example, I would expect an LLM to do fine and if you have access to the raw logits you can measure whether or not it was confused and assign a confidence to the answer it gave.
I do think that Jev handles more than this though and, in my early testing, does things that are not easily accomplished with guided decoding techniques.
I prefer Cursor at this point just because of Composer. Claude Code is passable but the lack of a good, fast, cheap workhorse sucks. Sonnet and Haiku aren’t it.
I also do like Codex and Sol, Terra, and Luna. They are decent but I don’t find they stand out enough to use over the others.
Additionally, I have not tried Astra and found Fable to really not worth the cost for the tasks I do.
Finally, the latest Grok is actually a beast of a model, but expensive enough to not be a stand out.
Recently, I made a dumb little app for my kids and decided to try marketing it on social media just to see what it is like. It is fascinating in a sense and disheartening as well. I have been very unsuccessful, but the most signal tends to come from the dumbest content I have tried.
In doing this, I have come into contact with the social media feeds I never felt the need to look at and man… they are like a drug. I find myself mesmerized by random IG reels. It is one thing to understand what they are on an intellectual level and a totally different to feel it first hand.
I miss MySpace.
Some of my use cases are very latency sensitive. What sort of overhead are you seeing?
So much content is just straight copy/pasted from the LLM now. Articles, blog posts, linked in posts, reddit comments, etc. Even just using the LLM for 'editing' tends to shift the voice to an obvious LLM voice when used naively. It is getting worse too. Last week a co-worker sent me a screenshot of Claude for me to review their "work", which was just whatever Claude made up.
Usually, if something is very obviously unfiltered LLM output, I just stop reading.
I do use LLMs for writing myself. They are useful, but are poor authors.
When I was in school, decades ago now, very few people went into CS compared to other majors. Everyone I knew going into it did it because they loved it. I would have done it regardless of the career opportunities because I want to build stuff.
Interviewing candidates over the years since then, my experience has been there are still very few of those passionate nerds and a lot of people who did it for other reasons, like the money or similar. There is nothing inherently wrong with this. I don’t fault people for it.
Maybe if we get very lucky, it will go back to a relatively few passionate people building stuff because it is cool?
Something like this was always inevitable. I just hope it doesn’t ruin a good thing.