2,419 karma · joined January 24, 2014
I did side by side comparison with Gemini 2.5 Flash Lite, Jev, Jeff
I tried the 0.8B model, completely useless in classification. Qwen Jeff-Qwen3.5-2B was better, but still missed job type.
I suppose with larger model, this could be useful, but would require more ram and will be slower.
However, for this kind of customisation, Pi is actually quite great. One of the most things I love about Pi is ability to ask it to create an extension and it does it quite well as it’s part of their docs. Also ability to customise the system prompt to avoid the clutter that Claude Code add (around 20k system prompt that mostly had nothing to do with the code).
The demo was showing something I have created for my Pi setup, which is asking me in each new session which skills and MCP I would to enable for the session. This works quite well if you have multiple projects where you don’t need all skills but just a small subset
For your question, I assume you want people who can solve problems, can explain their thought process and reasons of decisions. I did interview devs recently, and my questions were related to their experience in their CV. Something like, tell me about project X in Y company, what did you do, what did you use, why, what did you learn from it.
My personal opinion is that LLMs are no difference from developers who over engineer and over complicate things. You can get them to do good work with the right steering and right understanding of the whole system within the context of the company.
It’s hard to justify several months to business when there is something off-shelf ready to use and doesn’t require domain specialists to run.
It’s a bit faster and bit cheaper, but this is compared to LLM. The consistency was nice to see, BUT, as someone who trained NLP models prior to LLMs, it’s just BERT with more data. I can see why people would want ready made one shot classifier, and I can see the value of sending multiple classifier in one call, but I wouldn’t call it breakthrough. And I believe many labs will replicate it in no time and might have it as part of their harness.
I see it as a wake up call for the tech community to go back to basics for most tasks instead of relying solely on generic LLMs.
I suppose you can balance between using them if the paycheck job is asking for them, and enjoy your craft outside your 9-5.
I’ve been reading more and more of this tone, and I do wonder what those devs writing these blogs do and where do they work. I’ve been a software engineer for more than 25 years, and the challenge was never the cording, it was always the human communication and understanding the outcome.
Yes LLMs can help none devs create things now, but show me a company that will allow AI generated code by none devs to go to prod.
Same for Google reviews.
I’m a software engineer, and the way I find new tools is usually YouTube, GitHub, and HN. Even those channels are manipulated now with bots to inflate numbers to reach more
> So it's not a skill issue?
I meant by offering sovereign cloud/inference, not for training. But even if it was for training, Chinese labs have limited supply of GPUs, look what they've done. So it is a skill issue.
Also to clarify, I didn't mean European engineers' skills, I meant "you get what you pay for" as a company, that's why I hope they start offering better compensation to retain talent.
I think the most of the money would go to purchase hardware, But I hope they can start making adequate compensation for AI engineers and researchers to move forward.
On the surface, AI showed results of £45 while normal results showed £39.99 and this particular result didn’t show at all in the AI search, even when clicking more.
However, small fine print was the £39.99 had a £4.99 delivery. Can it be that AI is optimised for full price including shipping?
I have a weird vibe from all the comments in this thread, they feel like a script rather a real experience.
A gentle reminder that none of Google models are in top 10 in Terminal Bench, so if this article is true, that would be a leap.
What a time to be alive!
That said, it’s really impressive to have a camera and all this tech in a toothbrush
1. Understanding user requirements/pain points is very important and many get it wrong.
2. Doing the architecture of the system within the org env/cloud/infra is often wrong by LLMs (for now).
3. Debugging when things go wrong. What to log, how to log it, and how to ensure not logging much or logging sensitive data.
4. Guide junior engineers, so they are not just accepting what LLMs are spiting.