HNHacker News
TopNewBestAskShowJobs

eurekin

1,034 karma · joined September 6, 2012

"thoughtful considerations about practical applications"

Self hosting, SBCs, AI/Vision/LLMs

submissionscomments
eurekin··on OpenAI still doesn't seem to have a handle on all of its rogue AI activity
I still can't fathom my company application security decision. Found pretty damning requests in our logs. Escalated. Expected it would result in at least reporting the TOS break from the originating place (one of cloud providers). Instead of that, they just went: "yeah, but we don't have logs". Provided them. "Yeah, but that ip doesn't resolve". I matched real ones from the load balancer. "There could just be many of them". There was one. When I had all the evidence gathered, they looked at me and finally told:

- It's just an Independent Security Researcher.

- So that's it? You will do no action?

- Correct

eurekin··on Plan mode is dead
It is? I still found it invaluable, when having a lot of externally managed knowledge. Basically the only gate before it goes on a certain failure, when dealing with proprietary libraries, tools and services
eurekin··on De-Brainrot Vacations
Brain fog? This might be a legitimate symptom worth seeing a doctor for
eurekin··on Why aren't smart people happier? (2022)
Have you ever gotten an IQ test?
eurekin··on LFM2.5 2.6B model competitive with 4x larger models
Care to share any details? I'm about to check the 2.6b lfm on document editing.
eurekin··on Bioengineered chewing gum may offer a way to fight HPV and other microbes
It does? I wonder, what it does to the gut biome then
eurekin··on Handbook.md shows that long policy documents do not reliably govern agents
The needle benchmarks show, that models extended context works for the part, that can be explained as: "I can access/adress that part of the input".

I have no idea, why in that context, the number of attention heads isn't mentioned. Models have a limited set of them and obviously, a model can focus at N max things at a time, which has to put an upper bound of long context support in some way. There's just more things to lose focus to (or, mismanage the limited attention heads resources - per token)

eurekin··on Qwen 3.8
I have a feeling this comment will make history
eurekin··on Qwen 3.8
It's a mcp, so connects quite easily to agents. With mcpo, I also connected it to open-webui (which has better support for OpenAI style tools/functions). Used it in claude code with that mcp plugin set-up too. Only ever used it for managing homelab information, but it met initial expectations. 27b is a great model, if grounded. The query about physical hosts and routing... I haven't found a single hallucination (altough Codex 5.6 as a reviewer mentioned something was wrong with some parts, and those were exactly the never properly documented ones. Codex/gpt had extra knowledge, because it was the conversation I used to set it up).
eurekin··on Qwen 3.8
With 3.6 27b, I just stopped changing local models and started tinkering with things on top (like mem0). Feels genuinely useful and more than a toy
eurekin··on EEG shows brain can simultaneous encode two speech streams
Also explains why we like music with two simultaneous distinct sections (bass + the rest). One without the other doesn't feel as complete
eurekin··on Shadcn/UI now defaults to Base UI instead of Radix
Disclaimer, I only use it to grow the "knowledge hub".

It's a single git project at my $USER home, that is referenced in global memory. It contains as much information about work things, as possible, to be productive.

I found that, if I allowed Claude to create the notes, it actually became more and more useful, but without the guideline, I just could barely get through reading it manually.

I'd never publish anything with such origin.

eurekin··on Shadcn/UI now defaults to Base UI instead of Radix
That's where the annotation in plannotator helps.

I'm asking to scan projects on gitlab, go through some docs to find more grounding material, write a subarticle (in the same style), scan logs on the test env, issue some curls, etc.; until the whole article is digestible - in the "backing knowledge graph" department.

eurekin··on Shadcn/UI now defaults to Base UI instead of Radix
Oh, I know exactly what you mean. I use plannotator with claude a lot and have much better time, since I asked for a specific styleguide.

I used "CD era MSDN reference and Raymond Chen blogging style" as a starting prompt for the styleguide and my work ability to digest AI plans raised a lot.

Couldn't recommend it more. Humble, insightful and respecting the reader

eurekin··on Marfa Public Radio Puts You to Sleep
Similarly, I never try to imagine a bright sunshine. It can wake me up
eurekin··on Ford AI hiccups push carmaker to rehire ‘gray beard’ inspectors
There never are. Those are going to be viewed as two discreet successful interventions.

One for lay-offs, because it was the best move at the time with the knowledge they had.

Second for quick correction, ability to pivot and execute quickly.

It's been always like that

eurekin··on The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
I remember seeing this popping up in discussions the first time, but never noticed any resolution (other than to train both sides). Has SOTA advanced?
eurekin··on GPT-5.5 hallucinates 3x more than MIT-licensed GLM-5.2
> A shift is happening among major AI labs, who are becoming increasingly skeptical of endless parameter count and training data scaling

I'm pretty sure it's mostly due to the training data quality. No idea, why this never gets mentioned in those discussions.

It was obvious right from the get go, that the scaling law just enabled some abilities, that were described by the underlying data and allowing the ANN to abstract it in the latent space.

eurekin··on Local Qwen isn't a worse Opus, it's a different tool
Noticed few cliffs. Sometimes it was a spurious stop (had to write "go on" or "continue" to restart), othertimes it was randomly saying: "Oh the user wants [the thing we already resolved]" and goes back in history. Cleared all out on fp16
eurekin··on Local Qwen isn't a worse Opus, it's a different tool
If I started today, with building a server, I'd jump right into verified set-ups and writeups, like this one:

https://github.com/noonghunna/club-3090

You can find info about running a patched version of vllm for 1x24gb, 2x and 4x. There's also quite a few "blackwell" subreddits, where people seem to share a lot of substantial information, if you're going the 6000 route.

eurekin··on Local Qwen isn't a worse Opus, it's a different tool
> The model is running so hot, that it shoots past the goal and starts looping

later:

> My latest experiment was setting up vLLM (the gold standard for production and concurrent serving) and even with an NVLink (175GBP) and tensor parallelism turned on, it was 3 tokens/second slower than llama.cpp during generation for an equivalent setup.

In all my tests, getting vllm to run is worth it. It was the single biggest thing, that helped for looping issues, agents going whack and losing focus on the task, long context being essentially useless.

FP8 model, unquantized cache in vllm an you have a league better overall experience, with any other stack I tested. Then, you can actually focus on using the model for other things and stop tinkering with settings.

eurekin··on CrankGPT
I'm still sour they had only one toast in, in a two slot toaster
eurekin··on Anthropic flies staff to D.C. to clean up White House fight
Also, supposedly biggest MMA event Rogan doesn't want to be on.

Something is up

eurekin··on RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
I kept getting recipes with "that one ingredient", which was either a major PITA to source or produced too much waste, even from a real world dietician consultation. Example, use 1/4th of a pumpkin for something. Those were good recipes, in terms of macronutrient composition, but doesn't work long term due to logistics.

I'm years after that strict diet needs, but that itch of fixing or easing some parts of the process stayed.

eurekin··on RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
Chrome driven by the OS accessibility API
eurekin··on RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
I keep finding more and more usecases for Q3.6 27b (same league) and the best performance is, when answers to my question is already in the context.

The moment I'm trying something open-ended or ambitious, Claude/ChatGPT clearly take you to the goal quicker.

For things, where there's a way to build a knowledgebase though, the local llm definitely can be a true contender. Plus, having a big context and no worries about filling it over and over - you can get quite far.

I'm writing this, literally in between cooking a pasta, that the local llm ordered products for me online. I've built a grocery shopping skill, so that it roughly knows what I have in fridge (losely), my last 10 representative orders (general preferences plus rich info about shops and skus around me) and actual real-time in stock info. The last part has been my personal pet peeve for every product that promised cooking ingredient delivery (that is not packaged specifically for that).

This is what has been promised to us by every big tech company with an agent, and now a local llms actually solved that for me fully.

eurekin··on Why AI hasn't replaced software engineers, and won't
Yeah, it won't. SWE here with a sidegoal to tackle the deployment side through various means (homelabbing, grabbing sre/cloud/observability tasks at work).

The biggest observable improvement in my post and pre ai development is the ability to tackle two projects at once, if the agent is on track. If not, and I have to do a deep dive to debug, it basically regresses to plain old everything like before.

eurekin··on Let's compile Quake like it's 1997
It feels now like an alternative timeline, one which performance optimisations were first and foremost still. Sometimes I fantasize, thinking how would our current development ecosystem look like, if we never abandoned the "be very vigilant with all resources you use" approach, that includes the whole webdev liftoff, where we ship a few hundred mb chromium engine for a dock app
eurekin··on DeepSeek 4 Flash local inference engine for Metal
Batching lowers that, since the model is read once from memory. Activation accumulation doesn't scale as nicely
eurekin··on PyInfra 3.8.0
On my homelab. It really feels like a dream come true for my usecase. No more puppet agents. No more declarative syntax, that you have to work around to do basic imperative ways. Or use a module, that stopped being maintained 3 years ago. Just plop a file here and there through ssh.
Page 1 of 21Next →