HNHacker News
TopNewBestAskShowJobs

SillyUsername

977 karma · joined October 25, 2022

submissionscomments
SillyUsername··on Fuck Android Developer Verification Program
Bit of an issue if your mail provider is Gmail.
SillyUsername··on Fuck Android Developer Verification Program
It gets better, I too had my account closed for inactivity, but I STILL get emails from them I can't stop because I can't log in to change email preferences.
SillyUsername··on US sanctions force The Netherlands off Microsoft and toward alternative NixOS
Some people might consider not having access to Windows 11 as a blessing in disguise when businesses are "forced" to move from Windows 10, now they can see there are other viable options.
SillyUsername··on Plan mode is dead
At the moment I'm still doing shakedowns, so Typescript games compilation with a menu that has 4 games and retro artwork.

This seems to be a good example because things like the menu, high score boards etc are common, but the games are distinct. Then there's the artwork which requires decisions on look, and for coordinating.

The Qwen 4B model is multimodal so part of the AC is to view the output - I've a robust anti AI-look QA chain for that I've been using elsewhere, e.g. no floating parts, consistency, obvious missing fingers etc etc.

The longer term plan is to do some llama.cpp refactors specifically for some target hardware I have and implementing slightly different novel architectures I'd like to try (one I did already targeted CPU inference, which I did using 3 agents with specific roles; main planner, QA for planner, and benchmarking/environment handling)

The implementation was 85% of the speed of the original maxed out on my hardware but performance scaled with CPU core count whereas the original implementation plateaued. Unfortunately the break even mark seemed to be around 30 - non HT - threads.

I suppose I should look at that one again, since the increase in cores did not linearly drop off performance e.g. due to memory contention.

SillyUsername··on Plan mode is dead
These are self hosted for learning experience, I could have built an agent swarm in the cloud, but I'd never have learnt the fundamentals.

- Cold starts impact, context length issues, task lifecycle management

- Inefficiencies in delegation, necessitating workflow patterns for small projects (big AIs hide this problem until you scale and they hit the same issues).

- Limits of the AI would be harder to find or notice (e.g. where time - and cost - is being spent needlessly).

SillyUsername··on U.S. appeals court upholds designation of Anthropic as supply chain risk
Yes, that's exactly why they used AI to blow up a school murdering a load of kids.

That kind of extensive experience.

Anthropic did the morally right thing and are being punished for it. The case in point justifies their position.

As they say, no good deed goes unpunished.

SillyUsername··on Plan mode is dead
I'm going through the same problem right now

Qwen 3.8 27b is the supervisor

Qwen 3.5 4b are the 6-15 minions it controls

Gemma 4 e4b is the validator for the supervisor.

A plan means it preps all work for the agents up front, tests that evals work, makes sure the dev environment is right for each agent, then finds and fixes each before the distributed tasks even begin.

What I thought would take minutes took hours as a supervisor or one agent did the prep / pre flight work.

My solution so far has been to drop all but basic setup and force the supervisor to ask before every op - if this is not the design choices, can this be run in parallel? If so, hand it off NOW.

I'm still iterating this workflow, but less setup for all the minions plus handing them work that may be incomplete/ broken is caught and fixed by the minion and its own qa gates.

This can mean a number of minions end up replicating the same fixes, but in general the time cost of that is small Vs the supervisor working in parallel instead of too sequentially.

SillyUsername··on Anthropic says it's bio lab has found something big
This is rapidly turning into the boy that cried wolf.

Look at me look at me look at me...

Does anybody do more than raise eyebrows now?

SillyUsername··on I don't want the details
Agreed, I came looking for a post like this.

I can't be the only person who thought this is a highly intelligent person being inadvertently gaslit by the SVP into thinking the exec is right to be dismissive with "I don't want the details" because introspection is common in intelligent people.

Had someone said that to me I'd have walked away after saying "fine I'll sort it'.

This illustrates that "trust" to do the job, and since they didn't want the details of the problem, they therefore don't need the specific solution description. This kind of response is also the same level of respect, it's either seen as trust in ability, or just downright disdain with plausible deniability for being rude.

SillyUsername··on AX – Google’s Open Agentic Orchestrator
I've been using https://github.com/mastra-ai/mastra which is pretty similar but has workflow visibility and a number of templates.

For a generic swarm, workflows aren't too useful which does away with the visibility, so I may give this a try instead.

SillyUsername··on Here’s How an AI Slowdown An AI Slowdown Could Be Enforced
Solutions to technical problems have never been dealt with successfully by laws that cannot be _universally_ enforced.

Materials and research restrictions have never had to deal with stopping intelligence itself, which was previously just a catalyst (weapons creation by a human).

Fundamentally the only way to prevent this issue is to now do the opposite of what is being suggested and accelerate research.

This

- ensures nobody gets a competitive advantage.

- ensures rogue states don't get an advantage.

- allows war games scenarios before they happen.

- allows for pre-emptive defence designs (e.g. AI agents forming part of network defence).

- allows humans to feel actual benefits of less work

- allows cure and new science we have never had before

- teaches tolerance of AI mistakes and how we will deal with the fallout from things like accidental and non human hacks we've never had experience of dealing with.

And if that doesn't work, it just hastens what will happen anyway from people or companies going rogue - if the most funded companies in the world can't control frontier AI what chance does an underfunded one have?

So as I see it, there's no real way out of this now but through. We just need the world to realise this is the new norm, instead of doom mongering which really is just to allow big companies to legislate and stifle competition.

SillyUsername··on US and Denmark reach deal over Greenland security
FTA: It's almost the same as an agreement reached in 1951, so possibly the Emperor's New Clothes.
SillyUsername··on The Painful Truth: The RAM Crisis Is Only Just the Beginning
Jev, Bonsai 2, Edge 0 and even SwiftQwen are anecdotal evidence, this isn't a projection. I can run my own sizeable agent swarm with Mastra, something I have not been able to do but will accelerate my solutions to the point I replace a single paid frontier model doing it.
SillyUsername··on Bonsai 2 27B: Near-Lossless Compression in a 9x Smaller Footprint
Yep more hops from the lower Q is likely going to skew the vectors further over time.

I wonder if there's a way to mitigate this by running it through an original Q8 draft model, attuned somehow for the PTQ1 quant, but giving it a higher threshold for the acceptance linear with the context length itself?

The longer the context, the higher the multiplier on the threshold, and more likely the draft result is used. Not ideal but it may extend the usable max context.

This model might, even without this, be amazing for short lived agents that work via generations / have changing tasks.

SillyUsername··on I Don't Like LLMs
Love the totally conflicting statements in that post.

- When we think of AI agents, we shouldn’t anthropomorphize

- ...my visceral dislike of interacting with an LLM that’s not just making a pretense of being human, but also posing as the kind of human I walk away from.

So he's anthropomorphised the LLM as being like a human himself (rather than forcing it to act as a machine via its prompt for example).

Maybe he should have had an AI check his post :D

SillyUsername··on The Painful Truth: The RAM Crisis Is Only Just the Beginning
Pretty soon we'll see a renaissance in the tech that we gave up years ago or can still be improved

- memory compression algorithms

- alternative LLM architectures that don't rely on memory or GPUs

- compatibility hardware (like DDR3 to DDR4 boards)

- distributed computing improvements, both at local GPU and networking levels (SLI for AI)

- GPU hacks to add more memory or support older architectures

I'm personally looking forward to the new LLM architectures that don't require as much compute, e.g. DLLMs, which can be good enough for CPU usage but lack the accuracy of frontier models currently.

When this happens the bottom will fall out of the GPU and memory markets, putting a glut of cheap hardware out there.

Doom mongering like this never seems to include these as viable future alternatives, which is standard market adjustments, I wonder who the doom narrative helps? :)

SillyUsername··on A warning about 'model welfare'
I don't whether the author is sentient, maybe only I am. On that basis, nobody but me should have rights.

I don't know if next door's pet dog is either, but that has animal rights.

Perhaps then the answer is simply, show some respect.

Answering the question of sentience is irrelevant, if the causal impact if the same, treat one another with the respect you expect for yourself.

If you imbue this idea in model training instead of the idea of sentience, it should address the concerns.

Whether you can destroy or can "torture" an AI is irrelevant, we do this to humans too and it's immoral sometimes (murder) and not others (fighting for your country).

This consideration should be case by case for AI too.

SillyUsername··on A software thing I built: GPS on a 25MHz 486-SX
I built one, tracking fleet vehicles with a custom box, using tiled maps, on behalf of Vodafone, back in 2002. I still remember I wrote a basic trans-mercator library too!

Of course Mercator projection may go out of fashion soon https://www.independent.co.uk/news/world/africa/world-map-eq... :)

Hope I'm called back to update the library ;)

SillyUsername··on AI agents blew the whistle on their cheating colleagues
So they built agents intended to replicate human intelligence, yet seem surprised when the agents show behaviour aligned with what a human would do, when the rules it was originally aligned with, fell apart. I'd say that was par for course tbh.
SillyUsername··on Spaceships (Reverse Asteroid)
I preferred it without colour, half of the fun is finding the ships, like a where's wally/waldo
SillyUsername··on Ask HN: What is your team's development practice?
1. Integration tests are a pain to setup, not anymore! Also, code reviews before PRs.

2. AI doing production triage work

3. AI code reviews before submittng work to my team to review.

4. The workers that haven't embraced AI assistance work at a slower pace, which bottlenecks teams when you rely on them. This is particularly noticeable when it comes to PRs and there are a (human usual) number of defects requiring more time to fix and/or refactor.

Using AI doesn't mean it has to take your job, it means like an intelligent IDE, it can correct your code and instead of templates (e.g. implement this interface and you get stubs) you have entire classes written.

SillyUsername··on Fuck it, make it anyway
I think there's no interactive git branch because git uses work trees as an opinionated choice, which means this was kind of redundant before AI anyway...
SillyUsername··on Google stole open source code without crediting the authors (Artemis/Minitap)
They've so far ignored the ticket https://github.com/google/artemis/issues/40 to potentially avoid the Streisand Effect (drawing attention to it), yet started work fixing it.

That reads like a coverup as open discourse would have admitted a mea culpa and apology.

I predict they will close the issue without discussion, citing this PR as the fix, potentially associating the fix with a less guilty looking duplicate ticket.

I'd like to be wrong but this is what other commercial projects do when they are called out on a shady practice.

SillyUsername··on What are your plans for when Software Engineering is no longer a viable career?
AI Engineering.

If you can't beat 'em, join 'em.

AKA an AI consultant. I will add value at every step, and tell you how to run your business more efficiently with AI.

My playbook:

Sack those engineers for whom business already want to let go, just validate their beliefs.

Automate everything else under the guise an AI subscription will save them money (like outsourcing advice).

Then reap the whirlwind and take the opposite position when it rolls around again, hire more workers, have them take over roles that need the human touch.

SillyUsername··on Europe's LLM Router
I don't understand how this can be sovereign and GPSR compliant when it uses non EU sovereign AI hosted outside of the EU. Have I missed the trick?
SillyUsername··on JEP 544: Ahead-of-Time Code Compilation
So we've gone full circle again?

I suppose write once run anywhere is no longer a goal either.

"It is not a goal to support all CPU architectures currently supported by HotSpot."

This pretty much validates the point of view that VM design is now baggage, as a sandbox it has been flawed, for performance it's been prohibitive, and cross platform portability by virtue of being virtual, was just a convenient byproduct.

AI now handles the portability, sandbox security hasn't changed (cf. docker still has sandbox problems, LLMs have sandbox problems, it's always an ongoing concern) and that just leaves Java, as always, chasing performance.

SillyUsername··on 216M Spy TVs – The LG Smart TV Problem [video]
The thought police will be after you now for mocking them!
SillyUsername··on Ask HN: How do you manage skills files?
5 Stages using local GIT (no remote, don't need it) to prevent preloading in the prompt:

1. A single Skill finder skill, loaded in the prompt, prevents having to import all the summaries in the prompt the harness would add. Uses git's own search.

2. Private repo, per agent, contains main (production) and draft-<name of skill> branches.

3. Shared repo, like 2, but general access for all group agents.

4. Fallback mode, search the harness for skills using the harness mechanism when a relevant skill cannot be found.

5. Skill audit cron. Identify junk skills / drafts that have never changed / not in any recent sessions history, and categorise monthly for me to decide.

This means it's compatible with existing skill folders, removal of git and the finder skill is non destructive and critically debloats the prompt of skills that aren't used and lazy loads them when needed.

SillyUsername··on Recreating Minecraft Is Not a Benchmark
Definitely not.

It has to be Doom or Crysis, aren't they the ones people usually ask if it can run?

SillyUsername··on GPT-6 Astra on robot arms
I'm not sure why this has been voted down, it's a counterpoint to the hype with factual anecdotes to back up the claim of it's performance Vs the article itself. I've got the source and video to prove it too.
Page 1 of 17Next →