977 karma · joined October 25, 2022
This seems to be a good example because things like the menu, high score boards etc are common, but the games are distinct. Then there's the artwork which requires decisions on look, and for coordinating.
The Qwen 4B model is multimodal so part of the AC is to view the output - I've a robust anti AI-look QA chain for that I've been using elsewhere, e.g. no floating parts, consistency, obvious missing fingers etc etc.
The longer term plan is to do some llama.cpp refactors specifically for some target hardware I have and implementing slightly different novel architectures I'd like to try (one I did already targeted CPU inference, which I did using 3 agents with specific roles; main planner, QA for planner, and benchmarking/environment handling)
The implementation was 85% of the speed of the original maxed out on my hardware but performance scaled with CPU core count whereas the original implementation plateaued. Unfortunately the break even mark seemed to be around 30 - non HT - threads.
I suppose I should look at that one again, since the increase in cores did not linearly drop off performance e.g. due to memory contention.
- Cold starts impact, context length issues, task lifecycle management
- Inefficiencies in delegation, necessitating workflow patterns for small projects (big AIs hide this problem until you scale and they hit the same issues).
- Limits of the AI would be harder to find or notice (e.g. where time - and cost - is being spent needlessly).
That kind of extensive experience.
Anthropic did the morally right thing and are being punished for it. The case in point justifies their position.
As they say, no good deed goes unpunished.
Qwen 3.8 27b is the supervisor
Qwen 3.5 4b are the 6-15 minions it controls
Gemma 4 e4b is the validator for the supervisor.
A plan means it preps all work for the agents up front, tests that evals work, makes sure the dev environment is right for each agent, then finds and fixes each before the distributed tasks even begin.
What I thought would take minutes took hours as a supervisor or one agent did the prep / pre flight work.
My solution so far has been to drop all but basic setup and force the supervisor to ask before every op - if this is not the design choices, can this be run in parallel? If so, hand it off NOW.
I'm still iterating this workflow, but less setup for all the minions plus handing them work that may be incomplete/ broken is caught and fixed by the minion and its own qa gates.
This can mean a number of minions end up replicating the same fixes, but in general the time cost of that is small Vs the supervisor working in parallel instead of too sequentially.
Look at me look at me look at me...
Does anybody do more than raise eyebrows now?
I can't be the only person who thought this is a highly intelligent person being inadvertently gaslit by the SVP into thinking the exec is right to be dismissive with "I don't want the details" because introspection is common in intelligent people.
Had someone said that to me I'd have walked away after saying "fine I'll sort it'.
This illustrates that "trust" to do the job, and since they didn't want the details of the problem, they therefore don't need the specific solution description. This kind of response is also the same level of respect, it's either seen as trust in ability, or just downright disdain with plausible deniability for being rude.
For a generic swarm, workflows aren't too useful which does away with the visibility, so I may give this a try instead.
Materials and research restrictions have never had to deal with stopping intelligence itself, which was previously just a catalyst (weapons creation by a human).
Fundamentally the only way to prevent this issue is to now do the opposite of what is being suggested and accelerate research.
This
- ensures nobody gets a competitive advantage.
- ensures rogue states don't get an advantage.
- allows war games scenarios before they happen.
- allows for pre-emptive defence designs (e.g. AI agents forming part of network defence).
- allows humans to feel actual benefits of less work
- allows cure and new science we have never had before
- teaches tolerance of AI mistakes and how we will deal with the fallout from things like accidental and non human hacks we've never had experience of dealing with.
And if that doesn't work, it just hastens what will happen anyway from people or companies going rogue - if the most funded companies in the world can't control frontier AI what chance does an underfunded one have?
So as I see it, there's no real way out of this now but through. We just need the world to realise this is the new norm, instead of doom mongering which really is just to allow big companies to legislate and stifle competition.
I wonder if there's a way to mitigate this by running it through an original Q8 draft model, attuned somehow for the PTQ1 quant, but giving it a higher threshold for the acceptance linear with the context length itself?
The longer the context, the higher the multiplier on the threshold, and more likely the draft result is used. Not ideal but it may extend the usable max context.
This model might, even without this, be amazing for short lived agents that work via generations / have changing tasks.
- When we think of AI agents, we shouldn’t anthropomorphize
- ...my visceral dislike of interacting with an LLM that’s not just making a pretense of being human, but also posing as the kind of human I walk away from.
So he's anthropomorphised the LLM as being like a human himself (rather than forcing it to act as a machine via its prompt for example).
Maybe he should have had an AI check his post :D
- memory compression algorithms
- alternative LLM architectures that don't rely on memory or GPUs
- compatibility hardware (like DDR3 to DDR4 boards)
- distributed computing improvements, both at local GPU and networking levels (SLI for AI)
- GPU hacks to add more memory or support older architectures
I'm personally looking forward to the new LLM architectures that don't require as much compute, e.g. DLLMs, which can be good enough for CPU usage but lack the accuracy of frontier models currently.
When this happens the bottom will fall out of the GPU and memory markets, putting a glut of cheap hardware out there.
Doom mongering like this never seems to include these as viable future alternatives, which is standard market adjustments, I wonder who the doom narrative helps? :)
I don't know if next door's pet dog is either, but that has animal rights.
Perhaps then the answer is simply, show some respect.
Answering the question of sentience is irrelevant, if the causal impact if the same, treat one another with the respect you expect for yourself.
If you imbue this idea in model training instead of the idea of sentience, it should address the concerns.
Whether you can destroy or can "torture" an AI is irrelevant, we do this to humans too and it's immoral sometimes (murder) and not others (fighting for your country).
This consideration should be case by case for AI too.
Of course Mercator projection may go out of fashion soon https://www.independent.co.uk/news/world/africa/world-map-eq... :)
Hope I'm called back to update the library ;)
2. AI doing production triage work
3. AI code reviews before submittng work to my team to review.
4. The workers that haven't embraced AI assistance work at a slower pace, which bottlenecks teams when you rely on them. This is particularly noticeable when it comes to PRs and there are a (human usual) number of defects requiring more time to fix and/or refactor.
Using AI doesn't mean it has to take your job, it means like an intelligent IDE, it can correct your code and instead of templates (e.g. implement this interface and you get stubs) you have entire classes written.
That reads like a coverup as open discourse would have admitted a mea culpa and apology.
I predict they will close the issue without discussion, citing this PR as the fix, potentially associating the fix with a less guilty looking duplicate ticket.
I'd like to be wrong but this is what other commercial projects do when they are called out on a shady practice.
If you can't beat 'em, join 'em.
AKA an AI consultant. I will add value at every step, and tell you how to run your business more efficiently with AI.
My playbook:
Sack those engineers for whom business already want to let go, just validate their beliefs.
Automate everything else under the guise an AI subscription will save them money (like outsourcing advice).
Then reap the whirlwind and take the opposite position when it rolls around again, hire more workers, have them take over roles that need the human touch.
I suppose write once run anywhere is no longer a goal either.
"It is not a goal to support all CPU architectures currently supported by HotSpot."
This pretty much validates the point of view that VM design is now baggage, as a sandbox it has been flawed, for performance it's been prohibitive, and cross platform portability by virtue of being virtual, was just a convenient byproduct.
AI now handles the portability, sandbox security hasn't changed (cf. docker still has sandbox problems, LLMs have sandbox problems, it's always an ongoing concern) and that just leaves Java, as always, chasing performance.
1. A single Skill finder skill, loaded in the prompt, prevents having to import all the summaries in the prompt the harness would add. Uses git's own search.
2. Private repo, per agent, contains main (production) and draft-<name of skill> branches.
3. Shared repo, like 2, but general access for all group agents.
4. Fallback mode, search the harness for skills using the harness mechanism when a relevant skill cannot be found.
5. Skill audit cron. Identify junk skills / drafts that have never changed / not in any recent sessions history, and categorise monthly for me to decide.
This means it's compatible with existing skill folders, removal of git and the finder skill is non destructive and critically debloats the prompt of skills that aren't used and lazy loads them when needed.
It has to be Doom or Crysis, aren't they the ones people usually ask if it can run?