At some point, this is a distributed system of agents.
Once you go from 1 to 3 agents (1 router and two memory agents), it slowly ends up becoming a performance and cost decision rather than a recall problem.
4,101 karma · joined August 20, 2012
https://aidnn.ai
At some point, this is a distributed system of agents.
Once you go from 1 to 3 agents (1 router and two memory agents), it slowly ends up becoming a performance and cost decision rather than a recall problem.
Let's pick a simpler compression problem where changing the frame of reference improves packing.
There's a neat trick in the context of floating point numbers.
The values do not always compress when they are stored exactly as given.
[0.1, 0.2, 0.3, 0.4, 0.5]
Maybe I can encode them in 15 bytes instead of 20 as float32.
Up the frame of reference to be decibels instead of bels and we can encode them as sequential values without storing exponent or sign again.
Changing the frame of reference, makes the numbers "more alike" than they were originally.
But how do you pick a good frame of reference is all heuristics and optimization gradients.
Famous tweet about conversations with God.
[1] - https://x.com/WraithLaFrentz/status/1981404849305686219
This is a real problem when the "direction" == "good feedback" from a customer standpoint.
Before we had a product person for every ~20 people generating code and now we're all product people, the machines are writing the code (not all of it, but enough of it that I will -1 a ~4000 line PR and ask someone to start over, instead of digging out of the hole in the same PR).
Feedback takes time on the system by real users to come back to the product team.
You need a PID like smoothing curve over your feature changes.
Like you said, Speed isn't velocity.
Specifically if you have a decent experiment framework to keep this disclosure progressive in the customer base, going the wrong direction isn't a huge penalty as it used to be.
I liked the PostHog newsletter about the "Hidden dangers of shipping fast", I can't find a good direct link to it.
The model is not great, but it was the "least amount of setup" LLM I could run on someone else's machine.
Including structured output, but has a tiny context window I could use.
Ira Glass has a nice quote which is worth printing out and hanging on your wall
Nobody tells this to people who are beginners, I wish someone told me. All of us who do creative work, we get into it because we have good taste. But there is this gap. For the first couple years you make stuff, it’s just not that good. It’s trying to be good, it has potential, but it’s not. But your taste, the thing that got you into the game, is still killer. And your taste is why your work disappoints you. A lot of people never get past this phase, they quit. Most people I know who do interesting, creative work went through years of this. We know our work doesn’t have this special thing that we want it to have. We all go through this. And if you are just starting out or you are still in this phase, you gotta know its normal and the most important thing you can do is do a lot of work.
Or if you're into design thinking, the Cult-of-Done[1] was a decade ago.
[1] - https://medium.com/@bre/the-cult-of-done-manifesto-724ca1c2f...
That is the part of the post that stuck with me, because I've also picked up impossible challenges and tried to get Claude to dig me out of a mess without giving up from very vague instructions[1].
The effect feels like the Loss-Disguised-As-Win feeling of the video-games I used to work on at Zynga.
Sure it made a mistake, but it is right there, you could go again.
Pull the lever, doesn't matter if the kids have Karate at 8 AM.
I think the eggs aren't dividing as you age (you are born with them, so to speak) and the sperm is held "outside" the body.
One is in original packaging and the other is produced in a "cooler" enviroment by the billions with a heavy QA failure of 99.9999%.
Every strategy which worked with an off-shore team in India works well for AI.
Sometime in mid 2017, I found myself running out of hours in the day stopping code from being merged.
On one hand, I needed to stamp the PRs because I was an ASF PMC member and not a lot of the folks who were opening JIRAs were & this wasn't a tech debt friendly culture, because someone from LinkedIn or Netflix or EMR could say "Your PR is shit, why did you merge it?" and "Well, we had a release due in 6 days" is not an answer.
Claude has been a drop-in replacement for the same problem, where I have to exercise the exact same muscles, though a lot easier because I can tell the AI that "This is completely wrong, throw it away and start over" without involving Claude's manager in the conversation.
The manager conversations were warranted and I learned to be nicer two years into that experience [1], but it's a soft skill which I no longer use with AI.
Every single method which worked with a remote team in a different timezone works with AI for me & perhaps better, because they're all clones of the best available - specs, pre-commit verifiers, mandatory reviews by someone uncommitted on the deadline, ease of reproducing bugs outside production and less clever code over all.
For my kid who uses a Chromebook right now, Magsafe would've been improvement in how often the power cable pulls the it off the desk.
But otherwise, this checks all the boxes, including applecare.
It was easier to beat Parquet's defaults - ORC+zlib seemed to top out around the same as the default in this paper (~178Gb for the 1TB dataset, from the hadoop conf slides).
We got a lot of good results, but the hard lesson we learned was that scan rate is more important than size. A 16kb read and a 48kb read took about the same time, but CPU was used by other parts of the SQL engine, IO wasn't the bottleneck we thought it was.
And scan rate is not always "how fast can you decode", a lot of it was encouraging data skipping (see the Capacitor paper from the same era).
For example, when organized correctly, the entire l_shipdate column took ~90 bytes for millions of rows.
Similarly, the notes column was never read at all so dictionaries etc was useless.
Then I learned the ins & outs of another SQL engine, which kicked the ass of every other format I'd ever worked with, without too much magical tech.
Most of what I can repeat is that SQL engines don't care what order rows in a file are & neither should the format writer - also that DBAs don't know which filters are the most useful to organize around & often they are wrong.
Re-ordering at the row-level beats any other trick with lossless columnar compression, because if you can skip a row (with say an FSST for LIKE or CONTAINS into index values[1] instead of bytes), that is nearly infinite improvement in the scan rate and IO.
We're so deep in this hole that people are fixing this on a CPU with silicon.
The Graviton team made a little-endian version of ARM just to allow lazy code like this to migrate away from Intel chips without having to rewrite struct unpacking (& also IBM with the ppc64le).
Early in my career, I spent a lot of my time reading Java bytecode into little endian to match all the bytecode interpreter enums I had & completely hating how 0xCAFEBABE would literally say BE BA FE CA (jokingly referred as "be bull shit") in a (gdb) x views.
The original take was that "We need to make tools which an AI can hold, because they don't have fingers" (like a quick-switch on a CNC mill).
My $job has been generating code for MCP calls, because we found that MCP is not a good way to take actions from a model, because it is hard to make it scriptable.
It definitely does a good job of progressively filling a context window, but changing things is often multiple operations + a transactional commit (or rename) on success.
We went from using a model to "change this" vs "write me a reusable script to change this" & running it with the right auth tokens.
The huffman tree, LZ77 and LZMA explanation is truly excellent for how concise the explanation is.
The earlier Veritasium video on Markov Chains in itself is linked if you don't know what a markov chain is.
I expected Veritasium to tank when it got sold to private equity & Derek went to Australia, but been surprised to see the quality of the long form stuff churned out by Casper, Petr, Henry & Greg.
I love this approach, because this is a good litmus test of both criteria I am looking for in engineers.
Part one is "can you use these tools fluently" and the second section is more like "what are you without that suit, Tony Stark?" (or Peter Parker, depending on your movie preference).
The AI part has been moved to a take-home test section in my scenario, but the 40 minutes of the interview is "I want you to make a minor change to the tool you built" & see how a change in requirements bounces through the person's head.
The part 3 is the turn, where I pull the rug out of the assumptions in the original business case.
My company has probably hit max engineering size already, but I've found there are people who build "technical savings" in their architectures and not debt, which makes their next 3 steps velocity without compromising on pr quality.
That's probably the reason - we only need dressage horses and pure bloods now that the real draft horse is getting put to pasture.
These are no longer workhorses.
Six cylinders are the smoothest engines out there.
Honda used to have a 1L 6 cylinder engine for their bikes - the Gold wing has a 6 cylinder still.
The perimeter of the piston goes down in relation to its area (& multiplied by BMEP) when the radius goes up - looking at you Africa Twin.
The perimeter is where the unburnt fuel lives and gets caught up in the emission rules. So fewer larger cylinders is better according to EPA - 500cc each, maybe.
If we're only going to have hobby vehicles with internal combustion, then a six cylinder or doubling up to a v-12 makes sense.
They're toys for the weekend, not to put a 100k miles on it.
Imagine the scene from Ratatouille, where Remy explains "taste" and the brother finds it impossible to understand what it is ("Food is food").
The dad goes from being annoyed that Remy is a picky eater instead decides to put him to work as a taster. Gives him the job of approving forage that comes into the family & protect others from being poisoned.
The reason we say "taste" is because that's the closest parallel.
When it is even more vague, I call it a "code smell".
This is a bill with no votes - the first committee hearing is in March.
The purpose of the bill seems to be have some controversy & possibly raise the profile of the proposer.
The bill is written very similarly to how we enforce firmware for regular printers and EURion constellation detection.
The ability to hold two conflicting thoughts and yet continue to function is a test of intelligence - be able to see that things are hopeless and yet be determined to make them otherwise.
You don't need to lie yourself that the world is not falling apart, but being truly optimistic instead of nihilistic at the face of that is a difficult test for any intelligent human being.
In the scale of the universe and history, most of what you do is not important, but it is very important that you do it (rambles on about Gita, Ecclesiastes and Plato ...).
You can find all the possible tricks in making it debuggable by reading the y.tab.c
Including all the corner cases for odd compilers.
Re2c is a bit more modern if you don't need all the history of yacc.
Yes, this is not some sort of hard-fought wisdom.
It should be common sense, but I still see a lot of experiments which measure the sound of one hand clapping.
In some sense, it is a product of laziness to automate human supervision with more agents, but on the other hand I can't argue with the results.
If you don't really want the experiments and data from the academic paper, we have a white paper which is completely obvious to anyone who's read High Output Management, Mythical Man Month and Philosophy of Software Design recently.
Nothing in there is new, except the field it is applied to has no humans left.
Coherence requires 2 opposing forces to hold coherence in one dimension and at least 3 of them in higher dimensions of quality.
My team wrote up a paper titled "If You Want Coherence, Orchestrate a Team of Rivals"[1] because we kept finding that upping the reasoning threshold resulted in less coherence - more experimentation before we hit a dead-end to turn around.
So we had a better result from using Haiku (we fail over to Sonnet) over Opus and using a higher reasoning model to decompose tasks rather than perform each one of them.
Once a plan is made, the cheaper models do better as they do not double-think their approaches - they fail or they succeed, they are not as tenacious as the higher cost models.
We can escalate to higher authority and get out of that mess faster if we fail hard and early.
The knowledge of how exactly failure happened seems to be less useful to the higher reasoning model over the action biased models.
Splitting up the tactical and strategic sides of the problem, seems to work similarly to how Generals don't hold guns in a war.
I can believe SAS works great until the context has errors which were corrected - there seems to be a leakage between past mistakes and new ones, if you leave them all in one context window.
My team wrote a similar paper[1] last month, but we found the orchestrator is not the core component, but a specialized evaluator for each action to match the result, goal and methods at the end of execution to report back to the orchestrator on goal adherence.
The effect is sort of like a perpetual evals loop, which lets us improve the product every week but agent by agent without the Snowflake agent picking up the Bigquery tools etc.
We started building this Nov 2024, so the paper is more of a description of what worked for us (see Section 3).
Also specific models are great at some tasks, but not always good at others.
My general finding is that Google models do document extraction best, Claude does code well and OpenAI does task management in somewhat sycophantic fashion.
Multi-agents was originally supposed to let us put together a "best of all models" world, but it works at error correcting if I have Claude write code and GPT 5 check the results instead of everything going into one context.
We can optimize the hash function to make it more space efficient.
Instead of using remainders to locate filter positions, we can use a mersenne prime number mask (like say 31), but in this case I have a feeling the best hash function to use would be to mask with (2^1)-1.
We didn't pick this because it was super technical, but because the financial team is the closest team to the CEO which is both overstaffed and overworked at the same time - you have 3-4 days of crunch time for which you retain 6 people to get it done fast.
This was the org which had extremely methodical smart people who constantly told us "We'll buy anything which means I'm not editing spreadsheets during my kids gymnastics class".
The trouble is that the UI that each customer wants has zero overlap with the other, if we actually added a drop-down for each special thing one person wanted, this would look like a cockpit & no new customer would be able to do anything with it.
The AI bit is really making the required interface complexity invisible (but also hard to discover).
In a world where OpenAI is Intel and Anthropic is AMD, we're working on a new Excel.
However, to build something you need to build a high quality message passing co-operating multi-tasking AI kernel & sort of optimize your L1 caches ("context") well.
The trouble is that you need to specifically optimize for fsyncs, because usually it is either no brakes or hand-brake.
The middle-ground of multi-transaction group-commit fsync seems to not exist anymore because of SSDs and massive IOPS you can pull off in general, but now it is about syscall context switches.
Two minutes is a bit too too much (also fdatasync vs fsync).
Netflix is a different creature because of streaming and time shifting.
They don't care about people watching a pilot episode or people binge watching last 3 seasons when a show takes off.
The quality metric therefore is all over the place, it is a mildly moderated popularity contest.
If people watch "Love is Blind", you'll get more of those.
On the other hand, this means they can take a slightly bigger risk than a TV network with ADs, because you're likely to switch to a different Netflix show that you like and continue to pay for it, than switch to a different channel which pays a different TV network.
As long as something sticks the revenue numbers stay, the ROI can be shaky.
Black Mirror Bandersnatch for example was impossible to do on TV, but Netflix could do it.
Also if GoT was Netflix, they'd have cancelled it on Season 6 & we'd be lamenting the loss of what wonders it'd have gotten to by Season 9.
This was one of the bigger hidden performance issues when I was working on Hive - the default coercion goes to Double, which has a bad hash code implementation [1] & causes joins to cluster & chain, which caused every miss on the hashtable to probe that many away from the original index.
The hashCode itself was smeared to make values near Machine epsilon to hash to the same hash bucket so that .equals could do its join, but all of this really messed up the folks who needed 22 digit numeric keys (eventually Decimal implementation handled it by adding a big fixed integer).
Databases and Double join keys was one of the red-flags in a SQL query, mostly if you see it someone messed up something.
The nurture part of it is already well established, this is the nature part of it.
However, this is not a net-positive for the folks who already discriminate.
The "faults in our genes" thinking assumes that this is not redeemable by policy changes, so it goes back to eugenics and usually suggests cutting such people out of the gene pool.
The "better nurture" proponents for the next generation (free school lunches, early intervention and magnet schools) will now have to swim up this waterfall before arguing more investment into the uplifting traumatized populations.
We need to believe that Change (with a capital C) is possible right away if start right now.
Generally speaking, it is the hardware not the OS that makes it easier to build for Macs right now.
Apple Neural Engine is a sleeping giant, in the middle of all this.