787 karma · joined March 26, 2023
I did my honest best possible interpretation of what you really meant from what you wrote.
>> We do know that zero parts of human decision making are based on predicting the next most likely letter based on a giant internet-based database. We do know that’s what LLMs do.
I read this as "The decision making of LLMs are based on predicting the next most likely letter based on a giant internet-based database."
Is that wrong?
I understood that your meaning was something like "LLMs can't reason, they just output likely letters"?
> If you’re claiming that the underlying structure of digital so-called neural networks is comparable to biological neural networks
No, I don't claim that.
What do claim is this: Regardless of how the LLMs were trained, they show overwhelming signs of being able to reason, and not just recall memorized information.
This doesn't mean that they always reason perfectly about everything.
But if they only memorized things and output the next likely letter, you would see them answering very badly much more often.
If you're claiming that the training objective tells us what kind of internal mechanisms the training produced, then I think that's just plain wrong.
Next-token prediction describes the optimization target, not the internal mechanisms that the training produced.
In the same way for the natural evolution of humans, DNA replication is the evolutionary objective. It's not a description of the internal mechanisms that evolution has produced.
As an example, we know that neural networks can be trained to develop generalized algorithms for arithmetic.
They might first memorize the training examples, then with further training transition to a solution that generalizes correctly to unseen examples.
In some cases we've even reverse-engineered the evolved internal mechanisms and found structured arithmetic algorithms rather than rote memorization. Interestingly, for modular addition this can involve Fourier representations, which isn't an algorithm I would have guessed gradient descent training of neural networks would produce.
But that was the only thing I tripped on. I enjoyed reading the article in general.
That's a lot different from "general purpose processor which can act based on program logic, stored data, and input data".
But it seems to me that if the LLM can effectively "execute" the instruction of how to take an input IP packet and generate a response IP packet based on a set or rules, then that's effectively a general purpose processor. And not an "auto completer", right?
It looks to me like the LLM "executed" the logic in pure output tokens, not by using any kind of external tool calls?
Because this seems to disprove that claim pretty convincingly?
32:22 It's time to be honest. As engineers, we are not known to be effective communicators. That's just not a strong suit of an engineer.
32:29 It is time to lay out and systematize the communication so that every organization that's involved knows far more information than they need. We have actual targets and dates that are actionable and they don't depend on a miracle in technological innovation occurring at some point on the Gantt chart.
32:48 We actually have actionable things that have to happen in certain time frames. And if something doesn't happen, a critical technology isn't in development, we communicate that there's a schedule slip to everyone.
From what I've seen of recent very large engineering projects, saying "we need Gantt charts" is a good way to get sidelined/fired?It seems that modern organizations absolutely loathe up front systems engineering and planning?
https://de.wikipedia.org/wiki/Verdeckte_Gewinnaussch%C3%BCtt...
As someone else mentioned, the taxes are different.
Namely: Salary is taxed lower than dividends. So the German tax authorities checks very carefully that you don't pay salary instead of dividends. If they determine that you paid out dividends as a salary, then you'll be charged with tax fraud.
Now you might say, "I don't care about paying a bit extra in taxes, so I'll pay it as dividends as they wish"
The problem is that you can only pay dividends the year after you earn the money.
If you can set a fixed salary which you can keep paying throughout, and then wait for the dividend payments next year, that's fine.
But what if you want to pay yourself wildly different amounts of money each month based on how much you managed to charge your customers? You can't just keep adjusting your salary up and down every month with a corporation.
So here's where something like a sole proprietorship may be simpler from that aspect?
Another thing you want to look at is "how easy will it be to dissolve the operations?" With a GmbH/UG it takes several years and potentially many thousands of euros in accounting fees. Not sure about the foreign corps. I think German sole proprietorships are simpler in either case?
Also, Germany has a "Moving away tax" where you get taxed on the fictious value of your company if you move away from Germany. This fictious value can be quite a lot more than what you'd actually get if selling the company.
Yet another thing: Depending on your setup, you may be covered by different rules regarding health insurance and pensions. If you don't make a lot of money in the beginning, it may be best to stay in the government insurance. But if you think you'll make a lot of money, it can be better to be able to do private insurance instead? There are rules on how you can move back and forth between government/private, so this is another area to consider carefully.
This is my understanding as a layman, please check this with a competent local tax expert before acting on any advice here.
I'm fairly sure the German tax authority will claim that you have a local German branch office since you live and work there.
That might be OK tax wise?
But I'd recommend starting with the tax situation in Germany.
Having limited liability through some kind of corporation can be nice.
But on the other hand, it becomes harder in Germany to pay out a varying salary as profits fluctuates throughout the year since the German tax authorities will see that as an illegal dividend payment from your company.
From this perspective it can be easier to set up some kind of sole proprietorship. Easier accounting etc and can pay out profits easier. But you get the personal liability.
This is not hard advice, just some things to point out that it gets complicated fast. So I'd recommend spending a few hundred euros on getting advice from a tax professional to begin with.
But yeah, if the recruiters start asking for "10 years experience with Claude Code", then I guess a tongue-in-cheek answer would be "sure, I did 10 projects in parallel in one year".
As the same age as Linus Torvalds, I'd say that it can be the opposite.
We are so used to "leaky abstractions", that we have just accepted this as another imperfect new tech stack.
Unlike less experienced developers, we know that you have to learn a bit about the underlying layers to use the high level abstraction layer effectively.
What is going on under the hood? What was the sequence of events which caused my inputs to give these outputs / error messages?
Once you learn enough of how the underlying layers work, you'll get far fewer errors because you'll subconciously avoid them. Meanwhile, people with a "I only work at the high-level"-mindset keeps trying to feed the high-level layer different inputs more or less at random.
For LLMs, it's certainly a challenge.
The basic low level LLM architecture is very simple. You can write a naive LLM core inference engine in a few hundred lines of code.
But that is like writing a logic gate simulator and feeding it a huge CPU gate list + many GBs of kernel+rootfs disk images. It doesn't tell you how the thing actually behaves.
So you move up the layers. Often you can't get hard data on how they really work. Instead you rely on empirical and anecdotal data.
But you still form a mental image of what the rough layers are, and what you can expect in their behavior given different inputs.
For LLMs, a critical piece is the context window. It has to be understood and managed to get good results. Make sure it's fed with the right amount of the right data, and you get much better results.
> Nowadays I just paste a test, build, or linter error message into the chat and the clanker knows immediately what to do
That's exactly the right thing to do given the right circumstances.
But if you're doing a big refactoring across a huge code base, you won't get the same good results. You'll need to understand the context window and how your tools/framework feeds it with data for your subagents.
AI does a much better job of translating than the stuff I see on TV.
I get the impression that it's done by a lowly paid person who uses a computer dictionary to translate word by word, in a very rushed manner.
For the first time since 1783, there are now "Hessians" (German state troops) in North America with their guns pointed at the United States.
Quite the opposite. It has already led to a hard core of NATO countries shifting gears quickly.
If one or more other NATO countries attack them, it would push the hardcore NATO countries even closer together.
Small force, symbolic stand: "Remember the Alamo", but "Remember Greenland" instead this time.
By the way, can you tell me the background and meaning of these phrases?
"Don't tread on me!"
"Live free or die"
"Give me liberty or give me death"
"From my cold, dead hands"
Sure, you can convince a close friend of yours to take his home security much more seriously by telling him that you'll come by later and rob him at gunpoint.
But do you think he'll be even remotely friendly to you after that?
The whole southern part of Greenland was empty when Denmark landed there a thousand years ago.
Bad weather and the Inuit managed to kill off the Danish settlers after that, before they returned a few hundred years later.
So the Danish were one of the original settlers of Greenland. Not "colonizers".
Or do you call the Inuit "colonizers" too, since they spread to lands outside of the original home?
Counterpoint: ChatGPT came up with the new expression "The confetti has left the cannon" a few years ago.
So, your claim is not obviously true. Can you give us an example of a programming problem where the LLMs fail to solve it?
Fun fact: The VT-52 didn't have a loudspeaker for the bell sound. Instead, it had a electromechanical relay which was set up to self-oscillate.
"Typing a character produced a noise by activating a relay. The relay was also used as a buzzer to sound the bell character, producing a sound that "has been compared to the sound of a '52 Chevy stripping its gears."
Counterpoint: ChatGPT came up with the new idiom "The confetti has left the cannon"
How many software engineers with a good math education can do this?
> It's not. It's a query-retrieval system that can parse human language.
And humans aren't general AI either. They're just DNA replicators. It is very obvious when you realize that humans weren't designed to be intelligent. They were just randomly iterated through an environment which selected for maximum DNA replication.
Until you have a higher being which explicitly designs for intelligence, you'll just get things like LLM query-retrievals, or DNA replicators.
How well will the european countries survive with it if the US cuts off access to spare parts, SW maintenance links etc?