As for evolution, you can still go gradients, the problem is that you can't do gradients in a space with many false positives. You need some method of figuring out the true optimal point.
2,279 karma · joined April 8, 2019
As for evolution, you can still go gradients, the problem is that you can't do gradients in a space with many false positives. You need some method of figuring out the true optimal point.
* Start their own company
* Go work for a startup where they actually get paid options, and have a say in what the company does.
* Go chill at one of the big companies getting paid good salary while coasting because the work is so easy.
What they certainly don't do is go work at a company for less pay and harder work hours, all so that they can make a literal Nazi richer.
Most manufacturers already were working on hybrids, which to this day are still suprerior to EVs. Chevy Volt, outside of being Chevy, was still one of the best cars ever made for utilitarian purpose. Nobody wanted to foot the bill to do electric conversions until this was necessary.
Tesla only opened up a market segment for high end electric cars, which I guess is cool, but far from revolutionary. The model 3 was a big success only because again, it was subsidized. Meanwhile BYD actually makes cheap affordable electric cars, and we both know why they are not sold in US.
>If not for SpaceX there wouldn't be gigabit internet connectivity in the middle of the ocean.
Plenty of companies were doing geostationary orbits with satellite connectivity. SES for one.
Any more Elon slop? You realize you are defending a dude that is literally a Nazi, right?
If you think anything Elon doing is groundbreaking, you have no idea how the world works. Recent Space X ipo showed that the launches aren't cheaper, they are just heavily subsidized. Tesla was a piece of crap until they got their model 3, the only reason Tesla succeeded with their S model is because Elon was the edgy hype dude who managed to generate enough hype to carry them through the bullshit with the car. Self driving was supposed to be solved last year, and tiny companies like Comma AI manage to build self driving systems that are in someways better than Teslas.
I bet you think Steve Jobs was a visionary as well lol.
...You have weights matricies for K/Q/V, which when post multiplied with the input, give you the KQV ** matricies **...
In the contest of LLMs, you cant have these types of coded function. Your function has to be a mathematical equation that is smooth - i.e no discrete steps, no singularities. The reason for this is when any neural net is trained, you use backpropagation of the error to adjust weights, and how much you adjust them is directly proportional to the weights effect on the final output, and in order to compute this, you have to have smooth functions from start to finish.
So what you do instead is you add data to your 4 values, that capture different relationship between them. If your 4 values are x,y,k,and h, your first data point can be a1x + b1y + c1k + d1h. The second point can be a2x + b3y + c4k + d5h. And so on. You can have as many of those values as you want. And then you can add, combine, and scale those values in any way you chose.
This basically gives you a map of 4 values into a binary decision whether the ball will end up in a goal or not, after sufficient training. However, the total number of extra values that you chose has to be large enough to capture all possibilities - if you don't have enough, you will start to make mistakes for some initial conditions.
>Nirvana? Singularity? Paperclips? Vernor Vinge rising from the dead? I'm curious; please share!
Simulated evolution. Thats how you "solve" highly nonlinear chaotic systems. And generally, if you think about it, you have to have some secondary system on top of the knowledge embedded in LLMs to drive them to select certain tokens, which then starts to eerily resemble what humans call emotions in themselves.
most people in ML have no idea what transformers actually are.
Traditional networks, at every layer, used to be output = [weights matrix][input], where input is a vector, and weights matrix is the weights, where each row corresponds to the set of weights for each neuron.
Transformers upscale the dimension of the data. Instead of the above, transformers do [output] = [input][weights_matrix]. When you multiply an input by a matrix, you get an output matrix back. Thats all that happens. Nothing fancy. You have weights matricies for K/Q/V, which when post multiplied with the input, give you the KQV vectors, and then you just simply multiply them together and apply a scaling factor.
There is nothing magical about K/Q/V. There is nothing about any one doing any querying or any one representing some keys. The naming is just a carry over from how they that selection process is used in pre llm data science fields where you manually define the key and query matricies to define relationships between components.
The reason of why it works is because is an extension of something called kernel tricks from pre LLM machine learning days - you map a lower dimensional space to an extra dimension based on some equation, and it lets you apply some classifier on the combination of existing values and new value. Thats what transformers are doing - they are mapping the individual token to the dk x n_heads latent space, which allows for a higher dimensional representation of the data, capturing complex relationships.
You can do Transformers with 5 matricies instead of 3, you can do this with 4-dimentional tensors, and so on. The thing is, there really isn't any way to tell if any of that gives you more advantage - it certainly would give you more granularity, but as of right now, in terms of training to generate a specific token given previous ones before it, it seems that you don't need any more dimentions than dk x n_heads. Interestingly enough, you also can mathematically represent any such transformer including the starting one with a sequence of linear layers like in traditional networks, the only thing is that it becomes computationally inefficient due to having duplicates of data.
The reason why RNNs and others and others didn't work is because RNN training is effectively trying to linearly regress on chaotic effects - i.e what set of starting conditions would evolve with a given process into what you want. This is an NP hard problem, and you can't really do it linearly.
Transformer models on the other hand, use breadth instead of compute to capture interactions. In those learned weight matrices, you have a latent space of a bunch of "knowledge" compressed, and an algorithm to search on that "knowledge".
But, its very possible that an RNN can be smarter than a frontier model while being much smaller in size - in the same way that its very possible that you can have the right set of prompts for an existing local inference smaller model that can basically be very close to AGI in terms of being able to solve any problem across any domain. Right now, the space is about exploring those prompts, which is the frameworks and harnesses, to get to there, as well as making the compute portion more efficient so you can explore that space faster.
And the thing that comes after harnesses/efficiency in terms of progress should be obvious if you understand all of the above.
A) He literally says "I tested a different Qwen model for the comparisons between Mac and PC." The model he tested has to fit on one GPU, otherwise the inference is dogshit slow as you are offloading results to ram. If you ran any amount of local inference, you would know this. Considering that Qwen3.8-Flash-Next Q4 is still 100gb, there is no realistic way to run this with a 5090. The model that was run was this https://ollama.com/library/qwen3.8:27b. And the speed of that model on a 5090 in terms of tok/sec is not 60 lol.
B) If M5 ultra runs 40 tok/sec on qwen3.8:27b (and lets assume its the mlx version to gain a performance boost: https://ollama.com/library/qwen3.8:27b-mlx), you have to be delusional to believe it can run 100gb models at 100 tok/sec lol.
As a bonus, in terms of use, its pretty well known that Qwen models are RLed to chase benchmarks. Check out https://huggingface.co/Qwen/Qwen3.8-27B versus https://qwen.ai/blog?id=qwen3.8-flash-next, using different benchmarks the 27b outperforms the flash next on agentic coding. But it matches it in other areas pretty well. So tell me again why you need 100gb models running dogshit slow at peak ~20 tok/sec?
It is so incredibly sad how hard you try to sound intelligent. But thats on par for the course of any person hyping up apple products, throughout apples history.
Considering that Apple probably doesn't want you to engage in this level of pettiness for their advertising posts, you have outed yourself to be #2. And Im not angry at all lol, you keep doing what you do, people like you in the industry are the reason I can work 8 hours a week and still get get paid a lot while being reviewed highly.
Looking at the article, which you clearly didn't read,the m5 ultra runs Qwen3.8, which fits on one GPU conveniently, at ~20 tok/sec. This is a fucking joke. It will take roughly a minute to generate one code file. Congrats if you want privacy I guess, but for straight up coding, you are better just using cloud models.
Meanwhile, I have an $800 mini PC, $200 Occulink gpu dock, a $2000 3090 and a $300 power supply, and I can run Qwen at over 100 tok/sec prefill, not to mention insanely quicker during inference. So its pointless to spend Mac M5 Ultra prices on Apple shit when they can have something much faster for cheaper
The whole thing of "well I can run bigger models that don't fit on a GPU" is either paid Apple advertising, or you are just an igorant fanboy.
So I ask you again, which one are you?
Also keep in mind that https://www.obdev.at/blog/a-hole-in-the-wall/ and its subsequent patch https://support.apple.com/en-us/102445 refers to things that bypass the firewall. There are things on your Mac that simply won't function if you block them from phoning home, making your OS an unusable mess.
You can verify this yourself. Get an old laptop to act as a wifi hotspot and forward traffic over usb ethernet adapter to your actual router. Then run tcpdump on the computer. You will see the multitude of phone-home traffic.
Ive essentially followed that paradigm with Python and C. I start out writing Python code. If I need something to run fast, I build a standalone C application that either reads from a file or listens on a socket, and just invoke it from Python. No need to write the entire thing in Rust and deal with all its semantics when it will be at best like 2% faster.
Write python code, ask any llm to translate it to C, then compile the C code - if it produces errors or fails to run, ask LLM to fix it. Then take it a step further and ask it produce machine code, and repeat the procedure.
Then RL the llm on the above, and you basically have a Python -> Machine code compiler. If you cover every single possible python syntax, every single possible C syntax, every possible standard library call, and all the compiler optimization examples (all of which is a final set), you should get something that is extremely accurate.
No different than the electrical signals in the intermediate neurons in your brain that comprises the latent space where all the processing happens
>They communicate in natural language which is itself ambiguous
They lack one-shot precision, sure, but it doesn't matter. They are precise enough with refinement over multiple prompts.
>the weights and training inputs are being hidden as a "trade secret"
For frontier models that make the company money through api pricing, sure. There are plenty of open source models that can be used for the same tasks, which have open weights.
>This doesn't really mean much unless we understand what is the quality and relevance of these sources for the problem at hand.
All of the modern models are RL trained on specific tasks when it comes to coding. I.e the initial training run learns to predict the next token based on context, from all the available texts, but then the RL runs specifically train the model in a harness where it produces code and RLed to produce correct code with specific formatting.
Lets say that you have a country where slavery is legal. The slaves of course don't like it. But lets say that the populace is largely in favor of it, because slave labor keeps the barebones economy going and people can trade and have food to eat. Most people don't actively capture slaves, they just buy them at the market.
If you had the power to eliminate slavery in that country, you would free the slaves, but you would basically condem a lot of the people to die from starvation.
So whats the "correct" moral stance? Real world isn't black and white like you would want it to be.
In the case of a Safari, there is this moral middle ground - according to the countries policy, if an animal attacks a human, it generally is a candidate for being put down because of perceived danger to humans in the area, but on the flip side, its also morally wrong to kill an animal just because it does that.
In the case of the gang warfare, you are doing a service by reducing the number of gang members and crime, but its also morally wrong to take someones life in offense for personal entertainment? (If its not clear, the difference in doing that in PNG versus any other country is because specifically of the lack of law enforcement that is the way of life there)
Another hypothetical scenario - say you have a country where age of consent is low, and certain type of porn that is highly illegal in the rest of the world is freely available (which is not far from reality, scarily). You can't stop the creation or distribution of that porn without basically invading the country and changing its law. However, what if you used said porn and supplied to convicted pdf files in your home country so that they are less likely to exploit young kids irl? Is that morally wrong, or is it morally justified?
Most of the human written code, in places where that code needs to make money, is decidable either entirely or in large parts. I.e without running the code, you can take a domain of inputs and build a complete range of outputs solely by looking at the code.
The way that works in your head is that you are effectively doing a compilation to a logical like structure, which then you can use to infer what the output will be from what the input is, and its a direct mapping that is invertible and separable, so if you know what the output should be, you know what the input is, you can pinpoint the exact location where it breaks. Thats how humans write code.
If thats not clear, imagine a piece of code that splits strings by spaces, deletes the empty strings, and returns the number of words in a string. The fact that you can say that if you want 3 words, there should be maximum 2 sequences of continous spaces between words, is you effectively transpiling that program into a latent space inside your brain neurons and inverting it.
LLMs essentially do this, with the added advantage of having been trained on a HUGE number of codebases, so they can recognize patterns that a human cant.
Where LLMs struggle is complex behavior - they can't simulate things like a human can and choose the best course of action. Even harnesses for agentic loops that can auto run and debug code can't match what a human can do in this regard (hence why self driving still sucks rn).
So moving forward, being a good coder isn't going to be about writing code, or even about prompting LLMs. Its going to be all about whether or not you can design good custom agentic loops, which necessarily involves knowledge of the model at hand (i.e what words you have to use to get it to do the right thing). This will be especially true as investment into "private" inference grows where companies will be using smaller models that have less detailed RL and thus will need much more guidance to do the right thing.
So even if you end up with a mess of a codebase, as long as you define your test cases and they all pass, what is in the middle doesn't really matter.
I.e "brainwashed control opposition". You got people like the twitch streamer Hassan Piker being paid by Bezos (who donated to Trump) to spread communism under name of progressivism, which basically gives conservatives the ammo to attack Democrats on.
It should be clear to anyone that the problem with the current administration is not the conservatism, but the authoritarianism that comes with its own set of people who are stupid that are not actually conservative. Republicans would easily hold down massive support if they were about conservatism and not about authoritarianism - i.e small government without expanding overreach, absolute freedom of speech, following the rule of law, and focusing on economic expansion.
And communism is authoritarianism, what you want is no different than what conservatives want, you just think your version is better, which is exactly what they think. But its easy to sell that to you because you think that you are on the left and thus different from republicans.
There is no middle ground where you assume that people have rhe ability to think and make rational decisions, except for the times when theybmake bad decisions, but in that case its not their fault.