HNHacker News
TopNewBestAskShowJobs

ActorNightly

2,279 karma · joined April 8, 2019

submissionscomments
ActorNightly··on Transformers Explained Visually
Im saying you don't need to capture directionality if you capture all possible cases of sentence construction.

As for evolution, you can still go gradients, the problem is that you can't do gradients in a space with many false positives. You need some method of figuring out the true optimal point.

ActorNightly··on Did OpenAI solve the wrong Navier-Stokes problem?
Betteridge's law of headlines
ActorNightly··on Grok 4.7
The smart people do one of 3 things:

* Start their own company

* Go work for a startup where they actually get paid options, and have a say in what the company does.

* Go chill at one of the big companies getting paid good salary while coasting because the work is so easy.

What they certainly don't do is go work at a company for less pay and harder work hours, all so that they can make a literal Nazi richer.

ActorNightly··on Grok 4.7
>If not for Tesla there won't be EVs you see today.

Most manufacturers already were working on hybrids, which to this day are still suprerior to EVs. Chevy Volt, outside of being Chevy, was still one of the best cars ever made for utilitarian purpose. Nobody wanted to foot the bill to do electric conversions until this was necessary.

Tesla only opened up a market segment for high end electric cars, which I guess is cool, but far from revolutionary. The model 3 was a big success only because again, it was subsidized. Meanwhile BYD actually makes cheap affordable electric cars, and we both know why they are not sold in US.

>If not for SpaceX there wouldn't be gigabit internet connectivity in the middle of the ocean.

Plenty of companies were doing geostationary orbits with satellite connectivity. SES for one.

Any more Elon slop? You realize you are defending a dude that is literally a Nazi, right?

ActorNightly··on Grok 4.7
Nice Elon slop.

If you think anything Elon doing is groundbreaking, you have no idea how the world works. Recent Space X ipo showed that the launches aren't cheaper, they are just heavily subsidized. Tesla was a piece of crap until they got their model 3, the only reason Tesla succeeded with their S model is because Elon was the edgy hype dude who managed to generate enough hype to carry them through the bullshit with the car. Self driving was supposed to be solved last year, and tiny companies like Comma AI manage to build self driving systems that are in someways better than Teslas.

I bet you think Steve Jobs was a visionary as well lol.

ActorNightly··on Grok 4.7
Doing much better if you consider the ratio of outcome success/ money put in.
ActorNightly··on Transformers Explained Visually
edit:

...You have weights matricies for K/Q/V, which when post multiplied with the input, give you the KQV ** matricies **...

ActorNightly··on Transformers Explained Visually
Imagine you have a soccer field, a ball with position x and y, a kick strength, and direction in an angle. Your job is to write a function that determines if the ball will end up in a goal. So that is 4 values. However the function itself will contain many intricacies, like trig functions, simulated drag, and so on.

In the contest of LLMs, you cant have these types of coded function. Your function has to be a mathematical equation that is smooth - i.e no discrete steps, no singularities. The reason for this is when any neural net is trained, you use backpropagation of the error to adjust weights, and how much you adjust them is directly proportional to the weights effect on the final output, and in order to compute this, you have to have smooth functions from start to finish.

So what you do instead is you add data to your 4 values, that capture different relationship between them. If your 4 values are x,y,k,and h, your first data point can be a1x + b1y + c1k + d1h. The second point can be a2x + b3y + c4k + d5h. And so on. You can have as many of those values as you want. And then you can add, combine, and scale those values in any way you chose.

This basically gives you a map of 4 values into a binary decision whether the ball will end up in a goal or not, after sufficient training. However, the total number of extra values that you chose has to be large enough to capture all possibilities - if you don't have enough, you will start to make mistakes for some initial conditions.

ActorNightly··on Transformers Explained Visually
I mean, given sentence construction, you don't really need to capture directionality, you just have a mapping of how sentences are constructed to the latent space of some representation.

>Nirvana? Singularity? Paperclips? Vernor Vinge rising from the dead? I'm curious; please share!

Simulated evolution. Thats how you "solve" highly nonlinear chaotic systems. And generally, if you think about it, you have to have some secondary system on top of the knowledge embedded in LLMs to drive them to select certain tokens, which then starts to eerily resemble what humans call emotions in themselves.

ActorNightly··on Transformers Explained Visually
>how transformers work,

most people in ML have no idea what transformers actually are.

Traditional networks, at every layer, used to be output = [weights matrix][input], where input is a vector, and weights matrix is the weights, where each row corresponds to the set of weights for each neuron.

Transformers upscale the dimension of the data. Instead of the above, transformers do [output] = [input][weights_matrix]. When you multiply an input by a matrix, you get an output matrix back. Thats all that happens. Nothing fancy. You have weights matricies for K/Q/V, which when post multiplied with the input, give you the KQV vectors, and then you just simply multiply them together and apply a scaling factor.

There is nothing magical about K/Q/V. There is nothing about any one doing any querying or any one representing some keys. The naming is just a carry over from how they that selection process is used in pre llm data science fields where you manually define the key and query matricies to define relationships between components.

The reason of why it works is because is an extension of something called kernel tricks from pre LLM machine learning days - you map a lower dimensional space to an extra dimension based on some equation, and it lets you apply some classifier on the combination of existing values and new value. Thats what transformers are doing - they are mapping the individual token to the dk x n_heads latent space, which allows for a higher dimensional representation of the data, capturing complex relationships.

You can do Transformers with 5 matricies instead of 3, you can do this with 4-dimentional tensors, and so on. The thing is, there really isn't any way to tell if any of that gives you more advantage - it certainly would give you more granularity, but as of right now, in terms of training to generate a specific token given previous ones before it, it seems that you don't need any more dimentions than dk x n_heads. Interestingly enough, you also can mathematically represent any such transformer including the starting one with a sequence of linear layers like in traditional networks, the only thing is that it becomes computationally inefficient due to having duplicates of data.

The reason why RNNs and others and others didn't work is because RNN training is effectively trying to linearly regress on chaotic effects - i.e what set of starting conditions would evolve with a given process into what you want. This is an NP hard problem, and you can't really do it linearly.

Transformer models on the other hand, use breadth instead of compute to capture interactions. In those learned weight matrices, you have a latent space of a bunch of "knowledge" compressed, and an algorithm to search on that "knowledge".

But, its very possible that an RNN can be smarter than a frontier model while being much smaller in size - in the same way that its very possible that you can have the right set of prompts for an existing local inference smaller model that can basically be very close to AGI in terms of being able to solve any problem across any domain. Right now, the space is about exploring those prompts, which is the frameworks and harnesses, to get to there, as well as making the compute portion more efficient so you can explore that space faster.

And the thing that comes after harnesses/efficiency in terms of progress should be obvious if you understand all of the above.

ActorNightly··on M5 Ultra Mac Studio Review
Nice try.

A) He literally says "I tested a different Qwen model for the comparisons between Mac and PC." The model he tested has to fit on one GPU, otherwise the inference is dogshit slow as you are offloading results to ram. If you ran any amount of local inference, you would know this. Considering that Qwen3.8-Flash-Next Q4 is still 100gb, there is no realistic way to run this with a 5090. The model that was run was this https://ollama.com/library/qwen3.8:27b. And the speed of that model on a 5090 in terms of tok/sec is not 60 lol.

B) If M5 ultra runs 40 tok/sec on qwen3.8:27b (and lets assume its the mlx version to gain a performance boost: https://ollama.com/library/qwen3.8:27b-mlx), you have to be delusional to believe it can run 100gb models at 100 tok/sec lol.

As a bonus, in terms of use, its pretty well known that Qwen models are RLed to chase benchmarks. Check out https://huggingface.co/Qwen/Qwen3.8-27B versus https://qwen.ai/blog?id=qwen3.8-flash-next, using different benchmarks the 27b outperforms the flash next on agentic coding. But it matches it in other areas pretty well. So tell me again why you need 100gb models running dogshit slow at peak ~20 tok/sec?

It is so incredibly sad how hard you try to sound intelligent. But thats on par for the course of any person hyping up apple products, throughout apples history.

Considering that Apple probably doesn't want you to engage in this level of pettiness for their advertising posts, you have outed yourself to be #2. And Im not angry at all lol, you keep doing what you do, people like you in the industry are the reason I can work 8 hours a week and still get get paid a lot while being reviewed highly.

ActorNightly··on Grok 4.7
Give me one actual reason why a smart person would take an underpaid overworked position at any of his companies.
ActorNightly··on Grok 4.7
Hey hows Starship going? Is it on Mars like he said it would be by now?
ActorNightly··on Grok 4.7
In other news, people hired by Elon suck at engineering.
ActorNightly··on M5 Ultra Mac Studio Review
Since you clearly don't use local llms, allow me to educate you - anything under 100 tok/sec is USELESS. When you are coding, the idea is that you want to have a system that can generate files fast, hopefully correct on the first try. Cloud models do this. Local models, by nature of having less parameters and more quantization, often require more guidance and repeated inference to get it right. The antigenic harnesses that people set up around local llms leverage this.

Looking at the article, which you clearly didn't read,the m5 ultra runs Qwen3.8, which fits on one GPU conveniently, at ~20 tok/sec. This is a fucking joke. It will take roughly a minute to generate one code file. Congrats if you want privacy I guess, but for straight up coding, you are better just using cloud models.

Meanwhile, I have an $800 mini PC, $200 Occulink gpu dock, a $2000 3090 and a $300 power supply, and I can run Qwen at over 100 tok/sec prefill, not to mention insanely quicker during inference. So its pointless to spend Mac M5 Ultra prices on Apple shit when they can have something much faster for cheaper

The whole thing of "well I can run bigger models that don't fit on a GPU" is either paid Apple advertising, or you are just an igorant fanboy.

So I ask you again, which one are you?

ActorNightly··on Turn off and restrict access to Apple Intelligence features on Mac
.... I literally gave you a way to check this yourself. Not sure why you would need a 3d party source to find something you can discover yourself.

Also keep in mind that https://www.obdev.at/blog/a-hole-in-the-wall/ and its subsequent patch https://support.apple.com/en-us/102445 refers to things that bypass the firewall. There are things on your Mac that simply won't function if you block them from phoning home, making your OS an unusable mess.

ActorNightly··on MCP was always a bad idea?
The better solution for security is sandbox execution. Most agents used in practice, if given terminal access, can find ways to get around most of the MCP restrictions.
ActorNightly··on Turn off and restrict access to Apple Intelligence features on Mac
Apple OS will bypass any user space firewall.

You can verify this yourself. Get an old laptop to act as a wifi hotspot and forward traffic over usb ethernet adapter to your actual router. Then run tcpdump on the computer. You will see the multitude of phone-home traffic.

ActorNightly··on M5 Ultra Mac Studio Review
When you are doing matrix math, compute is compute. Apple cant be more efficient due to physics. The only reason Macs are more efficient in general is that they have tightly bundled hw and sw for specific tasks.
ActorNightly··on Nvidia announces native GPU programming in Rust
Yep.

Ive essentially followed that paradigm with Python and C. I start out writing Python code. If I need something to run fast, I build a standalone C application that either reads from a file or listens on a socket, and just invoke it from Python. No need to write the entire thing in Rust and deal with all its semantics when it will be at best like 2% faster.

ActorNightly··on Nvidia announces native GPU programming in Rust
Why is this even a question, of course they can.

Write python code, ask any llm to translate it to C, then compile the C code - if it produces errors or fails to run, ask LLM to fix it. Then take it a step further and ask it produce machine code, and repeat the procedure.

Then RL the llm on the above, and you basically have a Python -> Machine code compiler. If you cover every single possible python syntax, every single possible C syntax, every possible standard library call, and all the compiler optimization examples (all of which is a final set), you should get something that is extremely accurate.

ActorNightly··on Learning Programming in an Age of LLMs
>They use inscrutable internal language of embeddings

No different than the electrical signals in the intermediate neurons in your brain that comprises the latent space where all the processing happens

>They communicate in natural language which is itself ambiguous

They lack one-shot precision, sure, but it doesn't matter. They are precise enough with refinement over multiple prompts.

>the weights and training inputs are being hidden as a "trade secret"

For frontier models that make the company money through api pricing, sure. There are plenty of open source models that can be used for the same tasks, which have open weights.

>This doesn't really mean much unless we understand what is the quality and relevance of these sources for the problem at hand.

All of the modern models are RL trained on specific tasks when it comes to coding. I.e the initial training run learns to predict the next token based on context, from all the available texts, but then the RL runs specifically train the model in a harness where it produces code and RLed to produce correct code with specific formatting.

ActorNightly··on Our chance to take down the Epstein class
So you don't deny anything? Good to know!
ActorNightly··on I can't stop thinking about Papua New Guinea
>Please explain how this is "not universally bad".

Lets say that you have a country where slavery is legal. The slaves of course don't like it. But lets say that the populace is largely in favor of it, because slave labor keeps the barebones economy going and people can trade and have food to eat. Most people don't actively capture slaves, they just buy them at the market.

If you had the power to eliminate slavery in that country, you would free the slaves, but you would basically condem a lot of the people to die from starvation.

So whats the "correct" moral stance? Real world isn't black and white like you would want it to be.

ActorNightly··on I can't stop thinking about Papua New Guinea
There being a market for something doesn't make it morally right, but it highlights the demand aspect where someone is willing to pay money, which means that there is some level of it not being universally bad.

In the case of a Safari, there is this moral middle ground - according to the countries policy, if an animal attacks a human, it generally is a candidate for being put down because of perceived danger to humans in the area, but on the flip side, its also morally wrong to kill an animal just because it does that.

In the case of the gang warfare, you are doing a service by reducing the number of gang members and crime, but its also morally wrong to take someones life in offense for personal entertainment? (If its not clear, the difference in doing that in PNG versus any other country is because specifically of the lack of law enforcement that is the way of life there)

Another hypothetical scenario - say you have a country where age of consent is low, and certain type of porn that is highly illegal in the rest of the world is freely available (which is not far from reality, scarily). You can't stop the creation or distribution of that porn without basically invading the country and changing its law. However, what if you used said porn and supplied to convicted pdf files in your home country so that they are less likely to exploit young kids irl? Is that morally wrong, or is it morally justified?

ActorNightly··on Learning Programming in an Age of LLMs
Your analogy is poor, and you are missing a very important fact.

Most of the human written code, in places where that code needs to make money, is decidable either entirely or in large parts. I.e without running the code, you can take a domain of inputs and build a complete range of outputs solely by looking at the code.

The way that works in your head is that you are effectively doing a compilation to a logical like structure, which then you can use to infer what the output will be from what the input is, and its a direct mapping that is invertible and separable, so if you know what the output should be, you know what the input is, you can pinpoint the exact location where it breaks. Thats how humans write code.

If thats not clear, imagine a piece of code that splits strings by spaces, deletes the empty strings, and returns the number of words in a string. The fact that you can say that if you want 3 words, there should be maximum 2 sequences of continous spaces between words, is you effectively transpiling that program into a latent space inside your brain neurons and inverting it.

LLMs essentially do this, with the added advantage of having been trained on a HUGE number of codebases, so they can recognize patterns that a human cant.

Where LLMs struggle is complex behavior - they can't simulate things like a human can and choose the best course of action. Even harnesses for agentic loops that can auto run and debug code can't match what a human can do in this regard (hence why self driving still sucks rn).

So moving forward, being a good coder isn't going to be about writing code, or even about prompting LLMs. Its going to be all about whether or not you can design good custom agentic loops, which necessarily involves knowledge of the model at hand (i.e what words you have to use to get it to do the right thing). This will be especially true as investment into "private" inference grows where companies will be using smaller models that have less detailed RL and thus will need much more guidance to do the right thing.

ActorNightly··on Learning Programming in an Age of LLMs
The difference is, the total cost for a human to write/modify code files is much higher than for an LLM. A lot of time has been spent designing program languages so that human time is more optimized. LLMS dgaf if the code is structured cleanly or is a mess.

So even if you end up with a mess of a codebase, as long as you define your test cases and they all pass, what is in the middle doesn't really matter.

ActorNightly··on Our chance to take down the Epstein class
>Lol I'm a communist.

I.e "brainwashed control opposition". You got people like the twitch streamer Hassan Piker being paid by Bezos (who donated to Trump) to spread communism under name of progressivism, which basically gives conservatives the ammo to attack Democrats on.

It should be clear to anyone that the problem with the current administration is not the conservatism, but the authoritarianism that comes with its own set of people who are stupid that are not actually conservative. Republicans would easily hold down massive support if they were about conservatism and not about authoritarianism - i.e small government without expanding overreach, absolute freedom of speech, following the rule of law, and focusing on economic expansion.

And communism is authoritarianism, what you want is no different than what conservatives want, you just think your version is better, which is exactly what they think. But its easy to sell that to you because you think that you are on the left and thus different from republicans.

ActorNightly··on I can't stop thinking about Papua New Guinea
I agree, but there is a market for the former, hence the question.
ActorNightly··on Our chance to take down the Epstein class
Again, if money can be used to influence how people vote, people dont have personal agency and that idea should be taken to its logical conclusion where everyones lives are tightly controlled.

There is no middle ground where you assume that people have rhe ability to think and make rational decisions, except for the times when theybmake bad decisions, but in that case its not their fault.

← PreviousPage 2 of 34Next →