In india, my telco gives me google ai pro for free. And agy with flash goes a long way.
2,746 karma · joined March 26, 2023
Currently academia
In india, my telco gives me google ai pro for free. And agy with flash goes a long way.
Till then, transformers were being used primarily for stuff like translation and such and no one was even pretraining at scale, even tho transformers and attention existed.
Openai and google if you count T5 persisting with pretrained generative models was what led to the LLM boom. Yes they used transformers, but that's just one IMO minor aspect.
On the private test set, the right way to evaluate this type of model, is giving i it the test question Q, which it will first train to AR predict first, and then it will inference using the just-updated weights with Q as prompt, giving you back A, and then you compare A with A_true secretly.
In the test dataset's Q,A pairs, it was only trained to next-word predict the question itself, and it was not given the answer at all.
It was then evaluated by seeing if it is able to output A_test given the Q_test as prompt.
What would be cheating is training it to produce A_test (given Q_test as prompt) as well, since then you can always make a model that scores 100% by just memorising Q_test, A_test pairs.
The complaints online mostly stem from not reading that properly and assuming they trained on Q_test,A_test instead of just Q_test. This is further because these days large LLMs are inadvertently trained on many benchmark solutions even unintentionally due to the massive scale of data and the infeasibility of auditing it all. But none of that is the case here.
The reason you want to train on Q_test is because in these AR transformer models, they learn useful composable encodings of Q by simply learning to next-word predict Q. So you enable the model to learn composable encodings of the test questions, so that it can hopefully "connect it" to an earlier train problem it had seen, and adapt the solution it had seen for that, much like humans do in school exams.
Without this step, you are making it difficult for the model to "connect" the test question to a train question it had seen earlier, and then it still has to adapt the solution. This way, you precompute that "this test question is like this train question" and then during the exam you only have to do the adapting the solution part after a simpler "retrieval" process.
You can just think of next-word training Q_test as a "retrieval" process.
This practice often used in continual learning or "test time training" is not yet useful in general real world ML tasks due to the differences in memory and compute requirements, and more so the general fragility of training large neural networks in a streaming realtime way (as opposed to large data, batched), versus inferencing from a static neural network. It is due to that fragility that I believe (correct me if I am wrong) this guy had to train on a batch of Q_tests. If you enforced that you will not provide Q_test_2 before they answer Q_test_1, the performance will drop.
While the increased compute and memory is difficult to solve inherently, there are various efforts being made to fix the fragility, especially in reinforcement learning where this is called "streaming RL", there is revival of interest as seen in RLC 2026.
[Note] Arc-AGI-1 doesn't have any actual english words or such, but it's simpler to pretend it was a basic Q&A benchmark to explain the above
There is a lot of SPV backed financial engineering upstream, but not here.
Additionally, HBM requires 3+ times the capacity of DRAM, so it gets crowded out. They do command 5x margins, so the unit economics is ok. But this means that even new capacity will mostly be used to service HBM demand, and consumer or enterprise DRAM is unlikely to benefit at the same pace - it will benefit somewhat of course.
I wouldn't even consider that a vulnerability tbh, every personal laptop I had I add myself to docker group. Yes, you can not namespace pids, filesystem, etc, and get root, but it's never mattered.
If someone can run that docker command, they can already read your whole homedir, edit bashrc, etc etc,. and sudo is useless anyways.
Only on a system where you are a user without sudo access, does it even begin to make sense. And if you go to the trouble of intentionally setting up a user without sudo access, you wouldn't be adding that user to the docker group either. In the default install, I assume omarchy adds you to the sudoers as well, making this a perfectly ok thing to do
Even if you participate in the esteemed Red Hat Security Theater and use wayland, flatpaks, etc, most flatpaks can write anywhere in your home dir, so they can do this too.
On standard linux desktop, sudo is not really security, but it is a UX improvement as it adds friction to accidentally doing things to the "system".
[I don't use omarchy]
This slow transition period will also help answer the "where will they go" questions. We can't answer them right now.
Ultimately, everything we do is in service of social political and personal human incentives, and I think the effect of that is discounted when people make these takeoff predictions for AI and "AGI"
It has nothing to do with skill. Both were very skilled jobs. Draftsman as well despite many people going to it straight from school. But computers and CAD mean that it is now necessary for someone to do a STEM degree to be a draftsman. Recorded audio made many stage musicians redundant. It is cheaper to do it this way and gets superior results, that is all, there is no further agenda.
Now too, the next generation of voice actors and many other knowledge workers will have to go up the value chain one step and operate or potentially build these tools (in whatever form they mature to in a decades time).
The current generation of voice actors will face the same situation as many before in the performance industry - stage musicians/performers for example that were made redundant by recorded audio. The reality is that most of them just left and dispersed into the economy doing completely unrelated jobs.
For software and generally computer engineers, this new primitive happens to itself be software, so it's less of a transition and an easier upskilling path to learn to build it. And building it is one step higher in the value chain than simply using it. That is a structural advantage.
For general workflows, I agree. Doing some DAG system is just a waste of time if you don't have a grounded verifier for every node. If you are going to human verify, why make a dag of subagents and waste time? Keep the dag in your head, the way we all did before LLMs. The dag in your head is also much better.
For specific workflows however, even this is under engineering. If your specific task or family of tasks can truly be decomposed into multiple verifiable subtasks then you should spend the effort to build that system. It doesn't matter if it's technically worse than the other method because it is basically automatic and incredibly cheap. For complex systems this is a lot of upfront work, but if it is done, then the return is completely outsized.
I painted it black and white, but you can of course compose everything.
When you have a few example chats you want a model to emulate - say you made it from your proprietary data, you can train any open model on that data in this cheap way. You don't lose any quality versus not using lora since the models overall knowledge won't shift that much due to your data anyways, so it's a waste to make high dimensional updates.
However, only in some cases is it worth it and equal in quality to just making a good retrieval system and exposing it to claude code or whatever. If a retrieval system over the same data is very difficult, or if the data simply must be proprietary, then you should go for it.
[1] technically "rank", but I'm simplifying
I don't see how? it is a 50Wh battery. Even at 6W that is about 8h of usage? On what OS + workload on x86 are you seeing just 6W usage? Or is it that each day your usage is only 2 hours?
For indians, you also have the religious aspect because the immigration center you go to to visit Mt Kailash is there.
That immigration building was the white one that got completely shredded in the popular video going around.
There is also the other idea where you run your attention layer on the GPU/TPU/Trainium and the FFN on the SRAM accelerator. Because KV cache is more difficult on cerebras etc, while MOE latency is easier to deal with
Even in nvidia land rubin + LPU does a similar thing.
It has its downsides of course - if your traffic swings prefill heavy to decode heavy, you can't suddenly use your lpu for prefill. With GPUs they're totally interchangeable. Tradeoffs.
A couple of years ago, people I know got paid OK for relatively simple programming and logic RLHF tasks. But very soon it turned dark because that sort of data was required less and less, and the number of feedbackers has grown.
Today, the type of data the model developers pay for requires actual domain experience. E.g in software engineering they have people work in simulated environments with other LLMs, grade them, feedback, PRs, Jira everything.
This pays well and they treat you well, entice you with more money/task/hour, etc, for now. In few years when this gets drilled into LLMs, these guys will face the same painful hours and bad pay and less work and so on too.
In physical tasks, we are still in the early stages where basic packing clothes (in a textile factory setting) etc is being recorded and data is only now being used for training. Due to the problems with translating human hand data to robotic hands, these people do the factory work holding robotic grippers and operating that, you should see a YouTube video. But this also means that it's much less sweeping than data collection in SWE. In many cases it's not practical to collect data given that you have to use the specific gripper, wear a big gopro type thing, etc. So I am expecting much slower of an impact on physical tasks (of this kind) compared to how quick the uptake was in SWE/math.
I have yet to hear back from them on how it's going for teamwork white collar tasks, it's been a few months. Everyone is paying for the end products of those it seems - grok bot, perplexity computer, claure cowork, chatgpt work etc,. Not as much as their coding agents of course.