HNHacker News
TopNewBestAskShowJobs

golol

1,204 karma · joined January 19, 2023

submissionscomments
golol··on Genie 2: A large-scale foundation world model
Do you want household androids? Because this kind of stuff is on the level of research a bery large step towards that. Think as it as ab example where we can make a model understand a lot of physical common sense stuff, which is the goal for robotics right now.
golol··on Genie 2: A large-scale foundation world model
I disagree. With more computation you can train a bigger model on the same size training data and it will be better. There is a lot if knowledge on the internet that GPT-4 etc. have not yet learned.
golol··on Approximating mathematical constants using Minecraft
I learned that 1 / e is the limit probability of a permutation not having any fixed points, nice.
golol··on DynaSaur: Large Language Agents Beyond Predefined Actions
I don't like the way LLM papers are written. LLMs receive inputs and produce outputs that are best represented as plaintext with some special characters. Simply showing a few examples of the agent's core LLM text continuation job would explain the architecture much better than figures. I can't help but feel that the authors which do this are intentionally obfuscating things.
golol··on DOJ will push Google to sell off Chrome
>pollute Seriously? You can't blame Google for its emissions.
golol··on DOJ will push Google to sell off Chrome
Why? What disaster? There can be no disaster when the product is free and there are many free alternatives with equal capability except for small conveniences. If you don't like Chrome because Google is being shady you can immediately seitch at zero cost. There is no disaster.
golol··on Something weird is happening with LLMs and chess
My understanding of this is the following: All the bad models are chat models, somehow "generation 2 LLMs" which are not just text completion models but instead trained to behave as a chatting agent. The only good model is the only "generation 1 LLM" here which is gpt-3.5-turbo-instruct. It is a straight forward text completion model. If you prompt it to "get in the mind" of PGN completion then it can use some kind of system 1 thinking to give a decent approximation of the PGN Markov process. If you attempt to use a chat model it doesn't work since these these stochastic pathways somehow degenerate during the training to be a chat agent. You can however play chess with system 2 thinking, and the more advanced chat models are trying to do that and should get better at it while still being bad.
golol··on Something weird is happening with LLMs and chess
You can still do this. People just lost interest in this stuff because it became clear to ehich degree the simulation is really being done (shallow).

Yet I do wish we had access to less finetuned/distilled/RLHF'd models.

golol··on Something weird is happening with LLMs and chess
Sorry this is just consiracy theorizing. I've tried jt with GPT-3.5-instruct myself in the OpenAI playeground where the model clearly does nothing but auto-regression. No function calling there whatsoever.
golol··on Something weird is happening with LLMs and chess
You have to get the model to think in PGN data. It's crucial to use the exact PGN format it sae in its training data and to give it few shot examples.
golol··on Something weird is happening with LLMs and chess
I've tried it myself, GPT-3.5-turbo-instruct was at least somewhere in the rabge 1600-1800 ELO.
golol··on Something weird is happening with LLMs and chess
Because it's a straight forward stochastic sequence modelling task and I've seen GPT-3.5-turbo-instruct play at high amateur level myself. But it seems like all the RLHF and distillation that is done on newer models destroys that ability.
golol··on An embarrassingly simple approach to recover unlearned knowledge for LLMs
As I understand the whole point is that it is not so simple to tell the difference between the model forgetting information and the model just learning some guardrails which orevent it from revealing that information. And this paper suggests that since the information can be recovored from the desired forgetting does not really happen.
golol··on Boston Dynamics robot Atlas goes hands on [video]
Well the robot is that machine....and much more.
golol··on Our First Generalist Policy
You know, it doesn't really matter to the original point. If the middle class is doing okay or if it's struggling, either way household androids and more generally less labor scarcity are exactly the kind of thing that will improve the situation.
golol··on Our First Generalist Policy
It was not. Maybe measured in relative terms the middle class is shrinking due to income inequality, but in absolute terms I am fairly confident it is at worst stagnating in America and western Europe. In many parts of the rest of the world there has been an amazing growth of a middle class that didn't exist before in the last decades. Eastern europe and asia of course.
golol··on Oasis: A Universe in a Transformer
Between the first half and the last sentence of your post is a giant leap of conclusion.
golol··on Our First Generalist Policy
The ever expanding middle class of course.
golol··on Our First Generalist Policy
Inbetween the current world full of labor scarcity, and the philosophical dilemma "what do I even do" post-scarcity utopia, is a world similar to our current one with much less labor scarcity and much more quality of life. That's what we're aiming for right now. What comes afterwards we can worry about then.
golol··on Our First Generalist Policy
I saw your foundation model is trained on data from several different robots. Is the plan to eventually train a foundation model that can control any robot zero shot? That is, the effect of actuations on video/sensor input is collected and understood in-context and actuations are corrected to yield intended behavior. All in-context. Is this feasible?

More specifically, has your model already exhibited this type of capability, in principle?

golol··on Physical Intelligence's first generalist robotic model
Idk this is really promising, how many robot foundation models have you seen before that also work very well? I believe this is all quite recent.
golol··on Physical Intelligence's first generalist robotic model
This is a duplicate thread. Can some mod merge them oO? I don't know how this works on HN.
golol··on Π0: Our First Generalist Policy
Isn't this a real big deal? They manage to train a foundation model that connects the physical understanding gained from unsupervised training on images and videos with the real physical understanding necessary to control a robot. They have impressive videos to show for it. The approach seems indeed scalable and generalist. I've thought for a while that the keystone missing for household androids is connecting the understanding of large multimodal language models with physical hardware. To me this looks like exactly that. It actually makes me optimistic that we will see household robots within a decade. Now I wonder why Tesla, Figure etc. are messing around so much with Teleoperation if this indeed works. Maybe I don't understand what's going on.
golol··on Tesla's Cybertruck is outselling almost every other EV in the US
>tests to prove its safety should be done in a way where gathering statistics doesn’t mean killing people.

Honestly, I believe that in the noisy world this is exactly how you do things. You try it out. We've done it for hundreds of years. If you have good reason to expect drastically effects you are careful of course. I don't think there is a good reason to expect drastic effects from the Cybertruck. If you were wrong you will quickly see problems arise and can cancel your trial. If nothing bad happens, seems alright. This is basically how we've dealt with the safety aspect of engineering and medicine since forever. You can never perfectly predict what's gonna happen.

golol··on How I write code using Cursor
I think there is only one thing we should focus on: Measurable capability on tasks. Understanding, memorization, reasoning etc. are all just shorthands we use to quickly convey an idea of a capability on a kind of task. Measurable capability on tasks can also attempt do describe mechanistically how the model works, but that is very difficult. This is where you would try to describe your sense of "understanding" rigorously. To keep it simple for example, I think when you say that the LLM does not understand what you must really mean is that you reckon its performance will quickly decay off as the task gets more difficult in various dimensions: Depth/complexity, Verifiability of the result, length/duration/context size, to a degree where it is still far from being able to act as a labor-delivering agent.
golol··on Dropbox announces 20% global workforce reduction
I suppose it is the following: 1. Maybe they didn't over-hire but instead the economic conditions changed. 2. Maybe instead of gradually letting go of people it is best to wait and then strike once and hard.
golol··on How I write code using Cursor
This is a prompt I gave to o1-mini a while ago: My instructions follow now. The scripts which I provided you work perfectly fine. I want you to perform a change though. The image_data.pkl and faiss_index.bin are two databases consisting of rows, one for each image, in the end, right? My problem is that there are many duplicates: images with different names but the same content. I want you to write a script which for each row, i.e. each image, opens the image in python and computes the average expected color and the average variation of color, for each of the colors red, green and blue, and over "random" over all the pixels. Make sure that this procedure is normalized with respect to the resolution. Then once this list of "defining features" is obtained, we can compute the pairwise difference. If two images have less than 1% variation in both expectation and variation, then we consider them to be identical. in this case, delete those rows/images, except for one of course, from the .pkl and the .bin I mentioned in the beginning. Write a log file at the end which lists the filenames of identical images.

It wrote the script, I ran it and it worked. I had it write another script which displays the found duplicate groups so I could see at a glance that the script had indeed worked. And for you this does not constitute any understanding? Yes it is assembling pieces of code or algorithmic procedures which it has memorized. But in this way it creates a script tailored to my wishes. The key is that it has to understand my intent.

golol··on Tesla's Cybertruck is outselling almost every other EV in the US
yea I prett much agree
golol··on Tesla's Cybertruck is outselling almost every other EV in the US
Easy to say but does that really represent the statistical reality? Does higher acceleration cause more accidents? Sports cars have more accidents. Sports cars that are driven by young people that want to show off. Cars with a high top speed. But do you think people will try to corner with the Cybertruck like they're racing? Yes they will spend a few seconds after each stop sign at hogher speed than if they weren't driving a Cybertruck but then they're just coasting down the suburbs with 30mph or whatever. But it is not obvious to me at all that Cybertruck driving should occur significantly often in dangerous fashion just because you can get to your cruising speed and overtake a bit quicker. You need statistics to back up these claims and these don't exist yet.
golol··on Tesla's Cybertruck is outselling almost every other EV in the US
I originally thought that the nhtsa would perform some tests of this kind but I was disappointed to find out that this doesn't seem to be the case. Most focus on pedestrian safety seems to be on automatic emergency braking and other ADAS systems. In this regard I claim that by default one should expect the Cybertruck to be at least decent. Not the best but surely not terrible. While the shape of the hood and the visibility obviously matter of course it seems to be the case that such evaluations are basically done by eyeballing it. Pedestrian dummy crash tests with collection of force data etc. seem to not at all be standard. Which sucks, but it again shows that the Cybertruck is not an outlier on some well-established metric.

I predict if after some years a statistical analysis is done on pedestrian crashes by car model the Cybertruck will just be some kind of average. It has decent AEB and it is a new car driven by young people. Would other cars have caused less injury in the crashes that do occur? Probably. To a degree that it justifies calling it a pedestrian killer? I doubt it.

← PreviousPage 4 of 17Next →