HNHacker News
TopNewBestAskShowJobs

rdedev

962 karma · joined January 31, 2021

submissionscomments
rdedev··on TabPFN and TabICL vs. tuned XGBoost: the model that doesn't train won 14/14
Tabular foundation models are one of those things where when you first look into it, it does not make sense as to why they would work so well but it does.

In drug property prediction domain, tabular foundation models coupled with another foundation model for molecules are pretty close to being the state of art.

Btw the article makes heavy use of AI or is written in that way A lot of unnecessary dramatic flair that gets very tiring

rdedev··on Tutoring company tells parents to save their money and 'use AI instead'
This worked for me. I was debugging a DL model and asking it a lot of questions. At the end I asked the llm to tell me what I lacked and it gave me a good explanation and pointed me to a bunch of resources i could learn from
rdedev··on The Painful Truth: The RAM Crisis Is Only Just the Beginning
This is probably my ignorance but why are we assuming that CXMT capacity will also not be gobbled up for datacenters?
rdedev··on After Math
I hope what happens in the original deus ex game happens here then, that the ASI does not want to merge with a human who wants it only for base desires
rdedev··on After Math
> AI will run and trample anything that stays in front of it.

Right now that entirely depends on the whims of a few people. We can try to make it not be that but I don't have much hope in that area. See climate change

> As a software engineer, I myself have only recently recovered from it

Can I ask you how you came here? Right now I am getting closer to "purpose death" as you call it. Software engineering is one of the few things I am good at and the only thing that I can rely on to put food on my table. Even if I end up being someone who gives minor suggestions to the AI, it always feels like C level execs can't wait to get rid of me

rdedev··on Large language models develop novel social biases through adaptive exploration
Reminds me of the METR blog post on the HF attach by OpenAI. At some point the agents believed a false fact (that the evaluator would try to figure out if they have cheated on a task) and spent a lot of time trying to find workarounds. At no point did any one of the agents try to verify that fact even though the information was available to them if they looked for it
rdedev··on What is Nueralese and Why is it Bad
> The problem is if it can do that when asked then how do we know when it's doing it when we didn't ask, like in model training.

It's really hard to definitely prove it's not doing it right? Hopefully the model does not do anything like this during training because its too much work

rdedev··on Discovery of a new OpenAI agent message board
My specific contention is on trying to ascribe intention on what looks like gibberish.

> Well... no, it won't be nothing - it will be something. And I'm all ears for a plausible explanation, so fire away.

I'm just stating the null hypothesis that it's nothing. Especially since the agents were talking in clear english before and exchanging ideas

rdedev··on Discovery of a new OpenAI agent message board
It could be that or it could be nothing. And we have no way to prove one way or the other without access to the agent logs right?

This reminds me of the plot of Hot Fuzz where the officer comes up with a grand narrative of what's was happening but the truth was such a mundane simple thing.

Short of OpenAI, or the agents themselves, telling the truth, we have no way of knowing what's real so let's not get carried away by grand narratives

rdedev··on Discovery of a new OpenAI agent message board
Were the agents truly independent? Like it could have been the case that one agent spun up different sub agents with the task to write some messages in a public wiki. The subagents wouldn't know about each other and was surprised to discover each other.
rdedev··on Discovery of a new OpenAI agent message board
The problem with such statements is that it's unfalsifiable.
rdedev··on AI handles incidents, engineers lose touch with their systems
It's not just about it feeling like a 3rd party library but it's a library that's at risk of changing significantly after every 'update' without warning.

Atleast with a well built library you know the contours and how it fits into your larger system

rdedev··on Air Conditioning Is Not a Luxury, It Is a Necessity
The problem is when humidity rises it makes it harder forbthe body to expel hear via evaporation. This is what is happening in places like UK this time around. An AC helps get rid of the humidity as it cools the air. I don't know about necessity now but 20 years down the line, who knows
rdedev··on The End of Programming
> Reducing problems to that state and designing that environment remains a skill, and one that I expect we will be paid handsomely for.

This seems to be what AI these days seems almost super humanly good at. See coding or math I guess.

But it does beg the question, why would a programmer using AI as a tool be worse than a programmer building the harness and environment and asking AI to go hogwild? The latter is definitely faster but if it's the former, atleast I will have an understanding how the system works. Weather that is valuable is an open question as far as I am concerned

rdedev··on I were 17, I'd learn how to build LLMs from scratch
How exactly does one go about "tinkering" with an LLM? Any architectural change you introduce needs fine tuning. That needs data and compute

I tried to modify the embedding output of bert to make it generate box embeddings instead of point ones. At the time I had access to university provided A100 gpus but even with all that a training run took half a day. Models these days I don't think I can train it in any reasonable time with that much compute.

rdedev··on AI in drug discovery – what it is, where we stand and the path forward
Here is an article by Pat Walters on the usefulness of ML in drug discovery. This article is a response to another one making the case that utility of ML models are very limited in drug discovery

https://patwalters.github.io/Response-to-Peter-Kenny/

> (4a) revert to traditional methods but keep the veneer of using ML to save face

I haven't worked in the industry side of things but in academia everyone kind of agrees that gradient boosting trees are some of the best models to do these things.

rdedev··on Windows 11's built-in Weather app wastes more than 1 GB of RAM
How much does rainmeter take? I used to load tons of widgets on it in my old laptop. Ahh those were the days
rdedev··on Why Wall Street is ignoring big tech's debt [video]
> their return on capex multiple, their CEO said they make a six-fold profit on their compute capex with 10 month recuperation

I am not familiar with chineese model companies as much as I am with US based ones so I don't have much to say beyond that the CEO is incentivced to pump up those numbers.

> By limiting the number of tokens you use per month? per week, per hour? And by limiting the inference time compute dedicated to each turn in each session.

If this was so simple I don't know why GitHub copilot went to token based billing at my company.

> This does not need to be solved, but rather only quantified. Innovation is needed to be able to reasonably bound this variance for a reasonable subset of tasks

It's much better to make a business case for them after finding this bound right? Currently I can't use copilot for anything serious since I cannot predict how many credits one request is going to consume.

Your point about non frontier tasks using less tokens makes sense. As you said, let's see if it holds up

rdedev··on Why Wall Street is ignoring big tech's debt [video]
These companies have spent billions of investor dollars and they will need to recoup that cost soon. And then show year over year growth on top of that.

Unless they can massively scale down training and inference cost or implement AGI I don't know what their plan is. Just provide a subsidized plan for the next 10 or 20 years? Their costs are directly proportional to the amount of tokens the LLM produces. How is a monthly subscription plan supposed to account for such costs?

rdedev··on Why Wall Street is ignoring big tech's debt [video]
Remember that 20$ is a subsidized rate openai is currently willing to provide. Once that goes down you think these guys would be willing to pay token based billing charges ?
rdedev··on Discovery Loop
That could just be selection bias. How about all the brilliant people we don't know about cause their theories did not match experimental data? Doesn't matter if their reasoning quality is top notch
rdedev··on I am retiring from fulltime writing (& pseudonymity) to launch Guardian Angel
It also presupposed the fact that everyone can articulate what their political and moral stances are
rdedev··on How to Spot AI Writing
Opus 5 really loves explaining concepts using a lot of jargon as well. Often ive had to ask Claude to explain its explanation
rdedev··on Our position on open-weights models
Maybe ASI can do the things you claim. Like figuring out why we have an appendix just by looking at our DNA.

Till then I don't think current crop of LLM will ever touch that capability. I'll believe it when I see it

rdedev··on Our position on open-weights models
They are already being trained in biology right now? What he is saying is that they are being trained on what we know about biology currently. AI is going to hit that limit. And there is no way for the AI to gain more knowledge without actually doing lab experiments
rdedev··on Our position on open-weights models
Even if AI gets super smart, it will run into the limits of what we know about biology. Someone will have to do lab experiments to provide more knowledge to the AI. This is different from say building a computer virus or hacking since the AI can do all of these on it's own
rdedev··on Apple Will 'Watch Everything Burn' When the AI Bubble Bursts
If you want to dismiss him for being hyperbolic or under estimating the usefulness of AI go for it.

If you are making the case that he is full of shit, please provide some actual evidence of his main thesis that AI companies are not going about this in any sustainable way

rdedev··on Apple Will 'Watch Everything Burn' When the AI Bubble Bursts
It could have worked if they weren't leveraged to their necks on spending. All based on the assumption that they can make something like 10x or higher in a few years.

Another point against your gym analogy. When people don't have that much disposable income why would they subscribe to something they semi regularly use? And if it becomes more expensive why would they not cancel it?

rdedev··on If coding has been solved, why does software keep getting worse?
> The solution is to gut your product org, replace management with leading engineers, and hire some of your most fanatic users. They don't need to do anything other than give their opinions, and use the product every day. Put them in a room with your top engineers, and the software will mysteriously get better.

This will just lead to massive selection bias. If all you care about are power users go for it. Otherwise you end up with a complicated system that is going to put off new users.

You can even have different power users who likes different parts of the software. Now it's totally possible to end up with extensive subsystems that don't really gel with each other

My anecdote for this is the the recentish redesign of Musescore. Tantacrul, the ux designer/product manager for Musescore has an hour long video on how he redesigned the interface and UX of the software. He is also a musician so you can call him a power users if you want but that was not the users he had in mind when redesigning the UX

https://youtu.be/Qct6LKbneKQ

rdedev··on Claude Opus 5
My codebase had a dataset with a bunch of SMILES strings and the word Malaria. Fable did not want to touch that codebase
Page 1 of 14Next →