HNHacker News
TopNewBestAskShowJobs

sweezyjeezy

2,032 karma · joined December 30, 2014

submissionscomments
sweezyjeezy··on The threat is comfortable drift toward not understanding what you're doing
I think this was an unpopular opinion mainly because it's scary, rather than there being an obvious reason to think otherwise.
sweezyjeezy··on What category theory teaches us about dataframes
> but it's really not such a crazy thing to have row labels built into your data table.

Sometimes you need data in a certain order. Sometimes there is no primary key. And it is nuts how janky the pandas API is if you just want the index to mean the current order of the dataframe and nothing else. Oh you did a pivot? I'm just going to make those pivot columns a row label now if that's alright with you. I don't do that for all functions though, you're going to have to remember which ones. Oh you want to sort a dataframe? You better make damn sure you reindex if you're planning to use that with data from another dataframe (e.g. x + y on data from separate dataframes), otherwise I'm going to align the data on indices, and you can't stop me. Also - want to call pyplot.plot(df['column'])? Yeah I'm giving it the data in index order obviously I don't care about that sort you just did. Oh you want to port this data to excel? Well if your row labels aren't meaningful and you don't want "Unnamed: 0" you're going to have to tell me not to. You need to manipulate a multi-index? You're so cute. Have fun with that buddy.

There is a reason no other dataframe library does this - because it's confusing and cognitive overhead that doesn't need to exist. I've used pandas since ~2013, had this chat with colleagues and many recommend just giving in and maintaining an index throughout. Except I've read their pandas and it sucks because now _you_ need to reason about what is currently the index - because it actually needs to change a lot to do normal things with data. I just use .reset_index copiously and try to make it behave like a normal dataframe library because it's just easier to understand later. Pandas has not earned the right to redefine what a dataframe means.

At the absolute least, index behaviour should be opt-in, not something imposed on the user.

sweezyjeezy··on What category theory teaches us about dataframes
The pandas API is awful, but it's kind of interesting why. It was started as a financial time series manipulation library ('panels') in a hedge fund and a lot of the quirks come from that. For example the unique obsession with the 'index' - functions seemingly randomly returning dataframes with column data as the index, or having to write index=False every single time you write to disk, or it appending the index to the Series numpy data leading to incredibly confusing bugs. That comes from the assumption that there is almost always a meaningful index (timestamps).
sweezyjeezy··on Oracle slashes 30k jobs
> After careful consideration of Oracle’s current business needs, we have made the decision to eliminate your role as part of a broader organizational change.

That is being laid off, not being fired - big difference. Being fired means being let go for poor performance / bad behaviour. No severance or grace period is necessary there (will be written in the contract). Being made redundant, particularly a redundancy of this size is quite well protected in EU. Typically negotiations between HR and representatives of the laid off group are required, you will continue to work (officially at least) until negotiations are over, as you are not officially out yet. This usually takes a few weeks.

I can tell you this from personal experience...

sweezyjeezy··on Social media bans and digital curfews to be trialled on UK teenagers
Honestly that is happening, and I think it's an overstep. I have never heard anyone talk about this in real life (that could be a London bubble though). I will say - I am nearly 40, and I've spent the last 25 years online reading about the Orwellian hell my life is (or is about to become). It has never felt like it comes from a place of lived experience. For example we infamously have a lot of cameras in the UK. 90% of them are on closed circuits in shops and it doesn't affect anything.

Right now the biggest issue in the UK is the same as most places - lack of money. It's killing our services, poisoning our politics. Everything else feels abstract in comparison.

sweezyjeezy··on Social media bans and digital curfews to be trialled on UK teenagers
It's their attention span. My SIL is an English professor and she stopped assigning long texts. The kids won't read it, will get an AI to summarize, and then give her poor reviews at the end for making them read.
sweezyjeezy··on Social media bans and digital curfews to be trialled on UK teenagers
Is HN in complete denial about what is happening to the younger generations right now? My whole family are teachers, and they are all sounding the alarm. A majority of kids are basically unable to read books now. Not just children - young adults studying English literature at college...

Parents are up against some of the wealthiest companies on earth, and the fear of socially excluding their kids by limiting their usage. Systemic change is never going to come from parents on this one.

sweezyjeezy··on Social media bans and digital curfews to be trialled on UK teenagers
I reckon in 20 years most countries will be doing this - the effects of social media on kids is too strong, and too negative to deny at this point.

Also in what way is the UK a police state? The amount of police is falling - we're strapped for cash...

sweezyjeezy··on Epoch confirms GPT5.4 Pro solved a frontier math open problem
To be fair, it is still pretty remarkable what the human brain does, especially in early years - there is no text embedded in the brain, just a crazily efficient mechanism to learn hierarchical systems. As far as I know, AI intelligence cannot do anything similar to this - it generally relies on giga-scaling, or finetuning tasks similar to those it already knows. Regardless of how this arose, or if it's relevant to AGI, this is still a uniqueness of sorts.
sweezyjeezy··on What young workers are doing to AI-proof themselves
Don't get fixated on plumbing itself. The point is if a bunch of people rush into any profession it leads to wage depression. Unless the amount of plumbing needed increases, the overall amount of money flowing to the plumbing populace is likely to stay roughly the same.
sweezyjeezy··on Waymo robotaxi hits a child near an elementary school in Santa Monica
These systems don't discriminate on whether the object is a child. If an object enters the path of the vehicle, the lidar should spot it immediately and the car should brake.
sweezyjeezy··on Travel Is Not Education
The trivia approach doesn't even work for most people - ask the wikipedia reader and the person who travelled to Turkey about it a year later and see who has actually retained some knowledge.
sweezyjeezy··on Using AI generated images to get refunds
That's a stretch. One can hold the view that division of labour is a useful economical principle, but also that oligopolies represent a dangerous concentration of power.
sweezyjeezy··on Trump says Venezuela’s Maduro captured after strikes
I think one of the best arguments against US interventionalism when it comes to tyrants is just how 'variable' (let's say) the outcomes have been over the years. For every Panama, there's two or three Guatamalas, Irans or most recently Iraq. Generally the hard part is not the removal of the head of state, which for the US is usually pretty quick. It's what beurocratic structures remain functional and whether the power vacuum created brings something better and more robust, or just decades of violence.
sweezyjeezy··on Koralm Railway
It's also really hard to make the tunnel remain a tunnel over its expected 150 year lifespan - given that it basically runs through a fault line. They had to study and test local geology for about 15 years, build certain sections to expect some movement over time, as well as kit everything out with a lot of sensors.

Overall an amazing achievement, and unsurprising it took this long to figure out!

sweezyjeezy··on The "confident idiot" problem: Why AI needs hard rules, not vibe checks
Exactly - 'non-deterministic' is not an accurate diagnosis of the issue.
sweezyjeezy··on The "confident idiot" problem: Why AI needs hard rules, not vibe checks
You could make an LLM deterministic if you really wanted to without a big loss in performance (fix random seeds, make MoE batching deterministic). That would not fix hallucinations.

I don't think using deterministic / stochastic as a diagnostic is accurate here - I think that what we're really talking is about some sort of fundamental 'instability' of LLMs a la chaos theory.

sweezyjeezy··on Gemini with Deep Think achieves gold-medal standard at the IMO
It's not it being case by case that's my issue. I used do olympiads and e.g. for the k>=3 case I wouldn't write much more than:

"Since there are 3k - 3 points on the perimeter of the triangle to be covered, and any sunny line can pass through at most two of them, it follows that 3k − 3 ≤ 2k, i.e. k ≤ 3."

Gemini writes:

Let Tk be the convex hull of Pk. Tk is the triangle with vertices V1 = (1, 1), V2 = (1, k), V3 = (k, 1). The edges of Tk lie on the lines x = 1 (V), y = 1 (H), and x + y = k + 1 (D). These lines are shady.

Let Bk be the set of points in Pk lying on the boundary of Tk. Each edge contains k points. Since the vertices are distinct (as k ≥ 2), the total number of points on the boundary is |Bk| = 3k − 3.

Suppose Pk is covered by k sunny lines Lk. These lines must cover Bk. Let L ∈ Lk. Since L is sunny, it does not coincide with the lines containing the edges of Tk. A line that does not contain an edge of a convex polygon intersects the boundary of the polygon at most at two points. Thus, |L ∩ Bk| ≤ 2. The total coverage of Bk by Lk is at most 2k. We must have |Bk| ≤ 2k. 3k − 3 ≤ 2k, which implies k ≤ 3.

sweezyjeezy··on Gemini with Deep Think achieves gold-medal standard at the IMO
100% o3 has a strong bias towards "write something that looks like a formal argument that appears to answer the question" over writing something sound.

I gave it a bunch of recent, answered MathOverflow questions - graduate level maths queries. Sometimes it would get demonstrably the wrong answer, but it not be easy to see where it had gone wrong (e.g. some mistake in a morass of algebra). A wrong but convincing argument is the last thing you want!

sweezyjeezy··on Gemini with Deep Think achieves gold-medal standard at the IMO
Gemini is clearer but MY GOD is it verbose. e.g. look at problem 1, section 2. Analysis of the Core Problem - there's nothing at all deep here, but it seems the model wants to spell out every single tiny logical step. I wonder if this is a stylistic choice or something that actually helps the model get to the end.
sweezyjeezy··on Bitcoin passes $120k milestone as US Congress readies for 'crypto week'
BTC is a (roughly) net-zero enterprise, every dollar taken out of the system comes from someone else putting a dollar in. Sure, if you had a crystal ball you could have made millions, but if everyone else ALSO had that same crystal ball you couldn't, since traders are mostly just shuffling money between themselves anyway.

There's no point kicking yourself over not foreseeing a far-fetched future scenario, if you were at a casino and a roulette spin landed on 12 - would you feel bad for not betting on that happening, despite having no good information it would land on that?

sweezyjeezy··on François Chollet: The Arc Prize and How We Get to AGI [video]
FWIW the original ARC was published in 2019, just after GPT-2 but a while before GPT-3. I work in the field, I think that discussing AGI seriously is actually kind of a recent thing (I'm not sure I ever heard the term 'AGI' until a few years ago). I'm not saying I know he didn't feel that, but he doesn't talk in such terms in the original paper.
sweezyjeezy··on Being too ambitious is a clever form of self-sabotage
100% - the quality group only had one chance to impress the teacher, whereas quantity group had dozens. The conclusion drawn from this in the text seems to be based on assumptions. We don't actually know how many intermediate photographs the quality group took as well, and without knowing that and also checking the quality of those, it's hard to say anything useful.
sweezyjeezy··on A deep critique of AI 2027's bad timeline models
I don't think the author of this article is making any strong prediction, in fact I think a lot of the article is a critique of whether such an extrapolation can be done meaningfully.

Most of these models predict superhuman coders in the near term, within the next ten years. This is because most of them share the assumption that a) current trends will continue for the foreseeable future, b) that “superhuman coding” is possible to achieve in the near future, and c) that the METR time horizons are a reasonable metric for AI progress. I don’t agree with all these assumptions, but I understand why people that do think superhuman coders are coming soon.

Personally I think any model that puts zero weight on the idea that there could be some big stumbling blocks ahead, or even a possible plateau, is not a good model.

sweezyjeezy··on P-Hacking in Startups
Yes 100% this. If you're comparing two layouts there's no great reason to treat one as a 'treatment' and one as a 'control' as in medicine - the likelihood is they are both equally justified. If you run an experiment and get p=0.93 on a new treatment - are you really going to put money on that result being negative, and not updating the layout?

The reason we have this stuff in medicine is because it is genuinely important, and because a treatment often has bad side-effects, it's worse to give someone a bad treatment than to give them nothing, that's the point of the Hypocratic oath. You don't need this for your dumb B2C app.

sweezyjeezy··on P-Hacking in Startups
This is a subtle point that even a lot of scientists don't understand. A p value or < 0.05 doesn't mean "there is less than a 5% chance the treatment is not effective". It means that "if the treatment was only as effective, (or worse) than the original, we'd have < 5% chance of seeing results this good". Note that in the second case we're making a weaker statement - it doesn't directly say anything about the particular experiment we ran and whether it was right or wrong with any probability, only about how extreme the final result was.

Consider this example - we don't change the treatment at all, we just update its name. We split into two groups and run the same treatment on both, but under one of the two names at random. We get a p value of 0.2 that the new one is better. Is it reasonable to say that there's a >= 80% chance it really was better, knowing that it was literally the same treatment?

sweezyjeezy··on How Frogger 2’s source code was recovered from a destroyed tape [video]
It's teenager-ish, and the kind of thing that would make people uncomfortable if they saw it at our own workplaces. We can argue about whether companies 'should' punish people for stuff like this, but I can say for sure that I don't feel like I'm missing out on much here.
sweezyjeezy··on Gemini 2.5 Pro Preview
Only skimmed, but both seem to be referring to what transformers can do in a single forward pass, reasoning models would clearly be a way around that limitation.

o4 has no problem with the examples of the first paper (appendix A). You can see its reasoning here is also sound: https://chatgpt.com/share/681b468c-3e80-8002-bafe-279bbe9e18.... Not conclusive unfortunately since this is in date-range of its training data. Reasoning models killed off a large class of "easy logic errors" people discovered from the earlier generations though.

sweezyjeezy··on Gemini 2.5 Pro Preview
I'm not betting any money here - extrapolation is always hard. But just drawing a mental line from here that tapers to somewhere below one's own abilities - I'm not seeing a lot of justification for that either.
sweezyjeezy··on Gemini 2.5 Pro Preview
A couple of comments

a) it hasn't even been a year since the last big breakthrough, the reasoning models like o3 only came out in September, and we don't know how far those will go yet. I'd wait a second before assuming the low-hanging fruit is done.

b) I think coding is a really good environment for agents / reinforcement learning. Rather than requiring a continual supply of new training data, we give the model coding tasks to execute (writing / maintaining / modifying) and then test its code for correctness. We could for example take the entire history of a code-base and just give the model its changing unit + integration tests to implement. My hunch (with no extraordinary evidence) is that this is how coding agents start to nail some of the higher-level abilities.

← PreviousPage 2 of 17Next →