HNHacker News
TopNewBestAskShowJobs

syllogism

3,512 karma · joined October 18, 2010

NLP researcher, recently switched to independent developer.

http://honnibal.github.io/spaCy http://honnibal.wordpress.com http://scholar.google.com.au/citations?user=FXwlnmAAAAAJ&hl=en http://github.com/honnibal/ http://github.com/syllog1sm/

submissionscomments
syllogism··on Grim Fandango Puzzle Document (1996) [pdf]
It released with unbearable keyboard controls instead of the familiar point-and-click. If I remember right a remaster eventually came out that allowed mouse control.

I agree that I hated the 3d though. This was a bad patch for games when everyone wanted to go 3d for the sake of it, even though it looked and played worse.

syllogism··on More questions about whether researchers can trust OpenAI with unpublished math
You can't take any statement like this remotely seriously. We live in a world where NSA officials can testify before congress that they don't "collect" data, because that's true under some baroque definition of "collect" that they invented and didn't tell anyone else about.

Similarly you have no idea what definition OpenAI intend for terms such as "specific user data", "accessed" etc. And we have no idea what non-excluded possibilities actually did happen that they simply omit from their statement.

In practice OpenAI and many others have created a situation where they're actually unable to make any credible denial of anything really.

syllogism··on “I just chose words carefully”
If you're submitting an eight page conference paper you're very often hunting runts (short lines ending paragraphs) and rewording for better justification, especially on the first page and even more so in the abstract.

First-page space is at a premium because you want the reviewer to scan the introduction well, and it's great if you can fit a key figure that explains the concept on the first page.

The game's probably very different now that there's such a flood of submissions but when I was reviewing I always saw messy alignment on the first page as a sign the paper was probably rushed, and usually I could see that elsewhere in the paper too.

syllogism··on Early-life stress leaves a 'scar' inside brain cells in mice
The optimal amount of stress about that stuff is not zero, and most teenagers are going to be desperately biased towards "none of this matters, just peermax".
syllogism··on Where Human Sleep Went Wrong
This interview feels weird to me, maybe because there's lots of places where I'm thinking there's an obvious idea that isn't raised.

Like, it seems obvious to me that sleep has a higher opportunity cost in humans than most other animals. So naturally there's more evolutionary pressure in humans to sleep more efficiently.

Is this actually true? Dunno! But it was really weird that it doesn't get brought up.

syllogism··on It is time to give up the dualism introduced by the debate on consciousness
It comes up because people are mistakenly conflating consciousness with moral personhood.

People want to be talking about whether AI suffers in a morally meaningful way. In non-human animals this debate is often centered around the question of whether the animal has conscious experience, because there's little doubt that much of the emotional and experiential systems are shared.

The analogy goes wrong with AI, where definitions of "consciousness" would seem to apply in the sense that the model clearly has a category for itself in its world model, feeds back on its output, etc. However the analogy between how it works and anything we would recognise as emotion or suffering is extremely strained.

The solution is to just focus on ths question of what we really mean when we think of morally relevant suffering. It's a much clearer question than "consciousness" and it sidesteps the problem.

syllogism··on Ask HN: Founders of estonian e-businesses – is it worth it?
You're going to make life much harder for yourself, not easier, because you'll still need German legal advice but now you need an expensive multi-national lawyer/firm. Anyone cheaper will refuse to touch it.

Germany cares about where the management of your company actually happens, not just where the entity is incorporated. So you're not going to avoid German bureaucracy, it's going to be worse not better.

syllogism··on LiteLLM Python package compromised by supply-chain attack
Just keeping a lockfile and updating it weekly works fine for that too yeah
syllogism··on Tell HN: Litellm 1.82.7 and 1.82.8 on PyPI are compromised
Maintainers need to keep a wall between the package publishing and public repos. Currently what people are doing is configuring the public repo as a Trusted Publisher directly. This means you can trigger the package publication from the repo itself, and the public repo is a huge surface area.

Configure the CI to make a release with the artefacts attached. Then have an entirely private repo that can't be triggered automatically as the publisher. The publisher repo fetches the artefacts and does the pypi/npm/whatever release.

syllogism··on Judge orders government to begin refunding more than $130B in tariffs
The corruptions of this administration are legion, but this isn't one of them. Unless you can point to something Lutnick did to create this outcome, I don't see how he had a better view of the whole thing than anyone else.
syllogism··on OpenAI agrees with Dept. of War to deploy models in their classified network
Government's free to not like the terms and go with another provider. That's whatever.

Government's not free to say, "We'll blow up your business with a false accusation if you don't give us the terms we want (and then use defence production act to commandeer the product anyway)". How much more blatantly authoritarian does it get than that?

syllogism··on OpenAI agrees with Dept. of War to deploy models in their classified network
If you think a blue government would even consider threatening to falsely accuse a company of being a supply-chain threat in order to gain leverage in a contract negotiation, you're insane. There's nothing remotely normal about this, it's not something you see in any western democracy
syllogism··on OpenAI – How to delete your account
The actions of the US government here are openly corrupt.

The point of the supply chain risk provisions is to denote, you know, supply chain risks. The intention is not to give the Pentagon a lever it can pull to force any company to agree to any contract it wants.

Hegseth doesn't even pretend that Anthropic is actually a supply chain risk. The argument for designating them so is that _they won't do exactly what the government wants_.

People use the term "fascism" a lot and people have kind of tuned it out, but what do you call a government that deals itself the power to compel any company to accept any contract, and declare it a pariah on thin pretext if it objects?

By taking the deal under these conditions OpenAI is accepting this. They're saying, "Well, sucks to be them, life goes on". They're consenting to the corruption and agreeing to profit from it. But they'll be next, and if the next company in line has the same stand then yeah, the government can force any company to do anything. There's nothing normal about this.

syllogism··on OpenAI agrees with Dept. of War to deploy models in their classified network
I don't understand how any sort of deal is defensible in the circumstances.

Government: "Anthropic, let us do whatever we want"

Anthropic: "We have some minimal conditions."

Government: "OpenAI, if we blast Anthropic into the sun, what sort of deal can we get?"

OpenAI: "Uh well I guess I should ask for those conditions"

Government: blasts Anthropic into the sun "Sure whatever, those conditions are okay...for now."

By taking the deal with the DoW, OpenAI accepts that they can be treated the same way the government just treated Anthropic. Does it really matter what they've agreed?

syllogism··on OpenAI agrees with Dept. of War to deploy models in their classified network
You should quit because the only reasonable thing for your leadership to have done is to refuse to sign any agreement with DoW whatsoever while it's attempting to strongarm Anthropic in this fashion.

It doesn't even matter if OpenAI is offered the same terms that Anthropic refused. It's absurd to accept them and do business with the Pentagon in that situation.

If you take the government at its word, it's killing Anthropic because Anthropic wanted to assert the ability to draw _some_ sort of redline. If OpenAI's position is "well sucks to be them", there's nothing stopping Hegseth from doing the same to OpenAI.

It doesn't matter at all if OpenAI gets the deal at the same redline Anthropic was trying to assert. If at the end of this the government has succeeded in cutting Anthropic off from the economy, what's next for OpenAI? What happens next time when OpenAI tries to assert some sort of redline?

What's the point of any talk of "AI Safety" if you sign on to a regime where Hegseth (of all people) can just demand the keys and you hand them right over?

syllogism··on We Will Not Be Divided
They're labelling Anthropic a supply chain risk, without even the pretense that this is in fact true. They're perfectly content to use the tool _themselves_, but they claim that an unwillingness to sign whatever ToS DoW asks marks the company a traitor that should be blacklisted from the economy.
syllogism··on A Remarkable Assertion from A16Z
I thought it was a joke? Like the reviewer is saying, "I didn't finish these books".
syllogism··on I don't care how well your "AI" works
If I were a CTO or VP these days I think I'd push for a blanket ban on committing docs/readmes/diagrams etc along with the initial work. Teams can push stuff to a `slop/` folder but don't call it docs.

If you push all that stuff at the same time, it's really easy to get away with this soft lie, "job done". They can claim they thought it was okay and it was just an honest mistake there were problems. They can lie about how much work they really did.

READMEs or diagrams that are plans for the functionality are fine. Docs that describe finished functionality are fine. Slop that dresses up unfinished work as finished work just fucks everything up, and the incentives are misaligned so everyone's doing this.

syllogism··on I don't care how well your "AI" works
Well, if you take "review the LLM output" in its most general way, I guess you can class everything under that. But I think it's worth talking about the problem in a bit more detail than that, because someone can easily say "Oh I definitely review the LLM output!" and still be pushing work onto other people.

The fact is that no matter whether we review the LLM output or not, no matter whether we write the code entirely by hand or not, there's always going to be the possibility of errors. So it's not some bright-line thing. If you're relatively lazier and relatively less thoughtful in the way you work, you'll make more errors and more significant errors. You'll look like you're doing the work, but your teammates have to do more to make up for the problems.

Having to work around problems your coworkers introduced is nothing new, but LLMs make it worse in a few ways I think. One is just, that old joke about there being four kinds of people: lazy and stupid, industrious and stupid, smart and lazy, and industrious and smart. It's always been the "industrious and stupid" people that kill you, so LLMs are an obvious problem there.

Second there's what I call the six-fingered hands thing. LLMs make mistakes a human wouldn't, which means the problem won't be in your hypothesis-space when you're debugging.

Third, it's very useful to have unfinished work look unfinished. It lets you know what to expect. If there's voluminous docs and tests and the functionality either doesn't work at all or doesn't even make sense when you think about it, that's going to make you waste time.

Finally, at the most basic level, we expect there to be some sort of plan behind our coworkers' work. We expect that someone's thought about this and that the stuff they're doing is fundamentally going to be responsive to the requirements. If someone's phoning it in with an LLM, problems can stay hidden for a long time.

syllogism··on I don't care how well your "AI" works
I think LLMs are net helpful if used well, but there's also a big problem with them in workplaces that needs to be called out.

It's really easy to use LLMs to shift work onto other people. If all your coworkers use LLMs and you don't you're gonna get eaten alive. LLMs are unreasonably effective at generating large volumes of stuff that resembles diligent work on the surface.

The other thing is, tools change trade-offs. If you're in a team that's decided to lean into static analysis, and you don't use type checking in your editor, you're getting all the costs and less of the benefits. Or if you're in a team that's decided to go dynamic, writing good types for just your module is mostly a waste of time.

LLMs are like this too. If you're using a very different workflow from everyone else on your team, you're going to end up constantly arguing for different trade-offs, and ultimately you're going to cause a bunch of pointless friction. If you don't want to work the same way as the rest of the team just join a different team, it's really better for everyone.

syllogism··on I don't care how well your "AI" works
Software to date has been a [Jevons good](https://en.wikipedia.org/wiki/Jevons_paradox). Demand for software has been constrained by the cost efficiency and risk of software projects. Productivity improvements in software engineering have resulted in higher demand for software, not less, because each improvement in productivity unblocks more of the backlog of projects that weren't cost effective before.

There's no law of nature that says this has to continue forever, but it's a trend that's been with us since the birth of the industry. You don't need to look at AI tools or methodoligies or whatever. We have code reuse! Productivity has obviously improved, it's just that there's also an arms race between software products in UI complexity, features, etc.

If you don't keep improving how efficiently you can ship value, your work will indeed be devalued. It could be that the economics shift such that pretty much all programming work gets paid less, it could be that if you're good and diligent you do even better than before. I don't know.

What I do know is that whichever way the economics shake out, it's morally neutral. It sounds like the author of this post leans into a labor theory of value, and if you buy into that, well...You end up with some pretty confused and contradictory ideas. They position software as a "craft" that's valuable in itself. It's nonsense. People have shit to do and things they want. It's up to us to make ourselves useful. This isn't performance art.

syllogism··on Chess grandmaster Daniel Naroditsky has died
This has been discussed to death but to reiterate here: people have been much too polite about Kramnick's nonsense. Danya and Hikaru are (were) probably the two people in the world whose bullet play is least suspicious. Cheating isn't very powerful in short time controls, and they have streamed thousands of games playing them.

Kramnick's bullshit never made any damn sense at all.

syllogism··on Want to piss off your IT department? Are the links not malicious looking enough?
In Europe there are legitimate and extremely established services that require you to input your bank login details into something other than your bank's website. It's madness.
syllogism··on Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
What is "real reasoning"? The mechanism that the models use is well described. They do what they do. What is this article's complaint?
syllogism··on Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
It's interesting that there's still such a market for this sort of take.

> In a recent pre-print paper, researchers from the University of Arizona summarize this existing work as "suggest[ing] that LLMs are not principled reasoners but rather sophisticated simulators of reasoning-like text."

What does this even mean? Let's veto the word "reasoning" here and reflect.

The LLM produces a series of outputs. Each output changes the likelihood of the next output. So it's transitioning in a very large state space.

Assume there exists some states that the activations could be in that would cause the correct output to be generated. Assume also that there is some possible path of text connecting the original input to such a success state.

The reinforcement learning objective reinforces pathways that were successful during training. If there's some intermediate calculation to do or 'inference' that could be drawn, writing out a new text that makes that explicit might be a useful step. The reinforcement learning objective is supposed to encourage the model to learn such patterns.

So what does "sophisticated simulators of reasoning-like text" even mean here? The mechanism that the model uses to transition towards the answer is to generate intermediate text. What's the complaint here?

It makes the same sort of sense to talk about the model "reasoning" as it does to talk about AlphaZero "valuing material" or "fighting for the center". These are shorthands for describing patterns of behaviour, but of course the model doesn't "value" anything in a strictly human way. The chess engine usually doesn't see a full line to victory, but in the games it's played, paths which transition through states with material advantage are often good -- although it depends on other factors.

So of course the chain-of-thought transition process is brittle, and it's brittle in ways that don't match human mistakes. What does it prove that there are counter-examples with irrelevant text interposed that cause the model to produce the wrong output? It shows nothing --- it's a probabilistic process. Of course some different inputs lead to different paths being taken, which may be less successful.

syllogism··on GCP Outage
Like a backup generator for inputs. Makes sense.
syllogism··on Why agents are bad pair programmers
LLM agents are very hard to talk about because they're not any one thing. Your action-space in what you say and what approach you take varies enormously and we have very little body of common knowledge about what other people are doing and how they're doing it. Then the agent changes underneath you or you tweak your prompt and it's different again.

In my last few sessions I saw the efficacy of Claude Code plummet on the problem I was working on. I have no idea whether it was just the particular task, a modelling change, or changes I made to the prompt. But suddenly it was glazing every message ("you're absolutely right"), confidently telling me up is down (saying things like "tests now pass" when they completely didn't), it even cheerfully suggested "rm db.sqlite", which would have set me back a fair bit if I said yes.

The fact that the LLM agent can churn out a lot of stuff quickly greatly increases 'skill expression' though. The sharper your insight about the task, the more you can direct it to do something specific.

For instance, most debugging is basically a binary search across the set of processes being conducted. However, the tricky thing is that the optimal search procedure is going to be weighted by the probability of the problem occurring at the different steps, and the expense of conducting different probes.

A common trap when debugging is to take an overly greedy approach. Due to the availability heuristic, our hypotheses about the problem are often too specific. And the more specific the hypothesis, the easier it is to think of a probe that would eliminate it. If you keep doing this you're basically playing Guess Who by asking "Is it Paul? Is it Anne?" etc, instead of "Is the person a man? Does the person have facial hair? etc"

I find LLM agents extremely helpful at forming efficient probes of parts of the stack I'm less fluent in. If I need to know whether the service is able to contact the database, asking the LLM agent to write out the necessary cloud commands is much faster than getting that from the docs. It's also much faster at writing specific tests than I would be. This means I can much more neutrally think about how to bisect the space, which makes debugging time more uniform, which in itself is a significant net win.

I also find LLM agents to be good at the 'eat your vegetables' stuff -- the things I know I should do but would economise on to save time. Populate the tests with more cases, write more tests in general, write more docs as I go, add more output to the scripts, etc.

syllogism··on My experiment living in a tent in Hong Kong's jungle
Homelessness is a somewhat broad category though. There's lots of people couch-surfing between friends and their car. They're also in a very different position from people who are sleeping rough.
syllogism··on Retailers will soon have only about 7 weeks of full inventories left
Carville (DNC strategist) is advocating a "play dead" strategy. Let Trump implode so that he owns the inevitable failure. His base will desperately want to blame the left for not letting the policies work as intended. The less the Democrats do, the harder that is. I think a lot of Democrat politicians are going this way, and it's why Schumer rolled over on the budget.

Part of the logic here is that Trump is indeed different from other authoritarians. He's even less competent. He's blowing all his political capital on imploding the economy. He also can't understand the legal battles, so when Stephen Miller tells him they won the Supreme Court case 9-0, he believes him. This seems to have been a big wake-up call to Gorsuch, Coney-Barrett and Kavanaugh. The administration has shown its hand much too quickly, before it fully consolidated its power.

What the Democrats should be doing already is campaigning more. Run ads that are literally just Trump quotes. Show people Trump calling January the "Trump economy" before inauguration, then calling April the "Biden economy" now that he's crashed it. If Trump polls low enough, more senators will jump ship, and impeachment could be possible.

syllogism··on Gukesh becomes the youngest chess world champion in history
It wasn't even so much the blunders as the strategic decisions I think. Like, a blunder isn't in itself "baffling".
Page 1 of 21Next →