HNHacker News
TopNewBestAskShowJobs

nicklecompte

2,008 karma · joined August 9, 2018

submissionscomments
nicklecompte··on How Might We Learn?
It is frustrating to read through demos which feature a hypothetical AI that greatly exceeds the capacities of any actual LLM, and which does not consider the serious risks of learners getting misled by confabulations.

It is especially frustrating when I recently tested GPT-4o on a factual question and got 1000 words which were all completely wrong, including fake citations.

It is especially frustrating to read this sci-fi daydreaming after talking to a high school science teacher who was forced to use generative AI tutors in their classes this year, even though these tutors are poorly tested and seem to have a ~20% confabulation rate. This particular teacher is technically sophisticated but even they sometimes get confused and misled by the chatbot. Students don't have a chance.

I think Matuschak has valuable insights on learning in general. But it seems incomplete to go through this AI thought experiment without discussing how inadequate current AI is to the task. "Technology will get better" but what if it takes 50 years?

nicklecompte··on Statement from Scarlett Johansson on the OpenAI "Sky" voice
Let me start by saying I despise generative AI and I think most AI companies are basically crooked.

I thought about your comment for a while, and I agree that there is a fine line between "realistic parody" and "intentional deception" that makes deepfake AI almost impossible to defend. In particular I agree with your distinction:

- In matters involving human actors, human-created animations, etc, there should be great deference to the human impersonators, particularly when it involves notable public figures. One major difference is that, since it's virtually impossible for humans to precisely impersonate or draw one another, there is an element of caricature and artistic choice with highly "realistic" impersonations.

- AI should be held to a higher standard because it involves almost no human expression, and it can easily create mathematically-perfect impersonations which are engineered to fool people. The point of my comment is that fair use is a thin sliver of what you can do with the tech, but it shouldn't be stamped out entirely.

I am really thinking of, say, the Joe Rogan / Donald Trump comedic deepfakes. It might be fine under American constitutional law to say that those things must be made so that AI Rogan / AI Trump always refer to each other in those ways, to make it very clear to listeners. It is a distinctly non-libertarian solution, but it could be "necessary and proper" because of the threat to our social and political knowledge. But as a general principle, those comedic deepkfakes are works of human political expression, aided by a fairly simple computer program that any CS graduate can understand, assuming they earned their degree honestly and are willing to do some math. It is constitutionally icky (legal term) to go after those people too harshly.

nicklecompte··on Building an AI game studio: what we've learned so far
This is an unacceptably lazy and reckless comment: your very first example blatantly plagiarized from Star Wars. You even called the bombs "BB-8 bombs" to make them look like the robot! It's just a dumb demo, but it means your product has no safeguards against copyright infringement, and you don't even care.

I suspect you might say "users of our technology assume full responsibility for copyright violations" or whatever, but what are your customers supposed to do if your product plagiarizes something more obscure than Star Wars? They can't be expected to check every generative AI output against every IP. You and your team need to take responsibility for the copyright problems your tool is guaranteed to create as currently implemented.

I really hate this "better to seek forgiveness" approach to copyright OpenAI has encouraged.

nicklecompte··on Statement from Scarlett Johansson on the OpenAI "Sky" voice
Right - the only reason Murati's behavior was "a tell" was that rational observers had a good reason to believe her statement was evasive and misleading, regardless of how she said it. In that context her emotional response is a funny (yet only barely rational) indication that she wasn't being honest. But trying to go the other direction - "her emotional response seems dishonest, what part of her statement was a lie?" - is as scientific as astrology[1], and a disaster for real people to use in real life. You'll quickly end up accusing innocent people for mean and bigoted reasons.

[1] Modern quack science might be worst than modern astrology/etc, since the stance of "I know it's not real, but..." means astrology folks can freely disregard horoscopes that are socially or ethically objectionable. If you're claiming your BS is actually Science, I think for a lot of people there's a sort of vicious feedback loop. Especially with social stuff, where these claims are essentially unfalsifiable. ("Body language reading is Science and cannot fail, it can only be failed by my incompetence.")

nicklecompte··on Statement from Scarlett Johansson on the OpenAI "Sky" voice
From the Ars Technica story[1], this is very funny:

> But OpenAI's chief technology officer, Mira Murati, has said that GPT-4o's voice modes were less inspired by Her than by studying the "really natural, rich, and interactive" aspects of human conversation, The Wall Street Journal reported.

People made fun of Murati when she froze after being asked what Sora was trained on. But behavior like that indicates understanding that you could get the company sued if you said something incriminating. Altman just tweets through it.

[1] https://arstechnica.com/tech-policy/2024/05/openai-pauses-ch...

nicklecompte··on Statement from Scarlett Johansson on the OpenAI "Sky" voice
I think anti-deepfake legislation needs to consider fair use, especially when it comes to parody or other commentary on public figures. OpenAI's actions do not qualify as fair use.
nicklecompte··on Ask HN: Most successful example using LLMs in daily work/life?
I tried using GPT-4 as a better way to search papers - it can be very annoying when you know the gist of a result but not the authors or enough details about the methodology for Google. GPT-4 was pretty good at figuring out what citation I wanted given a vague description.

However, the confabulation/hallucination rate seemed highly subject-dependent: AI/ML citations were quite robust, but cognitive science was so bad that it wasn't worth using. Eventually I went back to the Old Ways. But there are a good number of academics that use it as an alternative to Google Scholar.

nicklecompte··on Cognitive reflection, intelligence, and cognitive abilities: A meta-analysis
This is just a distortion, and shows you haven't actually engaged with the scientific criticism of IQ tests:

> General intelligence was discovered empirically, not hypothesized and then searched-for.

General intelligence has been hypothesized - and assumed to exist - by almost all peoples since prehistory. The g factor is what you are referring to, and that was discovered "empirically." What was not at all empirical was the decision to call the g factor a measure of general intelligence. That is a hypothesis which seems flatly wrong, and I have yet to see an argument for its validity that isn't circular or specious. My point is that written cognitive tests don't measure intelligence, they measure something much more shallow. The fact that written cognitive tests have a correlating g factor does not mean that the g factor itself is any less shallow.

nicklecompte··on Recall is Microsoft's key to unlocking the future of PCs
"To ensure your privacy against malicious use, we seed about 15% of the logs with fabricated entries that don't actually correspond with what you were doing on your computer."

"Wait a minute... doesn't GPT-4 still have a 15% hallucination rate when it comes to document summarization?"

"Moving right along..."

nicklecompte··on Grothendieck’s use of equality
I think HOTT helps remove a lot of the footguns and computational hairiness around using type-theoretic mathematics. But the issues Buzzard is talking about are really more profound than that. This relates to a much deeper (but familiar and nontechnical) philosophical problem: you assign a symbol to a thing, and read two sentences that refer to the same symbol ("up to isomorphism"), but not necessarily the same thing. If the sentences contradict each other, is it because they disagree on the meaning of the symbol, or do they disagree on the essence of the thing? There is a strong argument that most debates about free will and consciousness are really about what poorly-understood thing(s) we are assigning to the symbol. Worse, this makes such debates ripe for bad faith, where flaws in an argument are defended by implicitly changing the meaning of the symbol.

In the context of """simple""" mathematics, preverbal toddlers and chimpanzees clearly have an innate understanding of quantity and order. It's only after children fully develop this innate understanding that there's any point in teaching them "one," "two," "three," and thereby giving them the tools for handling larger numbers. I don't think it makes sense to say that toddlers understand the Peano axioms. Rather, Peano formulated the axioms based on his own (highly sophisticated) innate understanding of number. But given he spent decades of pondering the topic, it seems like Peano's abstract conception of "number" became different from (say) Kronecker's, or other constructivists/etc. Simply slapping the word "integer" on two different concepts and pointing out that they coincide for quantities we can comprehend doesn't actually do anything by itself to address the discrepancy in concept revealed by Peano allowing unbounded integers and Kronecker's skepticism. (The best argument against constructivism is essentially sociological and pragmatic, not "mathematically rational.")

Zooming out a bit, I suspect we (scientifically-informed laypeople + many scientists) badly misunderstand the link between language and human cognition. It seems more likely to me that we have extremely advanced chimpanzee brains that make all sorts of sophisticated chimpanzee deductions, including the extremely difficult question of "what is a number?", but to be shared (and critically investigated) these deductions have to be squished into language, as a woefully insufficient compromise. And I think a lot of philosophical - and metamathematical - confusion can be understood as a discrepancy between our chimpanzee brains having a largely rigorous understanding of something, but running into limits with our Broca's and Wernicke's areas, limits which may or may not be fixed by "technological development" in human language. (Don't even get me started on GPT...)

nicklecompte··on Cognitive reflection, intelligence, and cognitive abilities: A meta-analysis
It seems to me like it is yet another study which refutes the idea that we're able to scientifically measure "general intelligence" (or "cognitive intelligence," using the author's term). Instead of anything which can be called intelligence, these tests keep measuring the same wrong thing. This thing is clearly amenable to short-term improvement with practice, and possibly assistance from a trained psychologist. And I don't think it has much to do with the actual problems human brains have to solve. The fact that human brains can be trained to solve IQ test problems or math problems is interesting. But surely it's a tiny slice of what human brains can do.
nicklecompte··on Llama3 implemented from scratch
Big Government Socialism won't let you build your own 25km-circumference particle accelerator. Bureaucrats make you fill out "permits" and "I-9s for the construction workers instead of hiring undocumented day laborers."

I am wondering if "CERN was pushed on the masses by the few" is an oblique reference to public fears that the LHC would destroy the world.

nicklecompte··on Llama3 implemented from scratch
The Big Dig (Boston highway overhaul) cost $22bn in 2024 dollars. The Three Gorges dam cost $31bn. These are expensive infrastructure projects (including the infrastructure for data centers). It doesn't say anything about how important they are for society.

Comparing LLMs to the Manhattan Project based on budget alone is stupid and arrogant. The comparison only "makes itself" because Ethan Mollick is a childish and unscientific person.

nicklecompte··on Meteor seen in Portugal
If I had to die in a plane crash, this seems like the least terrifying way to go out. Way better than hearing an ominous rumbling, scrambling for the oxygen masks…
nicklecompte··on Llama3 implemented from scratch
One other thing to add is large-scale RLHF. Big Tech can pay literally hundreds of technically-sophisticated people throughout the world (e.g. college grads in developing countries) to improve LLM performance on all sorts of specific problems. It is not a viable way to get AGI, but it means your LLM can learn tons of useful tricks that real people might want, and helps avoid embarrassing "mix broken glass into your baby formula" mistakes. (Obviously it is not foolproof.)

I suspect GPT-4's "secret sauce" in terms of edging out competitors is that OpenAI is better about managing data contractors than the other folks. Of course it's a haze of NDAs to learn specifics, and clearly the contractors are severely underpaid compared to OpenAI employees/executives. But a lone genius with a platinum credit card can't create a new world-class LLM without help from others.

nicklecompte··on AI doppelgänger experiment – Part 1: The training
I don't think there is a slippery slope in the foreseeable future, unless you buy into sci-fi views of AI being like a human. We need to update the laws around copyright in response to these machines. It's similar to why copyright laws exist in the first place: the concept was developed in response to the printing press.

At some point there will be a real I, Robot problem about an AI artist that actually understands what its drawing and doesn't depend on interpolated plagiarism of inhuman amounts of data. But we aren't even close to that yet.

nicklecompte··on OpenAI created a team to control 'superintelligent' AI – then let it wither
> which can be debated, but it's a different topic

No, it "can't be debated," it is clearly false! You said "by definition," but you used an irrational and bigoted definition of "general reasoning and logic" which conflates such things with performance on a standardized test. Humans aren't innately good at stupid logic puzzles that LLMs might get a 71st percentile in. Our brains are not actually designed to solve decontextualized riddles. That's a specialized skill which can be practiced. It's depressing enough when people claim IQ tests are actually good measures of human intelligence, despite overwhelming evidence to the contrary. But now, by even worse reasoning, we have people saying a computer is smarter than "average humans." (MTurk average humans? Undergrads? Who cares!) The complete lack of skepticism and scientific thinking on display by many AI developers/evangelists is just plain depressing.

Let me add that a truly humiliating number of those """general reasoning""" LLM benchmarks are fucking multiple choice questions! Not all of them, but a lot. ML critics have been complaining since ~2017 (BERT) that LLMs pick up on spurious statistical correlations in benchmarks but fail badly in real-world examples that use slightly different language. Using a multiple choice test is simply dishonest, like a middle finger to scientific criticism.

nicklecompte··on OpenAI created a team to control 'superintelligent' AI – then let it wither
I don't think any of us will live to see AI smarter than a rat. So I am not concerned whether this team was going to superalign anything.

The problem is that OpenAI's software, especially GPT-4o, is primed for dangerous misuse, and the demo videos of GPT-4o seemed "misaligned" with any reasonable standards of AI safety. Even if Leike/Sutskever have delusions of grandeur about AGI, at least they cared about the idea of AI safety. It seems like pushing the team out meant getting rid of a lot of internal critics (and implicitly threatening anyone else who might speak up).

nicklecompte··on LLM-generated code must not be committed without prior written approval by core
If my subordinate was copying code from StackOverflow without attribution I would be annoyed enough to send a grouchy email. Behavior like that is bad hacker citizenship, and bad for long-term maintenance. You should at least include a hyperlink to the SO question.

I also think SO is different about mindless copy-pasting. Outside of rote beginner stuff it’s infrequent that someone has the exact same question as you, and that the best answer works by simple copy-pasting. Often the modification is simple enough that even GPT can do it :) But making sure the SO question is relevant, and modifying the answer accordingly, is a check on understanding that LLMs don’t really have. In particular, a SO answer might be “wildly wrong” syntactically but essentially correct semantically. LLMs can give you the exact opposite problem.

nicklecompte··on OpenAI putting 'shiny products' above safety, says departing researcher
I somehow missed this:

> “Building smarter-than-human machines is an inherently dangerous endeavour. OpenAI is shouldering an enormous responsibility on behalf of all of humanity,” Leike wrote.

Leike clearly did the right thing by resigning, GPT-4o is dangerous and irresponsible. But if that tweet is how OpenAI employees actually think of themselves and their technology...... yeesh.

nicklecompte··on If you’re seeing this, I’m in jail [video]
Probably because the first Boeing whistleblower died by suicide and the second Boeing whistleblower died due to a stroke secondary to a respiratory infection. The idea that either of these were murders is a stupid conspiracy theory with no supporting evidence, and it deserves to be heavily downvoted.
nicklecompte··on LLM-generated code must not be committed without prior written approval by core
The are two salient differences here:

1) You are a thinking adult - or a thoughtful teenager :) - and therefore you understand copyright law well enough to take responsibility for copyright law, and in particular you are capable of having standing in legal matters. AI is not, but it is quite capable of creating infringing code (eg GNU stuff) that human reviewers wouldn't even know was infringing. So it is much better for any honest and competent organization to ban commercial LLMs entirely. (I am fine with in-house solutions with 100% validated training data...but those aren't very good yet, are they?)

2) I am a broken record on this, but the biggest problem with the "stochastic parrots" analogy is that transformer ANNs are dramatically dumber than parrots, or any other jawed vertebrate (I am not sure about lampreys). As applied to code generation: when I first tested ChatGPT-3.5, I was shocked to discover it was plagiarizing hundreds of lines of F#, verbatim, including from my own GitHub. Obviously that's outrageous in terms of OpenAI's ethics. But it is also amazing how dumb the AI is! Imagine a human programmer who is highly proficient in Python, and pretty good at Haskell, yet despite reading every public F# project in GitHub it can't solve intermediate F# problems without shameless copying.

It is a completely misleading comparison to say that humans reading source code is anything like transformers learning patterns in text. The most depressing thing about the current AI bubble is watching tech folks devalue human intelligence - especially since the primary motivation is excusing the failures of a computer which is less intelligent than a single honeybee.

nicklecompte··on LLM-generated code must not be committed without prior written approval by core
Isn't "mainly in how it destroys thinking" precisely what the parent comment was referring to with the Torvalds quote? I am not sure what your disagreement even is.
nicklecompte··on LLM-generated code must not be committed without prior written approval by core
The headline of this HN post is misleading, it isn't really a "ban" as such and doesn't have to be "enforceable." These are guidelines whose target audience is good-faith developers who want to make positive contributions to NetBSD.

The comments saying "heh, those stoopid NetBSD maintainers would never know if I slipped a bit of AI into their codebase" are missing the point. The point of the guideline is to say that people who use Copilot/etc for NetBSD development are dumb assholes, so don't be a dumb asshole.

nicklecompte··on Thinking out loud about 2nd-gen email
> And neither of these companies have any interest in innovating

I don’t think this is true, or at least I think they have a strong interest in standardizing. Enterprise and personal users are routinely frustrated with Outlook and Gmail for dumb UI problems which are largely due to a lack of standardization. The only solution requires collective action. In addition, a well-written technical specification outsources a lot of difficult or highly specific questions to a committee of experts (kind of like how the C specification is an excellent technical manual, or K&R was a good de facto specification).

Gmail and Outlook both have market lock-in on personal / business email because of how their email clients integrate with other personal / business software. (Gmail is also given a hand by rational consumer apathy; Gmail is fine and free, changing email addresses is a pain.) I don’t think either company would gain or lose any competitive advantage by standardizing things around email itself. But it would probably reduce a lot of technical management headaches.

nicklecompte··on Toon3D: Seeing cartoons from a new perspective
That's a good point. What's funny is that "The Great Mouse Detective" was actually the film I was thinking of this whole time - I believe the ending sequence took place in Big Ben, and it looks quite good by 2024 standards. But I forgot the name of the movie and assumed it was "Oliver & Company" because Oliver is a plausible name for an English mouse :)
nicklecompte··on Toon3D: Seeing cartoons from a new perspective
But Disney financed and distributed Tron. It wasn't made by a Disney Studio, and most of the animation was outsourced to a Taiwanese studio because Disney wouldn't lend any of their own talent. So I think it's fair to say that Oliver & Company is the first Disney-made film to use CGI.
nicklecompte··on Toon3D: Seeing cartoons from a new perspective
I am amazed they didn't seem to talk to any 3D animators before writing this. Because this is just plain wrong:

> The hand-drawn images are usually faithful representations of the world, but only in a qualitative sense, since it is difficult for humans to draw multiple perspectives of an object or scene 3D consistently. Nevertheless, people can easily perceive 3D scenes from inconsistent inputs!

It is difficult for human artists to maintain perfect geometrical consistency. But that is NOT why 2D animation of 3D scenes is geometrically inconsistent! The reason is that artists stylize 3D scenes to emphasize things for specific artistic reasons. This is especially true for something surreal like SpongeBob. But even King of the Hill has stylized "living room perspectives," "kitchen perspectives," etc. The artists are trying to make things look good, not realistic. And they aren't trying to make humans reconstruct a perfect 3D image - they are trying to evoke our 3D imaginations. It's a very different thing.

Pixar and other high-quality 3D animation studios intentionally distort the real geometry of their scenes for cinematic effect: a small child viewed from an adult's perspective might be rendered with a freakishly long neck and stubby little torso, because the animators are intentionally exaggerating visual foreshortening to emphasize the emotional effect of a wee little child. A realistic perspective would be simply boring. These techniques are all over the place in Pixar movies - it's why their films look so good compared to cheaper studios, who really are just moving a virtual camera around a Euclidean 3D space.

I don't want to comment on the technical details. But it really seems like the authors missed the artistic mark.

nicklecompte··on PHYS771 Lecture 17: Fun with the Anthropic Principle (2006)
I thought the writeup was convincing and addresses the heart of the paradox. The only nitpick: you can have uniform probability distributions on infinite sets, like [0,1]: https://en.wikipedia.org/wiki/Continuous_uniform_distributio... There p(x) = 0 for any x, but for fixed e, p(x +/- e) is the same for all x.

But you can't have such a distribution on an unbounded set, which is where the paradox fails. If we had a uniform distribution on an unbounded set, p(x +/- e) has to be the same for all x and therefore nonzero, but

  p(1 +/- e) + p(2 +/-e) + ...
has to sum to <= 1. It is an infinite sum of nonzero terms so this is a contradiction. (The same argument works if you drop the epsilon for thinking of a distribution on the integers).

I think your writeup was basically clear on this in terms of the math, just some of the language was a bit confused.

nicklecompte··on Elicit – AI Research Assistant
> I'm not sure I agree that those rule-of-thumb statistics are "arbitrary" or "fictional"… I guess it depends on what you mean by that.

Sorry for the confusion: I meant that fragmede's comment was arbitrary and fictional, not the 90% figure. I was talking about these numbers:

  if it takes 1 hour to get one answer by hand, but only 20 minutes for the machine, and 20 minutes to check the answer, the user still comes out ahead
← PreviousPage 2 of 16Next →