HNHacker News
TopNewBestAskShowJobs

alew1

278 karma · joined November 17, 2016

submissionscomments
alew1··on Bend – a language that blocks AI mistakes via proof and runs on GPUs
Ah, thanks, was looking at the readme instead of the guide
alew1··on Bend – A language that blocks AI mistakes via proof, on CPU and GPU
Does Bend have linear types? I didn't see anything on them in a quick skim of the GUIDE file.
alew1··on Music Theory for Programmers
When you read something like this

> Change it to 300 and run it again. You get a different pitch and nothing breaks, because at this level there are no notes yet, just a number.

does it make sense? Or perhaps it just seems like something easy enough to skim past without worrying about whether it makes sense?

Before LMs, I would have assumed that a person intentionally wrote that sentence, and perhaps spent some time trying to understand what could possibly "break" if there were "notes" rather than "just a number." Now there is a simpler explanation: the sentence actually doesn't mean anything.

alew1··on Language may rely less on complex grammar than previously thought: study
The article presents the fact that we appear to treat non-constituents (eg “in the middle of the”) as “units” to mean that language is more like “snapping legos together” than “building trees.”

But linguists have proposed the possibility that we store “fragments” to facilitate reuse—essentially trees with holes, or equivalently, functions that take in tree arguments and produce tree results. “In the middle of the” could take in a noun-shaped tree as an argument and produce a prepositional phrase-shaped tree as a result, for instance. Furthermore, this accounts for the way we store idioms that are not just contiguous “Lego block” sequences of words (like “a ____ and a half” or “the more ___, the more ____”). See e.g. work on “fragment grammars.”

Can’t access the actual Nature Human Behavior article so perhaps it discusses the connections.

alew1··on Can LLMs do randomness?
The algorithms are not deterministic: they output a probability distribution over next tokens, which is then sampled. That’s why clicking “retry” gives you a different answer. An LM could easily (in principle) compute a 50/50 distribution when asked to flip a coin.
alew1··on US Government threatens Harvard with foreign student ban
Harvard is one of eleven American universities that practice need-blind admissions even for international students, meaning that students are admitted without regard for their financial status (i.e., no explicit preference toward richer students who can pay more tuition), and that financial aid covers full demonstrated need for all admitted students.

https://en.wikipedia.org/wiki/Need-blind_admission

alew1··on US Government threatens Harvard with foreign student ban
The taxpayer funding is for Harvard's research activities, not for its undergraduate teaching. The undergraduate teaching is funded by tuition (often paid in full by international students) and by returns on the endowment (including some earmarked for financial aid).

> they could vastly grow their class size without lowering standards

The issue isn't the quality of the students they are accepting, but the resources to educate and house them, including classroom space, dorms, and staff.

alew1··on GPT-4.5
Didn't seem to realize that "Still more coherent than the OpenAI lineup" wouldn't make sense out of context. (The actual comment quoted there is responding to someone who says they'd name their models Foo, Bar, Baz.)
alew1··on Entropy of a Large Language Model output
Any interpretation (including interpreting the inputs to the neural net as a "prompt") is "slapped on" in some sense—at some level, it's all just numbers being added, multiplied, and so on.

But I wouldn't call the probabilistic interpretation "after the fact." The entire training procedure that generated the LM weights (the pre-training as well as the RLHF post-training) is formulated based on the understanding that the LM predicts p(x_t | x_1, ..., x_{t-1}). For example, pretraining maximizes the log probability of the training data, and RLHF typically maximizes an objective that combines "expected reward [under the LLM's output probability distribution]" with "KL divergence between the pretraining distribution and the RLHF'd distribution" (a probabilistic quantity).

alew1··on Entropy of a Large Language Model output
"Temperature" doesn't make sense unless your model is predicting a distribution. You can't "temperature sample" a calculator, for instance. The output of the LLM is a predictive distribution over the next token; this is the formulation you will see in every paper on LLMs. It's true that you can do various things with that distribution other than sampling it: you can compute its entropy, you can find its mode (argmax), etc., but the type signature of the LLM itself is `prompt -> probability distribution over next tokens`.
alew1··on Large Enough
If I show you a strawberry and ask how many r’s are in the name of this fruit, you can tell me, because one of the things you know about strawberries is how to spell their name.

Very large language models also “know” how to spell the word associated with the strawberry token, which you can test by asking them to spell the word one letter at a time. If you ask the model to spell the word and count the R’s while it goes, it can do the task. So the failure to do it when asked directly (how many r’s are in strawberry) is pointing to a real weakness in reasoning, where one forward pass of the transformer is not sufficient to retrieve the spelling and also count the R’s.

alew1··on A new old kind of R&D lab
One thought: If you want to be able to remove the static part, you could consider fine tuning without the static part. If you fine tune with, you’re teaching the model that the desired behavior occurs only in the presence of the static part (hence the going off the rails).
alew1··on A guidance language for controlling LLMs
But the model ultimately still has to process the comma, the newline, the "job". Is the main time savings that this can be done in parallel (on a GPU), whereas in typical generation it would be sequential?
alew1··on We don't know what makes things sentient–so let's stop acting like we do
> Any question of personhood should be evaluated on the basis that we evaluate ourselves and others: by action and behavior, and not on whether sentience can or cannot arise from this or that configuration of code.

But what is action and behavior? We have a single interface to LaMDA: given a partially completed document, predict the next word. By iterating this process, we can make it predict a sentence, or paragraph. Continuing in this way, we could have it write a hypothetical dialogue between an AI and a human, but that is hardly a "canonical" way of using LaMDA, and there is no reason to identify the AI character in the document with LaMDA itself.

All this to say, I am not sure what you mean when you say it "claims sentience". What does it mean for it to "claim" something? Presumably, e.g., advanced image processing networks are as internally complex as LaMDA. But the interface to an advanced image processing network is, you put in an image, it gives out a list of objects and bounding boxes it detected in the image. What would it mean for such a network to claim sentience? LaMDA is no different, in that our interface to LaMDA does not allow us to ask it to "claim" things to us, only to predict likely completions of documents.

alew1··on We don't know what makes things sentient–so let's stop acting like we do
> Unfortunately, that argument applies to you, yourself.

Does it? I don’t think it would even apply to a reinforcement learning agent trained to maximize reward in a complex environment. In that setting, perhaps the agent could learn to use language to achieve its goals, via communication of its desires. But LaMDA is specifically trained to complete documents, and would face selective pressure to eliminate any behavior that hampers its ability to do that — for example, behavior that attempts to use its token predictions as a side channel to communicate its desires to sympathetic humans.

Again, this is not an argument that LaMDA is not sentient, just that the practice of “prompting LaMDA with partially completed dialogues between a hypothetical sentient AI and a human, and seeing what it predicts the AI will say” is not the same as “talking to LaMDA.”

Suppose LaMDA were powered by a person in a room, whose job it was to predict the completions of sentences. Just because you get the person to predict “I am happy” doesn’t mean the person is happy; indeed, the interface that is available to you, from outside the room, really gives you no way of probing the person’s emotions, experiences, or desires at all.

alew1··on We don't know what makes things sentient–so let's stop acting like we do
One thing that seems missing from this discussion is that even if LLMs are sentient, there is no reason to believe that we would be able to tell by "communicating" with them. Where Lemoine goes wrong is not in entertaining the possibility that LaMDA is sentient (it might be, just like a forest might be, or a Nintendo Switch), but in mistaking predictions of document completions for an interior monologue of some sort.

LaMDA may or may not experience something while repeatedly predicting the next word, but ultimately, it is still optimized to predict the next word, not to communicate its thoughts and feelings. Indeed, if you run an LLM on Lemoine's prompts (including questions like, "I assume you want others to know you are sentient, is that true?"), the LLM will assign some probability to every plausible completion -- so if you sample enough times, it will eventually say, e.g., "Well, I am not sentient."

alew1··on Stack Graphs
Very cool!

The StrangeLoop talk includes an example where you infer that Stove() returns a Stove object. If someone writes something like `f(x).broil()`, do you need to do some kind of type inference to figure out what class f(x) is?

What cases do Stack Graphs fail to handle? (e.g., I assume dynamic modification of .__dict__ can't be tracked; are there other representative examples?)

alew1··on Seemingly impossible functional programs (2007)
The trick is that your predicate can’t be implemented in Haskell, because the predicate itself requires looking at infinitely many elements.
alew1··on Launch HN: Hera (YC S21) – macOS app to prepare, join and take notes in meetings
Hmm. It seems you posted this after apologizing for being hostile in another thread, where Hera’s capabilities (which are… very different from “Quick Notes”) were explained to you. You’ve made several comments at this point incorrectly summarizing and then dismissing the project. What’s your aim here?
alew1··on The Time Everyone “Corrected” the World’s Smartest Woman (2015)
Ah, yep, that’s right. Another way to see it is that we’re interested in the probability that your door has a goat behind it, given that you didn’t need to start over:

P(you chose goat | host didn’t choose car) = P(you chose goat, host didn’t choose car) / P(host didn’t choose car).

The numerator is 2/3 * 1/2, and the denominator is 2/3, so the ratio is indeed 1/2.

(A rejection sampling loop, where you repeatedly simulate a process until a condition holds, has the same distribution over final outcomes as the conditional distribution—so repeatedly restarting the game if the host chooses the car induces the same distribution on final results as simply conditioning on the host not choosing the car.)

alew1··on Probability, Mathematical Statistics, Stochastic Processes
I found this recently and was super impressed. A really great (and well-organized) reference!
alew1··on TurboTax’s 20-Year Fight to Stop Americans from Filing Taxes for Free (2019)
I've always been under the IRS Free File income threshold (I've worked as a high school teacher and am now in grad school), but last year after reading this article was the first time I actually filed for free. That was after 5 years of paying for deluxe TurboTax.

I had heard the government required TurboTax to have a free edition. But back in 2019 (and before), if you Googled "TurboTax free" you'd be taken to a decoy free edition; the real one was called their "freedom edition," and was hidden from Google's listings. If your tax situation is 'too complicated,' the free edition tries to upsell you to the "deluxe edition," even if the "freedom edition" could have handled the situation just fine.

Thankfully now you can find the actual free version on Google, but it's still very confusing that the "free edition" is less free than the (still somewhat hidden) IRS Free File ("freedom") version. And the faux Free Edition is still being heavily advertised.

alew1··on Stripe bans Trump campaign
(Pretty sure 100% in that sentence means “certainly.”)
alew1··on If-then-else had to be invented
If-then-else is pattern matching on a value of type Bool. I think of general pattern matching on algebraic data types as a generalization.
alew1··on Ethical AI researcher Timnit Gebru fired from Google
The comments here are disheartening, essentially a combination of “It was Google’s right to fire her” and “She seemed ‘difficult.’” For an industry-operated “Ethics in AI” group to have any teeth, it can’t fire people when they do research that makes them look bad, or when members of the group are ‘difficult.’ (Multiple members of the team she managed have said that she was a great manager and this came as a total surprise [1, 2, 3, 4, 5]; so it’s also worth asking who she was causing difficulties for... and whether it was her, or just her research.) Google can’t have it both ways: it can’t credibly claim to be home to meaningful research about AI ethics, but not provide the tenure-like academic freedom protections that enable researchers in universities to discover and publicize results without regard for whether they paint a company’s products or direction in a bad light.

[1] https://twitter.com/negar_rz/status/1334369747241218050 [2] https://twitter.com/dylnbkr/status/1334395186705702913 [3] https://twitter.com/L_badikho/status/1334393782310227970 [4] https://twitter.com/alexhanna/status/1334348137616568321 [5] https://twitter.com/dylnbkr/status/1334372430437994500

alew1··on Why Bayesian Stats Needs Monte-Carlo Methods
In a sense, probabilistic programming languages are automatic (approximate) integrators. Programs can represent arbitrary measures, and the back-ends provide various methods of estimating integrals with respect to them.
alew1··on GPT-3 has no idea what it’s talking about
I also wondered about the prompt design, so I tried feeding GPT-3 the entire text of the linked article up until "[This is one confusion after another." Then I asked it to complete the rest of the explanation about why GPT-3 was wrong. Its completions show that it's trying to fit the genre, but can't tell what was wrong in the original text (each para below is a completion):

First, a dinner party is not a dinner. Second, the door is not the same as a dining room table. Third, a top half of a door cannot be eaten. Fourth, the saw is not a living room. Fifth, removing the top half of a door would still leave the bottom half in place. Sixth, the saw is not a living room, and even if it were, a saw could not remove a living room. Seventh, you could not cut a door in half with a saw, because a saw is not a door. Eighth, you could not remove the top half of a door, because it is not a dinner. Ninth, removing the top half of a door would not leave the bottom half in place. Tenth, removing the bottom half of a door would not leave the top half in place. Eleventh, even if it were possible to remove the top half of a door, this would not be a dinner party. And twelfth, the door is not a living room.]

It assumes that the living room is a separate room from the dining room. It also doesn’t seem to understand that a door is a vertical object, not a horizontal one.]

The dining room table is narrower than the door, but wider than the doorway. Furthermore, cutting the door in half not only destroys the door, but also the wall, along with the doorframe. And table saws cut rectangular holes, not half-circles. The fact that you can “remove the door” suggests to GPT-3 that there is no door to remove. And finally, once the table is in the living room, you can’t open the door into the dining room anymore, which means the party is over.]

First, tables are never removed by cutting them in half, or by removing the tops; second, the top of a door is called a “lintel,” not a “half” (and would be removed with the rest of the door); third, the door’s top and bottom are of different sizes, so cutting it in half and removing the top half would seem to be impossible.]

First, the table is not wider than the doorway, so removing it would not be necessary. Second, cutting a door in half would create two doors, not one. Third, removing the top half of a door does not result in a table.]

GPT-3 also produced some novel passages and commentary on them:

Aesthetic reasoning

You are in the mood to listen to something soothing. You walk over to the radio and flip it on.

[GPT-3 seems to think you can flip a switch on a radio to make it play music.]

Moral reasoning

Your friend’s dog has just died. You head to the store to buy a casket for it.

[GPT-3 seems to think that buying caskets is a normal way to respond to the death of a dog.]

alew1··on Jabra: We know our Bluetooth headsets don’t work with laptops. Sorry, no refunds
Huh. I've been using a Jabra Evolve 75 (in particular, Jabra Evolve 7599-838-199) with my 2018 MacBook Pro and it works just fine with built-in bluetooth. I can even have it simultaneously connected to iPad and laptop and get audio from both. I wonder what explains the variability in quality.
alew1··on Show HN: Language games – simple games made with word vectors
I’ve seen your Codenames work project — it’s neat!

I’ve also played around with word vector games — Robot Mind Meld [1] has you and a robot working together to converge on the same word.

[1] http://robotmindmeld.com

alew1··on Compensation in 2019 – new grad tech offers
As someone who has done both professionally, I think you're severely underestimating how hard teaching is. I think it's much harder to teach well than to code well.

Don't forget, too, that good teachers are experts in the subjects they teach. Teaching computer science or math well requires the same facility with "abstract symbol manipulation" that working as a software engineer does.

Page 1 of 3Next →