HNHacker News
TopNewBestAskShowJobs

enum

509 karma · joined April 7, 2009

https://github.com/arjunguha
submissionscomments
enum··on PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
Heuristic search, not exhaustive search, is an essential ingredient of reasoning. Has been true since chess. Remains true with MCTS, LLMs, etc.
enum··on PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
The nature of the problems makes it relatively easy to follow along with the models' reasoning and reasoning errors. For example, on this problem (answer "New England"):

> Think of a place in America. Two words, 10 letters altogether. The first five letters read the same forward and backward. The last five letters spell something found in the body. What place is this?

R1 fixates on answers of the form "CITY, STATE" and eventually returns some confidently wrong nonsense. It doesn't try to explore answers that don't fit the "CITY, STATE" template.

enum··on PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
Some still exist, but this is less and less true over time. Even when a requirement is called an "exam", it isn't what a high-schooler would recognize as one.

Many of the exams described on this page are personalized and only require mastery of a handful of papers that really matter to the individual student, and knowledge of related work:

https://www.cs.stanford.edu/phd-program-requirements-qualify...

I've been through something that sounds similar to these, and I would not have called it an exam.

enum··on PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
There is a trend of building hard benchmarks for LLMs, such as GPQA or Humanity's Last Exam, with questions crafted by PhDs. Ergo, most people cannot understand what these benchmarks are testing. If you're an expert in something, and if you're on HN you may be an expert in some sub-field of CS, I recommend trying to solve the questions from your field on Humanity's Last Exam. I took a crack at the Programming Languages questions myself.
enum··on PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
I think what you mean is that once you've managed to recall, checking constraints is easy. Remarkably, a few people are much better at this than others. They are able to think fast and execute an explicit mental search over a very small number of plausible candidates. Other people take forever. Seems to be the case for models too.
enum··on PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
It's definitely U.S. centric.

But, I do think this is reasoning. It requires recall, but anything other than pure logic puzzles do. For example, on a competition math problem or a programming problem, No person or LLM is inventing well-known lemmas and algorithms from first-principles.

enum··on PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
Definitely a lot of letter counting. It's not not a factor. I think the real problem is that the search space for each problem is enormous. When it gets stuck, it just gets stuck enumerating candidates that meet some but not all of the constraints.
enum··on PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
There is likely a little bit of that, but I don't think that's the primary factor that makes these hard. If you look at the reasoning performed by R1 ("Show Thoughts" checkbox), it has no problem splitting and joining candidate words into letters, or even doing the letter-level manipulations. Just to give an example:

Take a familiar brand name, seen along roads and highways in the United States. It has five letters, two syllables. The first syllable, phonetically, is a word that is the opposite of the word spelled by the second syllable. What brand name is it?

<think> Okay, let me try to figure out this riddle. The question is about a brand name that's familiar on roads and highways in the US. It has five letters and two syllables. The key part is that the first syllable phonetically is the opposite of the second syllable spelled out. Hmm.

First, I need to think of common brands seen on highways. Gas stations, fast food, hotels, maybe car brands. Five letters, two syllables. Let's list some possibilities: Exxon, Shell, Chevron, Mobil, BP (but BP is two letters, so probably not). Maybe food places like Subway, but that's six letters. McDonald's is too long. Burger King? No. Maybe hotels: Motel 6? No, that's a number. Holiday Inn? Too long. Maybe Best Western? No, too long.

Wait, five letters. Let's think of two-syllable brands. Exxon (E-xxon), but Exxon is five letters. First syllable is "Ex", second is "xon". But what's the opposite of "xon"? That doesn't make sense. Maybe the first syllable is a word, and the second syllable when spelled is the opposite. So, for example, if the first syllable is "up", the second syllable's letters spell "down". But "updown" isn't a brand. Wait, the first syllable phonetically is a word, and the second syllable's spelling is the opposite. So when you say the first syllable, it's a word, and the letters of the second syllable spell its opposite.

So, for example, if the brand is "Updn" (made up), where "up" is the first syllable, and "dn" spells "down". But that's not a real brand. Let's think of real brands. Maybe "Comcast"? No, too long. "Verizon"? Also longer. "Sprint"? Five letters. Sprint. S-print. But "S" doesn't mean anything opposite of "print". Hmm. Maybe "Exxon" again. "Ex" as in former, and the opposite would be "xon" spelled as... no.

Wait, maybe the first syllable is a word that's the opposite of the word formed by the second syllable's letters. For instance, "Sunoco" is a gas station, but that's six letters. "Shell" is five letters but one syllable. "Hertz" is car rental. Five letters, one syllable. "Avis"? Four letters. "FedEx"? Five letters? F-E-D-E-X? No, that's six. Wait, FedEx is two syllables: Fed-Ex. Five letters? F-E-D-E-X? That's five letters? F, E, D, E, X. Yes, five letters. FedEx. The first syllable is "Fed", the second is "Ex". Now, "Fed" as in past tense of feed, but maybe "Fed" as in federal. The opposite of "Ex" could be "current" or "present". But that doesn't seem right. Alternatively, "Ex" as in former, so the opposite would be "current" or "new". Hmm, not sure. Maybe "FedEx" isn't the answer.

Another thought: "Amtrak" is a brand on highways, but that's six letters. "Greyhound"? Too long. "Dollar" car rental? Five letters? No, six. "Thrifty" is seven. Wait, maybe "Lyft"? No, four letters. "U-Haul" is hyphenated.

Wait, let's think of opposites. The first syllable's sound is a word, and the second syllable's spelling is its opposite. So, for example, if the brand is "Inote", where "I" (pronounced "eye") and "note", but "eye" and "note" don't have opposites. Alternatively, "Hi-Lo", but that's hyphenated. Maybe "Dayton" (a tire company?), but "Day" vs "ton" – not opposites.

Alternatively, maybe "Duracell" – no, too long. "Energizer"? No. "Eveready"? No. Wait, maybe "Goodyear". That's a tire company. Good-year. "Good" and "year". Not opposites.

Wait, the key is that the first syllable phonetically is a word, and the second syllable's spelling is the opposite. So, for instance, if the first syllable is "pro", the second syllable's letters spell "con". But "procon" isn't a brand. Or "up" and "down". Maybe "Updown", but not a real brand.

Wait, the brand "Exxon" – "Ex" and "xon". If "Ex" is the opposite of "xon" spelled as "xon" – no. Maybe "Ex" (as in former) and "xon" – but "xon" isn't a word. Alternatively, "Exxon" could be broken into "Ex" and "on". Wait, "Ex" (as in former) and "on" (the opposite being "off"). But "Exxon" is E-X-X-O-N, so the second syllable is "xon", not "on". Hmm.

Another approach: think of opposites like "yes/no", "hot/cold", "fast/slow", "high/low", "in/out", "on/off", "start/stop", "day/night". Maybe the brand incorporates one of these.

"Hi-Lo" is hyphenated. "Stop & Go" – no. "Speedpass" – not five letters. "7-Eleven" – no.

Wait, "Loews" hotels? Five letters. "Lo-ews". "Lo" as in low, and "ews" – not an opposite. Maybe not.

Wait, "Lucent" – a telecom company, but that's six letters. "Verizon" – seven letters.

Wait, "Motel 6" – no, includes a number. "Super 8" – same.

Wait, "Shell" – five letters but one syllable. "Chevron" – two syllables but seven letters.

Wait, "Pepsi" – five letters, two syllables. "Pep" and "si". "Pep" as in energy, and "si" – not an opposite.

"Cola" – not a brand. "Coke" – one syllable.

Wait, "Adidas" – three syllables. "Nike" – two syllables but four letters.

Wait, "Reebok" – six letters. "Fila" – four.

Wait, maybe a car brand. "Ford" – one syllable. "Chevy" – two syllables but five letters? C-H-E-V-Y. Five letters, "Chev" and "y"? Not opposites.

Wait, "Honda" – five letters, two syllables: Hon-da. Not opposites. "Toyota" – too long. "Lexus" – two syllables but five letters. "Le-xus". Not opposites.

Wait, "Nissan" – two syllables, six letters. "Mazda" – two syllables, five letters. M-A-Z-D-A. "Maz" and "da". Doesn't fit.

Wait, "Dodge" – one syllable. "Jeep" – one.

Wait, "Subaru" – three syllables.

Hmm. Let's think of gas stations again. "Shell", "Exxon", "Mobil", "Texaco", "Chevron", "BP", "Valero", "Citgo".

Citgo: C-I-T-G-O. Two syllables: Cit-go. "Cit" (like "sit") and "go". Wait, "sit" and "go" – those are opposites in a way. If you sit, you're not moving; if you go, you're moving. So "Citgo" – "Cit" phonetically is "sit", and "go" is spelled as "go". So the first syllable sounds like "sit", which is the opposite of "go". That might be it!

Citgo is a gas station brand, seen on highways, five letters (C-I-T-G-O), two syllables: Cit-go. Phonetically, "Cit" sounds like "sit", and "go" is the opposite. So the first syllable (phonetically "sit") is the opposite of the second syllable spelled "go". That seems to fit.

Let me double-check. The brand is Citgo. Five letters, two syllables. First syllable "Cit" pronounced like "sit", which is the opposite of "go". Yes, that makes sense. I think that's the answer. </think>

The brand name is *Citgo*.

Citgo is a well-known gas station brand commonly seen along U.S. roads and highways. It has five letters and two syllables: *Cit* (pronounced like "sit") and *go*. The first syllable phonetically represents "sit," which is the opposite of the second syllable spelled as "go." This clever play on opposites fits the riddle's criteria perfectly.

enum··on PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
Thanks, fixed. (Spaces rebuilding.) We have manually combed labelled-wrong answers and tweaked the predicates that check correctness. Sorry we missed this one.
enum··on PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
They are scraped from the web, and discussed on Reddit. So, they are definitely in the training data. Despite that, the non-reasoning LLMs struggle to solve them.

There are however new problems each week, and released every week. So, we can safely assume the latest problems are decontaminated. It remains to be seen if and how performance drops on the problems released in 2025. (Not enough problems yet to tell.)

enum··on PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
Thanks, and yes :)
enum··on PhD Knowledge Not Required: A Reasoning Challenge for Large Language Models
Results and dataset explorer here: https://huggingface.co/spaces/nuprl/verbal-reasoning-challen...
enum··on Northeastern's redesign of the CS curriculum
A discussion about this from faculty: https://bsky.app/profile/samth.bsky.social/post/3lf22654bks2...
enum··on The Illustrated Guide to a PhD
You'll do great. This will eventually turn into new discoveries if you keep at it.
enum··on The Illustrated Guide to a PhD
In the U.S., in most fields, it is virtually impossible.
enum··on The Illustrated Guide to a PhD
Any evidence that this approach works? Are people who do this able to move from the PhD to a solid position afterward that they could not have had without the PhD?
enum··on Principles of Educational Programming Language Design
Absolutely. Demonstrating how to Google (and now, how to ChatGPT) is important. The pervasiveness of Java makes it relatively easy to do.
enum··on Principles of Educational Programming Language Design
The abstract asks:

> Why do we not have a programming language that is designed for education and in widespread use across the world

It is important for a teacher to immediately demonstrate subject-matter mastery. If a student asks a question that goes beyond the planned lesson, you need to have an answer. You can't say, "I don't know how to do that." That would make you look incompetent.

When you're teaching programming, it is easiest to do this with a programming language that you know well and use everyday. That language is unlikely be a language designed explicitly for education.

enum··on The 70% problem: Hard truths about AI-assisted coding
Academic studies are finding the same thing. Although there are a handful of beginners who are great at prompting, when you study beginning programmers at scale, you find that the mostly struggle to write prompts and understand why things go wrong. Here is one of several example studies:

https://dl.acm.org/doi/full/10.1145/3613904.3642706

enum··on Phind-405B and faster, high quality AI answers for everyone
This is going to be renamed to Llama Phind 405B, right?
enum··on Coroutines make robot code easy
Very cool to see. I had worked on something similar but in the context of JavaScript a few years ago (https://arxiv.org/abs/1909.03110). Without coroutines/continuations, it really would have been impossible to get people up to speed in the time we had (one week).
enum··on StarCoder and StarCoderBase: 15.5B parameter models with 8K context length
I'm curious--what have you fine-tuned it on?
enum··on BigCode Project Releases StarCoder: A 15B Code LLM
There is a lot of subtlety here. The model is trained on 80+ languages, but the volume and quality of data varies significantly. We have results showing benchmark performance on 19 languages, which is a broader evaluation than most Code LLMs.

But, I would hesitate to say that BigCode supports 80 PLs. Any LLM that claims to support 80 PLs is not presenting evidence that it does.

enum··on BigCode Project Releases StarCoder: A 15B Code LLM
You can just about load it on a 32GB GPU in 16bit mode.

Quantized versions here:

https://huggingface.co/mayank31398/starcoder-GPTQ

they will be benchmarked on humaneval and released soon—maybe tomorrow?

enum··on BigCode Project Releases StarCoder: A 15B Code LLM
Keep in mind that StarCoder(Base) is just a pretrained LM. The extra stuff that makes 3.5/4 like RLHF gets built on this.
enum··on BigCode Project Releases StarCoder: A 15B Code LLM
Suggested fine-tuning code is here:

https://github.com/bigcode-project/starcoder

enum··on BigCode Project Releases StarCoder: A 15B Code LLM
You may (but do not have to) use <reponame>, <filename>, etc. as special tokens to prompt the model with extra metadata. These help you use the model go beyond just code completion.

Page 30 of the TR has a few examples:

https://drive.google.com/file/d/1cN-b9GnWtHzQRoE7M7gAEyivY0k...

Note that they were produced with StarCoderBase. E.g., here is one:

Input:

    <commit_before>def fibonacci(n):<commit_msg>add type hints to function<commit_after>def

Output:

    fibonacci(n: int) -> int:
(If you try this in the playground, dial down the repetition penalty to say 1.0 in Advanced Settings.)
enum··on BigCode Project Releases StarCoder: A 15B Code LLM
This was used for PII detection and will be public soon.

It could have other uses, but its not what you want code generation. For code generation, use StarCoder or StarCoderBase.

enum··on BigCode Project Releases StarCoder: A 15B Code LLM
Also note that StarCoder itself is open (but not replit-finetuned).
enum··on BigCode Project Releases StarCoder: A 15B Code LLM
StarCoder is a 15B LLM for code with 8k context and trained only on permissive data in 80+ programming languages.

Release thread: https://twitter.com/BigCodeProject/status/165417494197606811...

Model: https://huggingface.co/bigcode/starcoder

Paper: https://drive.google.com/file/d/1cN-b9GnWtHzQRoE7M7gAEyivY0k...

Also available in HuggingChat: https://huggingface.co/chat

← PreviousPage 2 of 4Next →