Are you an editor or involved with app review? I would rather read this honestly written, succinct blog post than a mountain of AI slop that half of the HN front page is.
Well, I can tell what sort of angle you most enjoy. Anyways -- I think there is GOOD writing and BAD writing, but only subjectively. So if you enjoy it, power to you. It's certainly not random, but it is the sort of verbosity that turns off 99 percent of the people that would read it given a comparison. I find the former rather eloquent.
Another thing that sort of puzzles me about benchmarks is that LLMs are not deterministic and do not always complete a problem. So what are the results actually representing? The best run? The average? It is all in some ways a falsehood
I don't think it is vague in the slightest. Take the most simple examples, how many LLM's have you tested making them? There are stylistic choices pertaining to games that is well beyond a 0/1 reward. Even something as basic as breakout or flappy bird can have wildly different quality between models. Yeah, you could call this animal on a bike benchmarking, but I don't think it is. IMO the problem space occupies an interesting area where you can ignore the pass/fail and focus on the actual level of the model to do something beyond that.
I doubt the OP meant something like creating the whole tech stack for WOW.
I have been in RDP land but I would never go back unless required (windows). While it could just be personal experience/failures, latency was not good and reliability was not good. Things are probably better as of late though I hesitate to assume software has improved... It is just much simpler to live in the terminal for dev work, especially with agents. That being said, I use MacOS and Windows daily for non-dev work. Either way, your points are valid, and whatever works for you works for you friend :)
From what I understand pre-training is totally irrelevant to this and as far as post training goes there will be multiple steps, for claude and codex and the like that ship with a harness, the harness is definitely included in evaluation. However, they will definitely include evaluation from a variety or even none, and settle on something that works the "best" for a release.
Tmux is good because:
1) sessions stay running if you disconnect
2) window management - you can split your screen, have an octobox a-la redzone style, and focus/unfocus etc.
This is based on developing on a remote server - but even locally, I find it invaluable. Multiple terminal windows are fine, but some times you want multiple windows. Even in a pre-ai world, you might want to run a process, see the code, edit, and maybe have htop or something like that. If you ever NEED multiple terminal windows for the same thing, tmux is really the answer.
I just wanted to echo your comment. Most people, in my experience, pursue an undergrad program with the intention of being employable. This is my personal experience, but also a learned opinion from being a teaching assistant for some years. Places like waterloo stress co-op, it is built in to the program. IMO the average person seeking an education are served under a model that places them in or near industry related work at some point. While I understand the original post is relating to PHDs, I think most programs (or rather, participants in programs) would love to have this sort of integration despite it being a hard problem.
I honestly just use GPT models nowadays, Claude models are too restrictive and more of a quitter and fable/whatever is just too expensive to be worth it.
I find it useful for code reviews (spawn a subagent with minimal/no context to review X commit). Of course, this is more or less a shortcut that could be done with a seperate agent. Another use is multiple reviews at once if tokens are not an issue, with seperate "personas" or focuses. As far as implementation goes I have not seen any major usecase.
A degree is not a bad thing. This forum is pretty biased on startup culture but I bet the vast majority would say its not worth it personally but worth it on a career level. Even then, the space to explore things outside of your immediate interest is invaluable and you WILL make connections beyond what you expect. Good luck in your studies.
I echo the sentiment. Most work is described as basic and unimaginative, yet we still have every large company having outages despite employing "the best". Even worse, they game uptime and outages in a way that mirrors gerrymandering.
Most likely, this (reverse engineering) is one of the numerous things these LLM companies target. You can also assume all of the internet has been slurped up in to any frontier model. That doesn't mean what you want will be a one shot prompt though...
> For a significantly shorter critque of the book, check out qntm's critique. I mostly agree with qntm assessment. But it's a bit too emotional and personal and doesn't cover the parts i find the most harmful.
This page seems like an actual critique while the blog post doesn't offer much of one, am I missing something?
edit: There is a github linked towards the bottom of the post... full of LLM emoji exclamations.
It is extremely easy to burn tokens if that is required.
Explore this codebase.
Team x wants y feature, research and generate a full plan.
What does feature x in codebase y actually mean?
Analyze code coverage in x.
Map out code flow and find concurrency bugs in y
and on and on...
Oh and my favorite: Use 5 independent subagents to review code change and summarize the findings, and for any finding determine if they are real concerns
People will reply to you calling you crazy, but SF/bay is the only place I have ever experienced where many people will literally leave their cars unlocked because a broken window isn't worth the hassle. Yes, locking your parked car is a hazard here... and the reason is obvious.
Most of my work has been in core infra at large companies. Having the code written faster does not change rollout velocity all that much... It does help with signals and idiot proofing on bugs but when things break and cost real (very real) dollars AI is not an explanation. In that instance, its not even close. Development might be 10-20 percent of the actual work to get a change out.