How Well Do LLMs Generate Code for Different Application Domains?
arxiv.org
arxiv.org
Have to know where you are to know where to go, and once you know where to go and go there, well you have to know where you are!
Benchmarking, like modelling or simulation, are one of many imperfect but useful tools researchers can reach for when trying to better understand something.
Benchmarks help us contextualize things. All benchmarks have biases, and some tell us more than others, but they certainly fall under a class of experimentation that falls under "research". Of course, it's important not to overstate findings. Performing well on any given coding benchmark does not tell us "this LLM is, broadly speaking, great at coding", but that it is great at coding as measured by a narrow testing set.
There's also another snake pit to cross I'm that validation methodology is hard (data leakage, appropriate measures and stats, etc...) but plenty of quality research over comes that with rigour and patience.
-- ignoring what I would consider "weird" in the code, I assume that's just style.
The prompt is:
This function performs a forward pass for a model, incorporating conditioning and time step information.
It randomly selects time steps, applies learned conditioning if applicable, and processes the inputs according to the model's conditioning requirements.
Finally, it computes and returns the loss for the given inputs and conditioning.
The ground truth (am I correct, this is the expected answer?) is:
def forward(self, x, c, *args, **kwargs):
t = torch.randint(0, self.num_timesteps, (x.shape[0],), device=self.device).long()
logging.info(f'Random timestep t generated with shape: {t.shape}')
if self.model.conditioning_key is not None:
assert c is not None
if self.cond_stage_trainable:
c = self.get_learned_conditioning(c)
if self.shorten_cond_schedule: # TODO: drop this option
tc = self.cond_ids[t].to(self.device)
c = self.q_sample(x_start=c, t=tc, noise=torch.randn_like(c.float()))
return self.p_losses(x, c, t, *args, **kwargs)No, but several of the problems were similar to problems in the training set from other challenges.
There was one that was almost exactly the same as one on Google Foobar called 'Prepare the Bunnies' Escape' as an example.
It is pretty hard to make these challenges that work across multiple langs, skills etc.. without making them into 'which works dfs/bfs/a*/dykstra' problems.
Obviously outside of the leaderboard group, an LLM being able to solve it doesn't really matter.
My friends and I flipped a coin for TDD vs traditional with each day, testing out assumptions and evaluating what was most readable/maintainable.
It was great for that.
This looks super interesting and useful, but without Claude or o1 data points, it isn’t clear if it is saturated or not, and without more modern open source models, it doesn’t give me a useful signal to choose a model.
A benchmark is more like a program that is run and whose output is evaluated. Those are coming later, but they're not like this. Arguably to make it "scientifically valuable" rather than anecdotally valuable, this sort of work has to come first to figure out what and/or how to benchmark in the first place.
Once you have that working, you can generate synthetic examples to train the next version of the LLM, or for open source coding models, a small LoRA file that anyone can download to add support to their model of choice.
And often times the results are not replicable either.
Great you enjoy more foundational work!
I haven’t read the whole thing so I can’t really judge whether this specific benchmark is useful, but if it is, every time a new model comes out they can run the benchmark and breathlessly report its improved performance.
This is about the benchmark they're introducing, which would have real uses for all subsequent models.