129 karma · joined November 12, 2021
Yes, this dynamic has indeed been screwed up in many teams!
In my perspective, you can’t be a good product manager for a technical product without being somewhat technical yourself. You simply won’t understand your product well enough.
The only dynamic where that works is if the product manager acts more as a supporter and challenger to a team lead by an EM or technical lead, but then product manager is probably not the right title.
Dissatisfaction is kind of expected (my cost goes up 2 orders of magnitude and I already cancelled since there are better options at market rate). Complaining without change won’t matter.
As I understand it, it is designed to evaluate the LM itself and not agentic systems with online access (very high likelihood of unintentional cheating/solution leaking). The paper and docs are not super clear on the concrete requirements (although reproducibility is emphasized which goes against online access). So I was hoping for someone with more familiarity to chip in.
Obviously not a problem for internal evaluations, but for fair scoreboard submissions it matters. It's not a matter of whether internet searches are useful, but rather what the benchmark is intended to benchmark.
I get that you may accidentally include something in local git history, but it feels off to me to run these kinds of benchmarks online.
The additional up-front cost for hardware designed to run an LLM in addition to normal workload is unlikely to be accepted by most consumers.
The scale will be very constrained (like Apples on-device models which are small, heavily quantized, and have a small 4K token context window). It’s also terrible for battery life.
AI as it is implemented today is simply just computationally expensive and unless you put in dedicated hardware (like the ANE) for only this purpose - a large cost driver - I don’t really see it getting large scale adoption.
Companies will probably need a server-backed solution as fallback if they want reasonable user experience, so why even invest in diverse hardware support.
I’ve worked most of my career in US tech satellite offices and I have not experienced EU team members to be less productive than US team members, nor spend less time on work (if anything, more really since they also need to be available for US time zone overlap).
It’s true there are chill jobs here, as there are in the US.
But ambitious people tend to work as much as ambitious US people (and it’s really more like 40 hours work weeks - 39,5 where I live since lunch is not work time). But again, many are not really counting, it’s just a full time job.
Vacations (typically 3 weeks summer holiday and additional weeks to distribute over the year) does create longer time on skeleton crew. Skilled tech labour is also cheaper so you can just hire more to make up for it.
My experience is that the same job also has a large range for meaningfulness depending on how well leadership manages to facilitate it.
The thread comes off as a pitch for a startup, hence my skepticism.
“Each setting corresponds to a time-based distance that represents how long it takes for Model 3, from its current location, to reach the location of the rear bumper of the vehicle ahead of you”
I think it might have been in the past, at least I was told the same thing years ago.
That said, it is actually an issue that it now keeps too much distance under some conditions on the lowest setting.