At this point, we think using AI and being able to use AI effectively is a skill in and of itself. When you're hired, you'll have access to AI. You'd be expected to be able to use said AI effectively.
So, we still give you a FizzBuzz. You can use AI. Even if we told you not to use AI, we know almost everyone would use AI. But you have to understand the FizzBuzz and be able to explain it to us and make changes to it "live". The amount of people that get weeded out just by having to explain the code they "coded themselves" is staggering (even pre-AI, even on a take home where you had no "OMG I suck at live coding" pressure).
You can likely control for that, if you either interview in person or via screen sharing. (Yes, it could be faked, but that's harder.)
The amount of people that can't even navigate "their own" code is astonishing. Never mind explaining what it does or making changes.
[0] The most reliable strategy I've found for that is choosing questions where the wrong answer is the right answer for some much more common question. Actually spending a few seconds and solving the problem easily lets a human pass, but an LLM with insufficient weights or training data (all of them) doesn't stand a chance.
I’ve mostly given up on all of the standard techniques for interviewing sadly, just because “using ai” makes a lot of them trivial, and have resorted to the good old fashioned interview, where I screen for drive, values and root cause seeking, and let people learn tech/frameworks/etc themselves.
But I was wondering, isn’t a take home question still good, if you give a more open ended and ambitious task, and let people vibe code the solution, review the result but ask for the prompt/session as well?
People will be doing that during normal work anyway, so why not test that directly?
One such question (obviously tailored to the role I'm hiring for) is asking whether SoA or AoS inputs will yield a faster dot-product implementation and whether the answer changes for small vs large inputs, also asking why that would be the case.
I typically offer a test with a small number of such questions since each one individually is noisy, but overall the take-home has good signal.
> why not test that directly?
The big thing is that you don't have enough time to probe everything about a candidate, especially if you're being respectful of their time and not burning too much of yours. Your goal is to maximize information gain with respect to the things you care about while minimizing any negative feelings the candidate has about your company.
I could be wrong, but vibe coding feels like another skill which is more efficient to probe indirectly. In your example, I would care about the prompt/session, mostly wouldn't care about the resulting code, and still don't think I would have enough information to judge whether they were any good. There are things I would want to test beyond the vibe coding itself.
In particular, one thing I think is important is being able to reason about code and deeply understand the tradeoffs being made. Even if vibe coding is your job and you're usually able to go straight from Claude to prod, it's detrimental (for the roles I'm looking at) to not be able to easily spot memory leaks, counter-productive OO abstractions, a lack of productive OO abstractions, a host of concurrency issues LLMs are kind of just bad at right now, and so on. My opinion is that the understanding needed to use LLMs effectively (for the code I work on) is much more expensive to develop than any prompt engineering, so I'd rather test those other things directly.
I get tons of spam that could be generated by even a basic LLM based on public information about me, but for positions that are not a reasonable fit.
Apparently, it is common for such cold calls to come from “recruiters” that are not affiliated with the hiring firm, but are trying to collect some sort of referral bounty.
I have no idea why an HR department would be dumb enough to set up such a pipeline (by actually paying for the third party “service”), but I guess once they have the program in place, they also need an LLM to screen spam applications.
"We saw your profile on github and thought you might be a suitable candidate for our open position at $CRYPTO startup.
PS you must be a US-citizen, and the job is 100% on-site"
Those things seem to be blasted out with no regard for my location - I'm not looking for a developer job anyway - but certainly not one in another country.
Spamming github users seems to be the latest growth hack, and it drives me nuts. I made all my repositories archived when I started getting hit with AI-PRs to review, but I'm reaching a point where I think my life would be easier if I just closed the account.
This happened before "AI" too. When all it takes is clicking an "apply now" button on LinkedIn some desperate people will spam any job they see.
Magicly, the spamming stopped, and we only had applications from good genuine candidates with a real interest in the role.
The job of any technology (like email and "apply now" buttons) is to make life easier and better. If it doesn't do that, then don't use it!
I recall seeing one where you had to send a specific payload to an https endpoint to apply (or it might have been an automated screen immediately after the application was submitted). Forcing potential candidates to briefly open the curl manpage seemed like a similarly elegant solution to me. I doubt it works as well in the era of LLMs though.
If I want to hire a driver, I can train someone who does not know how to drive, or hire someone who has experience as a driver. I can do either, but I'd prefer to do the latter in most cases.
I'm dealing with this all the time in recruitment. It can be done. People lie all the time or don't read the requirements. You need a way for the ones who really do know how to do the thing you need to demonstrate it to you.
Have you ever been in a situation where you had to hire someone to help you with something? Would you really follow your own advice? Your advice does not make sense for carpenters, cooks, or drivers. Why should it make sense for programmers?
Years ago, I hired people at closer to entry level. If they had experience that was a bonus but if they didn't we trained them. If they didn't respond to the training, they were let go after a probationary period.
So you waste the weekend on this project when you had no chance from the beginning. And the time restrictions they list mean nothing since if you actually stop after x hours, they will just pick the person who spent the whole weekend and did a more complete job.
I've done quite a few interviews and as long as the interviewee maybe said something like "it would be better to use a shadow DOM" and could explain what a shadow DOM is, I would be pretty happy with that
Expecting someone to build a full shadow DOM as part of their interview take home is excessive
The worst is when they basically ask how you'd build their product. Some people can't handle a different answer, even as they're busy hiring you to improve things.
It's not really bad to ask someone to do a design session with them and "build their product with them from scratch" isn't inherently bad. That's actually pretty neat if you ask me.
What's bad is if there's only a single answer and that's whatever they actually built themselves, which might be a pile of thrown together startup poo that was never cleaned up. But you have the same problem with all sorts of "needless trivia" type questions.
And then do you really want to work at a company, where you can't have a proper "pros and cons of different approaches" type of discussion? If you got hired, you'd have those kinds of discussions with them on an ongoing basis. Bad on the company for letting that person do the hiring but they got what they deserved so to speak.
Just to make an analogy:
If they simply ding you for using 4 spaces coz they use 8, that's bad.
If they ask you why you use 4 spaces, they use 8, give them pros and cons and are there any other approaches and what are the pros and cons of those? That's a good interview so to speak. As an interviewer I would give bonus points if the candidate says something like "I used 4 spaces because I thought that's what you guys were probably using coz everyone's moved away from 8 spaces but secretly I love usings tabs and setting tabwidth to what I want but in reality it really really doesn't matter as long as it's consistent across the codebase as humans can get used to almost everything and this one isn't worth fighting over. Linters and formatters exist for a reason".
Who still uses 8? Isn't that like a COBOL thing?
https://www.kernel.org/doc/html/v4.10/process/coding-style.h...
https://en.wikipedia.org/wiki/Zero-width_space
Btw, at an old job, some joker developer added or copied 1, and broke the whole testbed. It was quite funny. I came over to the sourcecode hosted in Gitlab, ran my regexes that look for naughty characters. Found it after it ate the devs for half a day.
Not because you use 2 spaces. You can argue 2 spaces and the pros and cons and how horizontal scrolling is an issue. One question back would be for example if that means you have huge run-on files where a single function does everything and that's why you need like 17 levels of indentation and that's why only using 2 spaces for each becomes important to you. And then you'd need to argue how that's better for visibility and what might actually be worse about it. If you can do all that, you're hired (if the rest of the interview goes well :P )
Who still uses 8? Isn't that like a COBOL thing?
That works as a flippant comment when we're joking about code indentation after working together for a while and we get along great. As the one and only answer in an interview, you're out. That's quite disrespectful and no it's not a COBOL thing, I've seen (and used) 8 spaces and argued for tabs or 4 much later than COBOL days. In fact I've never written a single line of COBOL.Is this one of the tests where I just need to throw together a five minute quickie to get over your “can you program” filter? or do you need me to put together something flashy and memorable to show off my ceiling? If o put together my flashy thing, would I get dinged for over-engineering something where a five minute hack solution was good enough?
Then they email me back and said the other candidate did the whole thing and they aren't sure if I know how to style a page now because I only completed the backend part.
- It was designed to be fast to complete (20min max -- not a huge imposition if being hired is likely, obviously very expensive if you're taking one for every job posting).
- I only gave them out after a resume screen. If you had a 0% chance then I didn't waste your time. If you had enough other proof of abilities then I skipped the take-home.
- Candidates were told that it was designed to be fast and that if they couldn't complete it quickly they were unlikely to be successful interviewing either. They still had the option to spend a lot of time if they thought my assessment of the situation was wrong, but part of the point was to allow candidates to gauge their own abilities and not waste their time interviewing without a chance of being hired.
- I did a lot of work behind the scenes calibrating and re-writing the questions individually and as a whole so that the test score correlated very well with interview performance (most interviews administered by not-me, removing a form of bias that's easy to creep in there).
If you want to make it more of a fair consideration of time, consider moving your take home to interviews, that way there isn't a time cost asymmetry. You can enforce your "20 min max" claim this way, you can judge a candidate's performance, thought process and filter out anyone who is LLMing or spending inordinate amounts of time on them.
You will also make a better impression on candidates by investing your time in them in the same way they are with you. Maybe you're hiring kids out of college without experience, but you only have to do so many take home tests before you realize that they're a waste of time, and pass on potential employers who throw them at you, or you learn to just send them your hourly rate for the test.
There is usually a huge disconnect between someone who knows that “this task should take 20mins” and doing it cold in a super high-pressure environment.
People sweat, panic, brain freeze, and are just plain out stressed.
I’ll only OK something like this if we give out a similar but not the same task before the interview so a person can train a bit beforehand.
I’ve heard it all justified as “we want to see how you perform under pressure” but to me that has always sounded super flimsy - like if this is representative of how work is done at this organisation, then do I want to work there in the first place? And if it isn’t, why the hell are you putting people through this ringer in the first place, just sounds inhumane.
If you give unlimited amount of time, you're giving an advantage to people with no life who can just focus on your assignment and polish it as if it were a full time job.
If you give a limited amount of time, then you're making the interview a pressure cooker with a countdown clock, giving a disadvantage to people who are just not great at working under minute-to-minute time pressure.
The ones we use have a clear scoring system and prepared inputs - all it matters is the generated output.
I started refusing take-home tests a couple of decades ago, but when I did them, this is 100% what I would have done.
In my experience this is the wrong game theory. Unemployed people can make job hunting their full time job, so a 20 minute take home doesn't select for "who delivers the highest quality solution in the least amount of time," it selects for "who is the richest applicant who can burn hours on a take home to deliver a higher quality result than people with less time they can afford to spend?"
Also, nobody should ever self-select themselves out of an interview process. Passing a resume review and getting a callback is about 10% likely: for every job hunt, in my experience , candidates get about 10 callbacks for every 100 resume sends. From there, it's about 20% chance to get to final stage, and from there, maybe 50% to get an offer (you're either their first choice or second; if second, your hiring hinges on whether the first choice accepts). Math is right there: once you pass a resume check, in terms of the volume of applications you've sent, it's optimal to spend far more effort into this gig than into firing off ten or twenty more resumes.
Therefore, even if the candidate doesn't think they're a good fit, they should do everything they can to stay in the game, including lying by omission.
After all they might be engaging in imposter syndrome, right? Why assume for the interviewer that your python skills aren't good enough - maybe the interviewer understands perfectly well that you've only used it for scripts and one off tools, but doesn't care because they personally believe your startup experience is more valuable to them and they believe you can up skill! Maybe the take home was designed poorly by someone who was tasked randomly by a lead to shit out a take home, and it's not an accurate indication of what the job would be like. Maybe they sent you the wrong take home? Maybe it's a good take home but you need money so fuck it, if you manage to sneak in despite not being a good fit, you can just bust ass to upskill and make up the difference before anyone notices. Or fuck it twice, it's a shit market and who knows how much longer you'll be able to sell your labor as an engineer, even if you can only fool them for two weeks, that's two weeks of income while you still keep up your job hunt.
GP specifically stated that this was the point of the takehome though. If the person handing it out specifically warns you that struggling with it means you aren't a good fit then if you struggle with it that's not imposter syndrome - you aren't a good fit! Not dropping out at that point is just refusing to acknowledge reality and insisting on wasting everyone's time.
Sure, but, there's nothing incentivizing the candidate to not waste time. At the very least, they can get free interview practice.
I get that that's very annoying when you're on the hiring side, but it's not all so bad in the end, you're getting paid for your time!
People who have really good jobs aren’t applying for jobs below them, so your applicant pool will always be people who are in an equal or worse position than your job.
No one at, Anthropic, for example, is applying to a job at Geico.
Also I know plenty of people in startup world who are phenomenal engineers that only have companies I've never heard of on their resumes - startups that for one reason or another simply didn't have a news-grabbing exit.
But they’re not. And they won’t. And that is my point. They’d make a ML Engineer post on LinkedIn and get a bunch of people for whom Geico would be a step up.
There will never be a job opening from Geico that someone at Anthropic would apply to.
That’s my point - your pool will always be people who are in a worse position than your job. Being laid off is a worse position than a job.
You’ll never see Anthropic candidates in a Geico hiring pool, unless they were laid off for being lousy and can’t find anything else.
The market is pretty efficient - people wouldn’t bid for jobs that are worse than their current situation.
This still seems like an oversimplification. It's easy to label FAANG, "frontier AI companies," whatever else, but the vast majority of jobs and the vast majority of engineers are in a soup that's maybe able to be split between "startup world" and "enterprise world" but beyond that, difficult to say one is "worse" or "better." And I've worked alongside FAANG people in startup world so, either that isn't a "worse" job and therefore your theory doesn't work because that means it's not really possible to accurately evaluate every single company as objectively worse/better, or, your theory doesn't work because people do apply to "worse" jobs.
Take home tests were never a worthwhile signal. Pre-AI, people would search for solutions or have another person complete it.
Employers can, with a somewhat high (?) degree of certainty, weed out candidates using AI.
The AI point is worth diving into a little. This was a year ago, so SOTA was worse, but I didn't find it terribly hard to write questions AI couldn't solve, whose answers you couldn't search for, and which good candidates could solve. The test was a few of those questions and a few which were easier to cheat, and almost nobody had good scores on just the cheatable section.
I don't think that moat will exist indefinitely, but today's AI just isn't very good at a lot of incredibly basic tasks unless the operator has enough outside knowledge to guide it in the right direction (and if a candidate did that I mostly wouldn't care because, by definition, they had the knowledge I was looking for). I use AI a lot, it's great at a lot of things, some even quite complicated, but it was weaknesses, and those are pretty easy to exploit.
> The test was a few of those questions and a few which were easier to cheat, and almost nobody had good scores on just the cheatable section
I also like how you allow/encourage self-assessment, where if a candidate can't do the test in ~20 minutes under zero pressure, they probably won't be a good fit in the role itself.
If you’re some random startup or no name company, I don’t bother. I already have a good job.
If you’re a top name hot company offering $600k+ total comp, I’m going to spend the hour shooting my shot. Even if it’s lower likelihood.
For random company XYZ, I expect humans to sink as much of their time into the process as I do.
I've worked at places that would send out 100. People would spend their weekends working on it and we often wouldn't even look at the submissions.
You're right the time commitment wasn't equal. Early on I spent much more time than the candidates designing and analyzing the test. Afterward, their 20 minutes would usually take me <5min (often <1min for obvious failures and obvious passes, the average brought up due to time analyzing edge cases).
I did read every submission though. It wasn't wasted time for candidates.
Wow, this is a great way of putting it. It's draining enough to go to third- and fourth-round interviews with other humans. Doing it with a series of AI chat bots would be devastating!
I hate it from the candidates' perspective, but it's not illogical from the employer perspective.
No, I don't know how to fix it.
It's quite rare for companies to have evidence to support their hiring methods, which unfortunately means it's heavily driven by trends.
I'm not sure that first sentence true. Let me play Devil's advocate:
What's the primary cause of not being able to find someone who meets your standard when you already get lots of applications? It's that your hiring process is bogged down by the masses of unwanted candidates you must evaluate to find the few wanted candidates in the crowd of applicants. And what's the fix? It's better screening. Which is raising your bar, isn't it? Even if it's only to add cargo-cult screens to your bar, it's making the bar more selective, isn't it? Fewer people clear it, right?
On the other hand if you "raise your bar" (let's say you do so by some method that makes it twice as expensive to judge a candidate; twice as likely to reject a candidate that would fit what you need, i.e. doubles your false negative rate; but cuts down on the number of applications by 10x, so that now 1 out of 100 candidates are what you need, which isn't that far off the mark for certain kinds of things), you cut down the effort (and time) you need to spend on finding a candidate by over double.
EDIT: On reflection I think we're mainly talking past each other. You are thinking of a scenario where all stages take roughly the same amount of effort/time, whereas tmorel and I are thinking of a scenario where different stages take different amounts of effort/time. If you "raise the bar" on the stages that take less amount of effort/time (assuming that those stages still have some amount of selection usefulness) then you will reduce the overall amount of time/energy spent on hiring someone that meets your final bar.
If someone has to pay for a stamp it will stop spam applications.
All companies attempt to give the same interviews, just have one centralized organization give two programing questions and two system design questions and some kind of proof once you pass it.
You filter every one that can't pass the interview in the first place, you get a better interview experience, and just focus on experience
Professional certifications are different
We already have such a credential. It's called "lasting two years at a FAANG+ without getting fired". If you do that you can get interviews anywhere.
The “good” news, was, that it was pretty easy to bin the spam.
Also, if you are having trouble hiring right now, that is 1000% a skill issue. It is easier to hire good talent right now than ever before. So I have absolutely 0 sympathy for this POV. Go down to your HR department if you want to see who is at fault.
PS You fix it by charging $1 to apply for jobs. Took me all of 30 seconds to figure that one out.
Yeah, I don't see anyone lining up to game that system. Maybe you ought to think about that a little longer than 30 seconds.
That way I know I'm not giving money to some huge corporation and they know I think applying to their job should at least cost me Y amounts of currency.
And if they waste more than an hour of my time with the hiring process, they could similarly pay a charity some money per hour.
That was neither me nor the company will feel cheated and in the end, no matter how the hiring turns out, a charity will have benefited.
This could also be used for combating spam elsewhere, like posting in forums, comment sections and so on. To preserve privacy, something like zero-knowledge proofs could be utilized. I don't know how the cryptography would work exactly, but if you can't double spend a credit and you can choose whether to keep it anonymous or not, it could work, too. It would be best if for a given credit spent, you could only disclose your identity to the entity you want access to, not the credit issuing entity.
For spam, it seems like the cost of maintaining a forum like the servers are much lower than the cost of the mods that deal with spam. So instead of paying the forum directly, we lower the need for human mods to spend their time. That way we lower resources to the forum indirectly. The credits could be per post or per account creation. I assume the HN mods' time is worth a lot more than the servers and power HN runs on.
Also, we won't have the issue that PoW and other proofs-of-X's have of being easier to do on some devices, but harder on others (like the power and time it takes to run PoW on a beefy desktop with AES-NI vs an on old phone).
But we'll still have the issue with different standards of living in different places making the credits more or less expensive for the user subjectively. Companies hiring worldwide could require different amounts of credits for applicants from different countries, but for forums this wouldn't work.
A solution to that could be issuers giving credits for local volunteering work. Clean up some garbage from the shore and get a credit regardless of whether you're in the USA or Bangladesh. But if you want to prevent credits from being traded (do we? idk) and, at the same time, have some amount of privacy, how would you do it?
But now you'd have to make sure that credit issuers all over the world only issue credits for real charity-like work. And who's to say how to value picking up garbage vs volunteering at an animal shelter vs donating 1$ to a charity.
It's interesting to think about this, even though I don't have any resource to implement anything like that.
Check all that apply.
Your post advocates a
(X) technical ( ) legislative (?) market-based ( ) vigilante
(X) Requires immediate total cooperation from everybody at once
- for the specific forums, jobs and other things that may use something like this
Specifically, your plan fails to account for
(X) Public reluctance to accept weird new forms of money
- if the credits are treated as money
(X) Armies of worm riddled broadband-connected Windows boxes
- that will always be an issue, but I doubt it's too relevant here
(X) Extreme profitability of spam
- if someone spends a credit for spam and they think it's worth it, it might be an issue. But most spam wouldn't be worth it, IMHO, especially if it will be deleted from a forum, anyway.
and the following philosophical objections may also apply:
(X) Ideas similar to yours are easy to come up with, yet none have ever been shown practical
- well, yeah :)
(X) Sending email should be free
- this isn't about email, but I don't necessarily like having to pay to post. However, lots of forums will remain free, as not everyone will use this idea if it's implemented. And some forums have paid accounts now, anyway.
(X) Why should we have to trust you and your servers?
- why should we trust the credit system - important question, as we haven't thought out how it could be gamed or abused.
(X) Sorry dude, but I don't think it would work.
:)Obviously there are lots of things to figure out, but I don't see how any one of those would be a deal breaker.
Aren't you ignoring the reports of companies receiving thousands of ChatGPT-written resumes, bots sending applications, and interviews with applicants being live coached by AI?
This is a breakdown of trust on both sides.