179 karma · joined March 28, 2012
Feel free to get in touch directly: alexyang.personal@gmail.com
I agree with you that it's disrespectful to candidates. My point is that it's not just a cultural problem - it's also a tooling problem.
I'd love to add GitHub's take-home to the library, but I feel the article describes the exercise without sharing the actual prompt. If there's a public link to it, lmk and I'll add it.
2. The library is organized and the prompts are generally clear and explicit about the requirements to complete.
3. You offer support for candidates with questions about the assignment.
4. Kathy Keating told me about the grading environment you all use with rubrics, blind grading, and rotating cohorts of graders. We actually created similar tooling for the teams we work with.
I also want to explain why they're rated as 3 stars (and why I think this undersells how good they are). When we designed the rating criteria, it was really important to recognize tests that could extract hiring signal without requiring a ton of time from candidates, since that's one of the biggest issues that candidates face. So one of our criteria was "setting clear expectations for candidates (e.g. time expectations)" and another was that the time requested was "reasonable...(<4 hours)". So a test that stated upfront that it would take 8-10 hours would meet only one of those criteria. You can see the full rubric if you hover over the "5-star scale" text in the sub-header.
Unfortunately, the Ad Hoc tests don't specify time anywhere (technically missing both of these criteria) while doing many other things well that aren't captured in our rubric (e.g. candidate chat tool, blind grading). I admit that our rubric isn't perfect and it feels like the Ad Hoc tests are being doubly penalized for something that's easy to fix. In fact, if you add this to the tests, I'd happily update these to 5-stars.
Finally, if you're up for a chat sometime, I'd love to meet you. I really appreciate the work that your team has done to improve the hiring experience for candidates beyond those at Ad Hoc! You can reach me at alex@trytapioca.com
I also think there's a surprising number of edge cases that need to be considered for this challenge. Consider cars that were already parked at the start of the time range, ones that were entered during the range and never left, etc. So I think it does require familiarity with SQL.
Did you disagree with the tagging, the stars, or both?
One of the ideas my team brainstormed is to create and maintain a library like this that engineering teams can rely on.
- Provide starter code and setup instructions so candidates don't waste time on boilerplate.
- Abbreviate requirements to what actually matters. E.g. do you really need 100% test coverage on a take-home? Ask candidates to write a few tests and then tell you what else they'd do given more time.
- Use an open-ended, time-boxed format instead of having end-to-end expectations. IMO a hybrid format where a short (1 hr) take-home is followed by a live discussion/pairing afterward can be the core component of a hiring process.
I'd love to hear more about the Ramp process. Do you mind sharing what sort of practical problems they used?
There will always be a few who say no (IMO this could be a warning sign), but it doesn't hurt to ask :)
Love that your approach focuses on understanding thought process! IMO many companies focus too much on raw technical skills, when softer skills like attitude may be more predictive of on-the-job performance. My team is working to elicit the same signal in a take-home format (since it's more scalable), but I think the best is a combination of the two: short (~1 hr) take-home + follow-up live session on the work that was already started.
No candidate wants to be entered in the Hunger Games for who has the most time to sink into a take-home.
1. Bring some work that you did previously (>4 hrs) and we'll discuss it.
2. Complete a 1-hour take-home.
3. Complete the 1-hour take-home live with one of our engineers.
He said that 98% of candidates chose option #2. I was surprised at how high it was!
My favorite is a combination: short (1 hr) take-home followed by live discussion/pairing with anyone who does a half-decent job. It reduces stress because candidates will already be familiar with the code (they wrote it!) while being efficient with time.
My team is doing our best to create the "perfect" take-home experience and reduce the frequency of these horror stories.
It's usually not malicious, just disorganization. From a hiring manager's perspective, it's too easy for things to slip through the cracks. When the volume of candidates increases, it's hard to keep track of all the zip file submissions + individual repos while making sure the eng team reviews them all.
This is one of the problems my team is hoping to solve through software. So far, we're seeing teams working with us getting back to 100% of candidates (often with personalized feedback), with median times as fast as 1 day. I hope we can reduce the frequency of terrible take-home experiences.
I happen to be an Applied Math major and spent years teaching competition math classes, though this was a while ago. If I can be helpful, shoot me an email: alex@trytapioca.com.
While it's not compilable, I did reach out to a few of the 5-star test designers and asked them this question. How teams use their tests (both in weight and stage) varies a lot, but the well-designed tests usually featured as a central component of the hiring process. They were often given at the stage right before the final round of interviews, but occasionally earlier. The 5-star test designers I spoke with weighted the test heavily because it: 1. helped them reduce bias, identifying great candidates even when they didn't have the backgrounds they expected 2. gave them a foundation for the future interviews. It's less stressful on candidates when they're already familiar with the code being discussed live (because they wrote it!)
Anecdotally, on-the-job performance of the candidates they ultimately hired seemed to correlate with performance on the test. I realize this is very hard to compare against other evaluation approaches though!
FWIW my team believes that test designers should strive to scope tests to no more than 1-2 hours (ideally 1).
The reason why there aren't any 1-star tests isn't because they don't exist, but because we didn't think anyone would want to see them. Cataloguing these was a ton of work (we sifted through hundreds of tests). Including the 1-star ones seemed like it would only shame companies who used them.
To encourage more thoughtful test design (and hopefully save future candidates from the worst offenders), my team compiled the largest library of non-“whiteboard” take-home tests that real engineering teams have used. You’ll find the challenges that Stripe and Microsoft gave to their full-stack candidates, front-end tests from Tailwind and Rivian, and back-end ones from Basecamp and Revolut. Whether you’re looking to evaluate an Android, DevOps, or Data Science candidate, a bootcamp grad, or senior engineer, we found a few options for each.
Having built 20+ tests ourselves, we also rated the design of each test. The criteria for a 5-star rating:
1. Tests for skills highly relevant to those required for the position
2. Includes a well-written description of the prompt and even motivation for using a take-home test
3. Sets clear expectations for candidates (e.g. time requirements, evaluation criteria, submission details)
4. Asks for a reasonable time commitment from candidates (<4 hours)
A few notes: - We found most of these test prompts in public GitHub repos, usually owned by the hiring team but occasionally in the candidate-owned submission. We sifted through hundreds of tests and filtered out those overly focused on algorithms (aka LeetCode), leaving us with 142 tests in the library.
- The larger and more recognizable companies didn’t always have the best tests. Some of the most interesting prompts we found were from smaller teams (e.g. YC startups). This shouldn’t be surprising. Startups need to design candidate-friendly hiring experiences to compete for talent against more established players.
- There were common themes among the tests we found. For example, front-end candidates were often given a Figma design + content feed to implement, while back-end candidates had to implement an API given a set of requirements. Data scientists were usually given a data set to clean, analyze, and submit a Jupyter notebook with their findings.
- We’ll continue to update this library and add descriptions of each test so it’s easier to compare.
Have feedback, or another take-home test we should add? We’d love to hear from you!
[1] The Validity and Utility of Selection Methods in Personnel Psychology (https://www.researchgate.net/publication/232564809_The_Valid...)
HackerRank questions typically don't meet criteria 1 or 3. The best tests are often ones that the engineering team has designed themselves.
For example, this is an example of a test that meets all the criteria above: https://app.mightyacorn.io/flow-club/founding-front-end-engi...
1. Design tests to be open-ended and demonstrate how a candidate thinks, rather than ones that drive toward a binary pass/fail outcome. Asking why a candidate worked on the pieces they did and where they would focus if given more time gets you high signal without requiring as much time from the candidate.
2. Provide clarity on the grading rubric. When candidates understand how they'll be evaluated, they'll spend less time on work that doesn't result in higher signal to the company. For example, having a candidate write some tests or talk through how they would approach testing is great! Expecting full test coverage isn't as useful.
3. Scoping is the toughest part about test design. Try to choose topics that closely mirror the day-to-day work in the role itself but that can be boiled down to a 1-2 hour task. You want to make sure that candidates who perform well are more likely to succeed if hired, but you also want to keep it short. From interviewing 50 hiring managers who designed the hiring processes at Basecamp/Medium/Mailchimp/etc, they revealed that the candidates who were ultimately hired consistently spent 2x the recommended time, which means that designing a 3-hour take-home is asking top candidates to commit 6 hours. It's a tough ask in the current hiring climate.
Here are 3 examples for inspiration:
- Full-stack role for Flow Club (YC S21) that runs virtual WeWork sessions: https://app.mightyacorn.io/flow-club/founding-front-end-engi...
- Front-end role for Playhouse (YC S21), building TikTok for real estate: https://app.mightyacorn.io/demo/playhouse-hiring-challenge
- Back-end role for zero5, a parking automation startup: https://app.mightyacorn.io/zero5/challenge-fullstack-softwar...
Happy to share more details if you have specific questions!
Source: my team spent the past 6 months designing personalized take-home tests for 20 VC-backed startups hiring founding engineers. We evaluated ~500 candidate submissions.
If you want to check out an example of automated Craigslist scraping, you can check out a search tool I built (craigslist-scraper.herokuapp.com) and the accompanying tutorial (baserails.com/apiscraper). Both are best viewed via laptop.
If you have any specific questions about BaseRails or Rails/coding in general, I'm happy to help. Feel free to reply back here or shoot me an email at alex@baserails.com.