Interviewing Engineers in the AI Era: Lessons from a Year of Rebuilding
coinbase.com
coinbase.com
I’m not saying that all code should be typed by hand in 2026 but there are certain subtle things you learn only when you get into nitty gritty details especially related to security.
Also today’s AI is notoriously bad at ownership. When you ask it to give you a concrete answer, it will still give you options with pros & cons of each so that ultimately you own the decision and not it. So how do you decide between the two (or more) when you never learned to do it yourself?
I think upper management are hoping they can replace a possibly reticent or AI skeptical "old guard" with AI native Gen Zs. In practice, the young out of school kids are much less excited about AI than the 50 year olds, so I don't know how well it will work.
https://futurism.com/artificial-intelligence/zoomers-ai-sabo...
Shouldn't this be a pretty fundamental part of a degree? If you're in school right now for a CS or SWE related degree, I would imagine you're learning how code actually works, how the math actually works, etc.
If an LLM makes a mistake, you should be able to call that out.
You might be arguing that a junior engineer may not have the RWE to make those judgements, but by the time you've completed a 4-5 year degree, you should have done the following:
- Completed coursework where you've learned the fundamentals of programming and software engineering
- Have done dozens of projects (building everything from basic web apps to more advanced pieces of software) where you've seen what works, what doesn't work, etc. This also gives you real world exposure to the latest and greatest frameworks, tools, etc.
- Have done several (at least 2) internships where you've worked at a real company writing real code, and have seen/been mentored into what AI is good at, where it fails, etc.
I’m sorry to say but colleges don’t teach the kinda stuff you’re gonna use day-to-day. They’ll focus on architecture, data structures, algorithms, databases, OS internals theory that only a minuscule number of systems engineers would get to work on. They won’t teach you that in your /login endpoint, if a user is not found, you should verify the password against a pre-computed dummy hash so the response delay matches a real user account workflow to avoid timing attacks. You only learn this on the job under the supervision of a senior mentor.
This was not my experience in college. A lot of it was VERY applied. Granted, that was 13 years ago.
But even so, if you're at a college where you feel that you're not getting enough exposure, that's why I also called out "Have done dozens of projects" and "Have done several (at least 2) internships".
> They won’t teach you that in your /login endpoint, if a user is not found, you should verify the password against a pre-computed dummy hash so the response delay matches a real user account workflow to avoid timing attacks. You only learn this on the job under the supervision of a senior mentor.
Actually this kind of timing attacks were taught to me in information security and cryptography class, alongside other side channel attacks.
But this is about what comes next i.e. seniors of tomorrow
If AI gets so good at software that vibe coding is the new norm and as Elon Musk says that AI will generate the machine code directly without any intermediate compilation or interpretation then we wouldn’t need seniors or juniors but I don’t think that’s coming any time soon, if at all!
Regarding seniors of tomorrow, I think there are big differences between good devs and sub-mediocre ones, and I think their proportion won't change, and the good ones will have the drive to use the AI to learn and understand because they are simply curious and want to know. But most programmers don't care at all, and that's the case even today, and they will be less forced to learn. But I don't think we will lose much with this. Important things are held up by a small proportion of engineers, a Pareto-like principle.
Though why would companies make that nontrivial investment if they can use AI instead?
Furthermore, I’d argue that learning good judgement takes around a decade of full-time dev experience. In particular, experiencing the long-term consequences of one’s design and implementation decisions. You don’t get that just by college education and a few internships.
CS and SWE are radically different subjects. Science vs Engineering. I'm sure it varies by school but many CS grads will have approximately zero exposure to engineering concepts or any of the latest and greatest frameworks.
It always baffled me that US colleges seem not to offer such programs. Or maybe they do, they are just not prestigious enough.
LLMs are, by now, pretty good at not making "code doesn't work" and "math doesn't work" kind of mistakes when writing code. These are also "easy" things to get good at, you learn what each part does, understand the abstractions, ensure it makes sense, and go on with your day. Unit testing helps here.
LLMs are not that good at not making "this works wrong" kind of mistakes. Maybe the code compiles and does what it has to do, but maybe it's 2 lines of code with 6 lines of comments (looking at you Claude), maybe it defines three helper functions it doesn't really need, maybe it does something "here" when it should be doing that something "there" instead, maybe it finds itself in a framework and completely disregards how the framework is supposed to do things, etc etc. These are harder things to get good at and you WILL end up with an unreadable mess if you disregard caution and let the LLM go at it.
I would fully expect a brand new college grad to call these things out. There's nothing in here that requires extensive experience to understand. These are basic principles any SWE should know. Nonsense comments are common sense to pull, if the LLM is pulling in a framework you should look into that framework and understand how it works. Don't know the framework during the interview? Say that. Tell the person interviewing you "looks like it's pulling in XYZ. I'm not entirely familiar with that. I understand at a high level what it's doing, but I'd want to dig deeper and understand if the LLM is doing this part right"
Today we have AI, which is basically like having infinite libraries available that do what you ask for. But you will not learn if you just take that code, similarly to how you don't learn if you just call scikit-learn to train your SVM.
And of course students grumbled back then also and said why do we need to do this when all those libraries exist?
Learning often requires not taking the most efficient path for every project.
we will tell ourselves that we should wait eagerly to see if we will be graced by "tibo's" largesse, like peasants clamouring to see a royal throwing bread among the masses.
or that it is right and good that claude should refuse to answer on account of safety, but hack other companies on behalf of anthropic. surely we are too simple, and sometimes naive, truly we know not what is best for ourselves.
no more.
Vive la révolution!!
Meanwhile, how many jobs have AI critics created? Crickets.
I did find my interviews (with humans) to be quite tight too
Where are they getting all those free tokens from?!? I don't like this rhetoric. It still costs money to write code, only now that wealth gets transferred to Anthropic instead of to individual contributors.
Sure, for some cases, the cost is going down. I and others on my team "one-shotted" impressive features in an hour of agent work that would have taken humans probably a month if done completely manually.
I've also seen agents going in circles for an hour on a fix that would have taken even a junior 3 minutes to get right (after it was struggling for 10 minutes, I wanted to see if it can ever get it right as an experiment).
And when taking a look at the whole organization, orgs still just aren't really shipping that much more quality features to their users as the impressive one-shot demos would make you think.
Writing code was never the bottleneck at large enterprises.
So I'm increasingly uncertain about what and how to test. My default for now is still to rely on ability to write basic code fluently, but I'm open to changing this perspective.
I really want to know how this existing repo AI-assisted live coding test works, with example problems.
It seems the standard data structure puzzle type thing won't be feasible if you are using an LLM.
Also the latency for these agentic coding/prompts seems like it would make the interview a bit awkward.
Anyone been conducting or taking interviews with this kind of thing with thoughts to share?
I ask them to implement xyz thing. What I’m looking for is how they interact with the AI agent. Do they ask the agent to plan first? Do they review the plan? Etc. It’s pretty typical stuff that you might expect an experienced engineer to do if they effectively use such tools daily.
There are a series of follow-up questions about how to productionize the toy system, which gives some additional signal about how well they understand what they’re making. I sprinkle these in when there’s dead air waiting for the bot to think.
I think we’ll need to evolve and refine this problem and process as the models continue to improve.
It's also the kind of stuff that somebody can learn in a week, so not hiring the right person who just didn't spend this week of time yet for whatever reason is a loss.
I agree. A technology professional who views themselves as "one of the cool kids" (read: easily manipulated by social media) is a legitimate security threat, as are many of the popularly promoted approaches to "LLM-assisted software developement".
If I merge a PR that was opened by Claude do all the changes in it automatically count as AI-generated? How about the individual commits? If I'm reviewing locally and make manual changes but then have Claude create the commit it appears to be AI-generated but might not have been.
The point is that these metrics are easy to game and I've definitely wasted time and tokens refining code with Claude that would have been easier to just edit by hand. It can be kind of fun and when the goal is just "use AI" I don't find it surprising that graphs showing 100% switch to AI-generated code could be defensibly generated without really saying anything about how much manual intervention is happening or how efficient the process is.
Now I don't even know! Colleagues send me graphs they made with Copilot and then we discover that the LLM did mental arithmetic or whatever to produce the results and they are wrong. Computers used to be able to do math !
- does the company provide the harness and specific model for this? Or the interviewee use whatever they have access to? If they don't do a good job, is it the fault of the harness or the model or the interviewee not knowing how to fully utilise the harness and model or alternative harnesses and models?
- if every interviewee uses different harnesses and models, how does the company ensure that it's a fair comparison between interviewees for the same role?
- will the interviewer go through all lines in the code generated or let AI do it?
- We have embraced everything about the AI Era at our company
- Arbitrary topic about how that changes something
The basic problem is, no one on our side has actually doubted the ability of candidates to us the AI tools in conventional ways.
Consider the design section here. I think we are all confident the candidates can query some AI tool for a design. It's basically copy-pasting the description in. The real goal is to have a discussion, to see if they understand the relevant issues. But then, we could just hand them a design.
It just seems like you waste time having them fiddle with prompts and everyone trying to read the output. Just skip to the discussion.
https://xcancel.com/brian_armstrong/status/20516167591451857...
- You get a leetcode question and if you're lucky is an easy medium that you can solve, if you're really lucky you already solved it and can pretend you are approaching the problem the first time. Good luck if you get a hard question and you never saw it before.
- You get a home assignment, in a framework you might not know but you're expected to be fluent with it, then waste 1 hour setting up the project structure, and one more hour to find out how the framework expects you to define the CORS allow list. You are expected to deliver the project in 3 hours.
The good I see in AI is that it completely removes the need to study just for interviews, and you can also delegate all the project setup to the AI. Then you can focus on what you would test (e2e? integration? what are the boundaries? what do we mock?), how to keep the documentation, how to structure your code. You have an expensive endpoint, do I make it sync or add an async jobs framework?
Imagine you're an expert in C++ interviewing for a Django position and the interview consists of fixing a big in a repo. The bug is that a function without type hinting is modifying what is expected to be a list, but the caller is passing a tuple. Trivial after a week you work in python and you have your environment set up for type warnings, also trivial with AI and definitely not an interesting problem that shows expertise with software engineering in general.
We also did this in our last interview at work, and it was a really good indicator to see if someone just copy pasted code, or understood it after it was generated. Some candidates had a unit test fail and couldn't debug it for his life, even if he "wrote" all the code himself. Others simply did not understand the architecture they wrote, and assumed that a function defined with "async" and awaited would run in parallel from the code that called it (as if you spawned a thread)
We explicitly said -- feel free to use whatever framework and AI assistant, just show us how you do it. The practical part had no leetcode too. Just build something really basic, then explain a snippet of code (3 lines) and generalize it. A trick question (with a disclosure it's a trick question) if the candidate did it fast enough that we didn't have to go into the overtime. A bit of theory about protocols, all in all an hour and we leave another 30 minutes on top to answer questions.
At the end of the day we just filter out with confidently bad takes, people who can't do 2+2 and ones that can't understand the question without rephrasing it three times.
The most bizzarre candidate didn't know anything at all, but was so relaxed and confident, that he spent all of the 30 minutes asking about the company and how his day would like and all that, while he clearly bombed it.
The difficult part is how to not filter out a competent person who doesn't necessarily agree with all of your takes, uses all the same tools and had all of the same experiences as both of the interviewers.
We resorted to filtering candidates in-person with 5 basic technical questions on pen and paper. And I mean _really_ basic questions.
This was surprisingly effective because it filters for many non-technical skills like being able to read and write English, follow instructions, and show up to the office, on time, and appropriately dressed.
The number of candidates that failed these basic skills was astounding. We had candidates show up 20 minutes late, or email 2 minutes before the "interview" asking for a Teams link even though the invitation stated the meeting was in-person (highlighted in yellow). Others couldn't write their own name legibly on the paper.
A candidate that cannot answer basic technical questions has no hope of being able to prompt an AI effectively or review the code it produces.
They have to earn it, as the tokens are not free.
Given that deskilling and over-reliance in AI assistance will continue to happen, putting a hard token limit <100k tokens in the interview process serves as a great filter to prevent the vibe-coders and "tokenmaxxers" out and forces a higher bar for quality, with clean code and reasoning across well maintained software with less tokens rather than increasing the slop.
Do you want a candidate that knows when to use AI and carefully uses tokens with in their limits, or do you want a candidate generating incomprehensible AI slop to be tokenmaxxing out your company limits and then draining your company bank account?
The whole point of Bitcoin is sanctions evasion. Tether moved from Deltec Bank (CIA) in the Caribbean to Lutnick's Cantor & Fitzgerald and El Salvador. That is where the action is.
So promoting the AI bullshit is for investors because that is what they want to hear. They are not going to vibe code financial transactions and get another $500M EU fine.
As an aside, it is interesting that the new "taste" talking point was already on Coinbase in mid-July.