The Machine Learning Job Market
evjang.com
evjang.com
^ this is increasingly a choice in your career -- 'where can you go to solve big problems', and 'big problems are increasingly complex'
scale is real, and tools matter. you can spend your whole project burn at the wrong company building something that you could buy somewhere else, or which already exists at a competitor
slight grain of salt here is that G's logging system, from my perspective as a gcp user, is slow as balls and the UX is the incarnation of scroll jank. and also this (very good) article led to the outcome of the author building soft hands happy-ending robots
I don't think the FAANGs will ever come up with AGI and I am glad for that. If the private sector gets there first, I am going to "accidentally" die of hypothermia on a hiking trip because I will do anything not to live in their world.
Writing TFX code at Google is like having your soul-sucked through your rear-end! Imagine TF1 with all the broken APIs, but now it's all distributed! Fun.
I laugh-snorted reading that! I am stealing that phrase.
A common thing I do is say company X has a fleet of Y assets. They repair them every N years. A good solution would be to predict which ones need repairs. Do those ones more often. Do the healthy ones less often. Pay more attention to the ones that are valuable.
Better outcomes, millions less spending. Probably don’t even need a live model in prod. Just a semi annual manual export run to excel for some planner guy who’s been keeping the schedule for 2 decades
For those of us who don't work at such scale, can you (maybe with a little fuzziness to avoid telling too much about an internal project) give a few examples of the kind of projects where a fairly simple model can have a 1M+ impact?
Interesting examples, and yes, they're all the kind of thing that might not justify the effort for an ML model (and might not have enough data to train) for a small website or operation, but can easily justify the cost and effort when you have a huge number of transactions.
On another note, this is why I often like lightening talks. So many people think that what they're doing falls below the threshold for what is an interesting presentation, when in fact it's the most relevant thing a lot of people will see at a conference.
What I refer to for my work is the low hanging fruit. Old problems that businesses solve with manpower or overly unspecific rules. Something where just a little clarity can help them hone their efforts on the 80/20 of it all. I made a slightly more detailed post in an adjacent response
Google has massive data and scaling advantages that can never be duplicated or fixed by smaller companies.
does GCP use G's internal tools?
Inquiring minds would like to know
Obviously the principles and theory behind scalability is still important for properly structuring your app, but there won't be many novel problems to solve, and increasingly obvious architecture choices as time goes on
Granted, they won't have to think about designing a solution, but they still won't have the computational power which can be afforded by larger companies (and cloud computing is ridiculously expensive), unless they have a lot of money .
On the other hand, I also worry about getting sucked into the bureaucracy of FAANG sized companies & not having any accountability or agency over what I work on. (I realize this is a sweeping generalization of FAANG, but some of my peers have had this experience even a few years into their jobs)
At that compute level, you should be able to at least replicate SOTA in optical flow, structure from motion, speech recognition, text to speech, translation, text summary, sentiment analysis, image classifications, image segmentation, and of course playing video games or optimizing processes with reinforcement learning.
I mean thanks to KiCad even custom sensor hardware is cheap these days.
Can you give more details about what you tried to do and why that wasn't possible?
Where I do agree with you is that transformer-style text generation models in the billion parameter range are off-limits for hobbyists. But that's only a tiny part of the useful applications of AI. And you can train them with gradient checkpointing, it's just 100x slower than what Google can do.
But back to my original point: training stylegan takes a lot of resources at higher resolutions. For something like thispersondoesnotexist on stylegan 3 you not only need to have lots of quality data, you also need many cards for many days.
The point is, Google is a mature company. Mortals like most of us don't get to break ground in a technically advanced but mature company like Google. Instead, we find fast growing new problems to solve, to hone our skill, and to get to scale.
P.S., I personally know a number of prominent professors used to work on the Borg projects just to optimize for a few percent of gains. It's deep and interesting work, but nonetheless hard for mortals like me to get much out of.
That is, it's a more sure bet to work for a baby Google than to work for a middle-aged Google.
I agree that sometimes our actions have unintended consequences. These are often hidden effects, which makes it easier for someone to overlook, even when they are trying to do good. But that’s just a generic idea - the specifics matter.
Let me ask you a question -- what proves your ability to work on scale as an engineer? Is it the scale of the problem, or the scale of your solution?
Outside of "Google scale" logging is a nearly trivial, solved problem. If the only thing that creates the "challenges of scale" is your data cardinality, guess what. You're only solving problems of scale in the most literal (and maybe trivial) sense. Working on that at Google isn't going to prepare you for how to architect a platform for a startup that scales from 0 to 1 million active users over night without breaking a sweat.
I am biased in that I care about the latter kind of scale and effectively couldn't care less about the former, because the latter is generally an existential problem to have, while the former is a nice problem to have.
Oh look it's me. I'm sitting here building a worse version of TestComplete and Ranorex for automated testing of a Windows desktop application.
Looks good on my resume though.
The glorified pattern matching can only take us so far. You know it's working as long as there is a pattern. I wouldn't call it a general intelligence per se. There is no "juice" in these algorithms.
If we use these tools, we can immediately see where they fail and where they do not. These are just a new tools in software engineer's box.
This argument is becoming less and less convincing year by the year. We're amazed that things we were sure couldn't be done are actually done.
I just didn't want to call it as "intelligent" and use this as basis for defining "intelligence." We can call them something else. It's learning to do a specialized job as intended and in "intelligent" manner. But it is not intelligence. Even a small ant is intelligent than our current AI systems though they aren't sophisticated and can't perform human task, they are intelligent than AI system.
I hope that made sense.
TL;DR We're also mostly brute forcing our way to discoveries. We're not that smart.
People too are relying on cultural handouts, maybe most of our intelligence is also "something else". Before electricity was discovered we had superstitious ideas about electrical phenomena. Before germ theory was discovered we were getting sick and dying like animals, helpless. Not so smart, even though it was a life and death situation for us.
It's easy to be "intelligent" when you're given the solutions beforehand by culture. ML learns from the same culture, like 99.99% of us who can't discover new things even to save our lives. And many of our discoveries are a gradual work of trial and error, we don't go directly to the target but stumble/brute force our way to it.
There was a news story recently title "Elegant Six-Page Proof Reveals the Emergence of Random Structure". The funny part is how the authors stumbled onto the amazing solution after many many unsuccessful trials by all the math community. Not a great sign of intelligence when you have to rely on chance so much and so many fail before one succeeds.
This tells me we're also mostly doing "something else". Intelligence means solving novel problems with few attempts, not spamming our attempts to death until something comes out. ML research looks more like spamming than intelligence too.
You know what else looks like spamming? Evolution. It's a blind search process brute forcing the problem of self replication for billions of years. It created us and everything else in one run but it's not very intelligent, it just spams a lot.
You're absolutely correct that patten matching AIs won't ever be truly intelligent. But then again, many humans also never exceed what can be simulated with good pattern matching. And an AGI household robot only needs to be as smart as the maid that it's replacing.
I'm optimistic that pure pattern matching will get us to usable AGI AI.
There's never going to be training data for "how things are going to be next year". A lot of large scale systems involve emergence [1], patterns which previously were not visible suddenly appearing. I think even today's AI can do things that a bit beyond pattern matching (learning to learn, etc). But pure matching as such is inherently limited.
What we need is a pattern detector + a the ability to create basically infinite ANNs (or be able to multitask on them) + an event loop that takes input feeds (from cameras, microphones etc,) does some kind of reasoning and then pushes to its output feeds (wheels, etc.)
I think you use pattern matching to extract unique objects, store these objects as a node with its own simple neural net + long term storage where it only stores pictures of this object plus a dataset about it e.g, how often you see it. You then you organize them into an object hierarchy. Each new object is compared against all other objects we’ve stored using their pattern marchers. The higher the output the more weight we give their “connection.” Each object is made up of of sub objects so they are the top of their own tree as well, so you can run this pattern finder on the dataset of individual objects itself and if you find new objects the tree recurses. You can then check these objects against existing ones etc.
A general intelligence does this constantly, in real time. Then it’s a quick algorithm.
1. Have I seen this object before
2. No, but it shares characteristics with animals (an object that groups together all things that look like animals.)
3. It’s much larger than my dog, and I’ve seen large animals attack small ones more often than not.
4. My dog is also a dick, and attacks other animals more often than not
5. It’s probably a threat
Just scaling modern compute won’t get you there unless you’re willing to dedicate a few orders of magnitude more energy than a human being to do so. You need a completely different, distributed, architecture if you are going to be able to compare billions of objects against billions of objects every time you see something new and in real time.
Machine Learning is great but it’s only the learning part. Intelligence is reasoning about multiple things in relation to one another not detecting a pattern. You might trick yourself into thinking you’re getting there because pattern matching is powerful but it’ll get you to the intelligence of microbe at best. Even then you need something that’s driving the actions.
I worry quite a lot about malevolent humans using enhanced technology (note that most things in technology, once accomplished, cease to be AI) will do. Authoritarian states and employers can already learn things about you (that may or may not be true) that no one should be able to know from a basic Google search. This is going to get worse before it gets better, and if corporate capitalism is still in force 50 years from now, we will never achieve AGI in any case because we will be so much farther along our path to extinction.
I worry quite a lot about malevolent humans using enhanced technology (note that most things in technology, once accomplished, cease to be AI) will do. Authoritarian states and employers can already learn things about you (that may or may not be true) that no one should be able to know from a basic Google search. This is going to get worse before it gets better, and if corporate capitalism is still in force 50 years from now, we will never achieve AGI in any case because we will be so much farther along our path to extinction.
It doesn’t need to be malevolent. You’re made of atoms and if the AI has uses for those atoms goodbye you. There’s no reason to believe consciousness has any impact on the ability to maximize an objective function, i.e. try for a goal.
This article is representative of an attitude I'm seeing around the tech industry, and if this is indeed the level of "confidence" in the Bay, I don't think that's a good sign.
In marketing, they say 5 years when it's actually 20 years away.
We know how to do fusion. We know the physics behind it. We haven't yet figured out how to build profitable fusion plants, and we probably won't for a long time, if for no other reason than improvements in fission--modern fission plants are the best .
When it comes to AGI, we have no clue. It's a constantly moving target, because our conceptions of intelligence evolve. Most things that were once "AI" became "non-AI" solved problems after we got good at them using a couple key insights, e.g. that image processing could be sped up with CNNs due to the existence of a topology on the inputs. We still have no idea what makes us tick, and moreover there is not a strong economic incentive to replicate all of our intelligence... although, of course, automation will continue and that itself will be disruptive enough.
Has nature ever proven that nuclear fusion can be sustained at human scale?
I would bet that agi comes first.
We're fairly certain it can be done given unbounded resources, we have some idea of the principles involved, but then there's a rather significant element of "draw the rest of the fucking owl" between where we are and where we imagine we could go.
I'm not aware of anything he's accomplished but can see the delusion. ML people seem to think the output of their work is not mediocre. Yeah, you bred monkeys till something resembling shakespeare appeared to some reproducible consistency and it is better than something someone can code - but that's an incredibly low bar.
Acknowledge that were still very much in the stone age of AI and what were doing is large scale analytics at best.
He's exposed to enough corporate work. https://www.linkedin.com/in/evjang/
If he doesn't like this place, he can just make another post like this and I am sure ML startup CEO and ML division heads will be flooding his inbox.
Does the opportunity exist to transition into any particular ML roles then grow from there?
For example when I hire MLEs (which I am doing now if anyone wants to apply - supportlogic.io) I am willing to look at people who are solid Python/backend engineers and who have been "ML adjacent" or who we believe could learn the ropes of ML enough to contribute. The stronger an engineer, the more flexibility we have in ML knowledge. Some ML engineering is task-specific but a lot of it is automation, data engineering, and improving data scientist code (for which you do need ML experience
I've found it's a lot easier to teach an engineer enough DS/ML fundamentals to do ML Engineering than it is to teach a data scientist engineering skills. A lot easier...
The other side is deploying it efficiently, and that becomes a more routine software engineering problem. Fundamentally you have some code that you want to run as fast as possible on the cheapest hardware you can feasibly use. Large companies like Google have the luxury of splitting this out into several distinct roles - from pure researchers (people publishing papers), to people who train models for business purposes (eg the Google Lens, computational photography, Translate), to people who optimise the ML library code underneath, to people who build out the end user application with the ML model as a black box service.
Most of those people don't need to know much ML, but the exposure can help you transition into a more ML focused role.
I don't know where OP is getting these figures from, but I doubt that FAANGs offer 7-figure comps to Staff-level people. It's probably more in the higher 6-figure level (400K - 600K).
https://www.levels.fyi/Salaries/Software-Engineer/Machine-Le...
"You can only be level X with compensation Y after Z YOE" is one of the greatest infohazards in tech.
In those jobs, you don't even have to show up (although most of themdo, in order to stay relevant but also because these ex-academics tend to be very driven people). You're getting paid well for letting the company say you work there, because this makes it easier for the rest of the company to fill out their chain gangs of early-20s Jira jockeys who'll be used for three years and then PIP-raped because "yellow zone".
In all honesty, after 6 years of studying, with 4 of those years studying ML, to be told that I lack experience with some particular stack as the ~~excuse~~ reason for rejection feels like a slap in the face.
And all that ignoring everything that expects 3+ years experience for entry level positions.
Tl;dr I moved to Japan and worked in ML (ish) job. Once you start working it becomes remarkably easier to not be scoffed for lack of experience (even if you learned very little in that job)
Job security is high in Scandinavian countries and as an consequence people hire very risk-averse. risky/not so established jobs such as ML positions, they'll be actively looking for reasons NOT to hire you
There is this interpretation of Bayesian networks that the parents are the true causes of a node and to predict what happens if you change a variable you need to remove the edges from the parents to that variable from the network. And then I study what you can do with that method
ML is the sexy wing of the tech industry, so it tends to attract the people who are willing to put in the hours (for interviews as well as towards work).
It enables the business to say they invest in R&D (tax writeoffs, marketing) and it also makes it easier to hire the grunts who'll do the scut work, thinking they'll one day be working on something more interesting (which they won't be).
Competition for bullshit "data science" jobs isn't that tight. For genuine research that actually matters, it is, and that's because most of these positions exist as recruiting tools (you're getting paid to let the company say you work there) and that's always going to be a thin ledge to try to perch on.
Today, the competition for research positions in a "sexy" STEM field (think quantum computing, black holes, DL, RL) is quite high. DL just happens to be a field where it is possible to get into research without investing too much time while also making a lot of money, so this is a dream job for many.
- The fact that it isn't representative is what makes the article an interesting read.
- The fact that they claim to have a plan for solving AGI in 20 years really detracts from their credibility.
I see they're going the cold fusion route.
> I’m not like one of those kids that gets into all the Ivy League schools at once and gets to pick whatever they want.
Followed by "FAANG + similar" and a deluge of options. Also, I feel like their message is pretty liberal with using future projections and implying it to be the present. For instance, the author has 6 years of experience with 2 at the senior level. This is pretty far from "staff level" (at 1M+ compensation, I think this is L8) which they imply is/was an option at a FAANG company. I don't doubt that in 5 years they would be at that level, but they almost certainly did not get offered a "staff" position at a FAANG.
It's different if the level isn't represented in the title - if they went from one band to another but the title was the same, I don't see a problem with putting down something like "Software Engineer, 2016-present" without wasting space on each promotion.
Now, I don't object to adding a line like: "promoted twice from Software Engineer 2 to Staff Software Engineer" or whatever, which I think is a good middle ground (and I would put this as the very last, least important entry for that company)
How much does that matter to you, anyway?
Thus I have to make sure either way and probe a lot unfortunately. Did they hold the title for the last 2 months and are jumping soon after? Will they do the same here? I want to know about and see the progression. There are situations where it is sort of irrelevant but in others it is detrimental if I have to probe.
If you are say in year 5 of your career and at senior level at just one company I want to know if you were a junior when hired out of college, were super awesome and made intermediate after one year and have been senior since year 2.5. You can show that to me right on the CV by listing it individually. If you just put the end title and that's it I will assume that you made senior this week and are trying to jump ship. This doesn't mean we can't figure it out together in the interview if it gets to that stage. But it sets a certain tone and connotation for the entire conversation. A bias to overcome.
I knew someone at Google who was hired as L3 straight out of college (as all non-PhDs are) and got promoted once a year to L6 (Staff) so 3 years. He got promoted to L7 2 years after that.
It's a rare combination of talent and the right circumstances but it does happen.
Low 7 figures is realistic for L8s without the share price noticeably appreciating. With discretionary grants some L7s may squeak into that club. But L6? No way.
That might matter.
As someone who has worked at FAANG for 5 years right out of school, getting to staff is less about raw intelligence and more about being lucky with working on projects that did not get canned and finding supportive managers. My friends much smarter than me have not had a good growth purely because they were unlucky with initial team assignment and PA / reorgs cancelling their projects.
I'm around 15 years of experience, and my appreciation for my own lack of knowledge and ability to make predictions still grows with every year.
On the definition of experience, I agree that education and life experience counts. I said "around 15" years for myself because the definition is fuzzy. I got some very specific career preparation and training in college, so that sort of counts, and I probably spend more personal time than many of my peers learning about relevant history and current events.
If I ever left my job I might have to quit DS/ML and do something else entirely.
At the same time, nobody is going to do all that well if they just apply online and cross their fingers. You need some kind of human contact, either though an introduction to an insider through your network, or through a recruiter of some kind. It takes a bit of time to develop the relationships, but it's quite doable and worth it, even for introverts. Best to start before you're interested in changing jobs.
that's what I call an understatement.
Odd choice of level, since the author worked at one of these companies and was not at that level, certainly did not make 7 figures.
(Not saying anyone "deserves" that or that's how it should be, but that's just how it is here in the valley.)
I'm a bit sceptical on this 10,5,1 whatever year ahead metric he pulled from wherever.
Interesting read regardless. My opinion about the next few years is that most value will come from finding your niche, creating Datasets, iterating and building your ML model (a bit like he wrote but without this AGI...)
The author apparently does not understand regulatory capture and is throwing around catch phrases to sound smart. Regulatory capture would imply that healthcare encounters less regulation than it should due to influence over the relevant government agencies. This should increase product impact and reduce time to market, the opposite of what he suggests.
I think the author's point still stands.
I believe that is unlikely. Here's metaphor for that, that I often like to use when I speak to engineers doing empirical ML work: In ancient times people build a lot of nice structures, such as pyramids, cathedrals etc. By trial and error many rules of thumb were devised and they more or less worked. But it's safe to say things like earthquake-resistant skyscrapes and modern bridge cannot be build without deep theoretical insights into structural mechanics. These are highly optimized, intricate structures. The same is probably necessary to build highly optimized, intricate models that deliver what we now consider to be "AGI" - but then again, the world of ML is full of surprises.
And the publicly traded crypto companies compete with FAANG on compensation too.
Non-crypto startups are the only ones sitting in the doldrums left out to dry right now.
Would love to hear more about the startups; I tend to turn down such opportunities far before we talk non-cash comp.
Outside of publicly traded crypto companies you need to talk to a third party recruiter in that space
Solana Labs, for example, one of many, was paying engineers $650,000 back in 2019-2020 (and still is) to mostly write in Rust. Compensation was ~$200k cash and $1.6-$2 million in Solana tokens vesting 3-4 years with 1 year cliff. Solana tokens were $.10 cents back then, so those engineers are sitting on like $100 million+ as Solana trades at $100/sol now, down from $250/sol.
For more typical results, companies that pay in crypto only have a few employees so giving them all a few million dollars in their much smaller less successful crypto still results in being able to liquidate close to the notional value they started with, derisking your time and coming out ahead in general. A “tiny” crypto is still like a $30 million marketcap. Even the $300 million marketcap ones are considered tiny. Market depth / liquidity is usually enough to support a few million dollars of periodic employee sell pressure.
tech sector is fast, crypto subsector is like an order of magnitude faster. its similar to tech employment in the 90s where there was fast vesting (mostly due to quick exits), liquidity at super low valuations that then rose extremely quickly and attractive compensation. the main difference now is that the valuations are much much higher. you can tap in sometimes/often at very low valuations - of the token - and also ride them up all the way to billions valuation very quickly. if they solve a market need (within the crypto space) then they attract value very quickly, sometimes that market need can just be the entertainment coming from hype, but most times its bandwidth since there is not enough blockspace to go around, periodically.
Team and advisor allocations have been this, and have been my best trades. Vesting grants for employees can be lucrative too. Often times these are also discounted prices to whatever any buyer can get. So things amplify very quickly, and there are less ways to lose.
> For more typical results,
I always thought that human shaped robots are a terrible form factor. Why limit yourself to the awkward design that 3.77 billion years of evolution accidentally landed on?
If you have a more specific goal in mind, e.g. solving a small set of industrial/commercial use cases, that changes the calculus dramatically.
Its rare that an article loses my faith in the first sentence.
https://images.squarespace-cdn.com/content/v1/5de799b06bb59b...
Are humanoid robots just around the corner? Musk claims Tesla will have a "prototype" humanoid robot this year. I dismissed that as Elon hype, but have I missed this coming?
You will have better luck working with one of the larger companies who have a good history with the FDA, and more importantly, have good relationships with hospitals and physicians. They are aware of the time and resources it takes to push something out. Pay will not be FAANG level or anywhere close, but they usually have great work culture and WLB.
That is why all the big like Google no longer push things forward, even with the best engineer and Phds and with so much money.
That’s why I think that Optimus at Tesla will crush all the other robotic platform.
On the positive side success of Optimus will help startups to get funds or get acquired by corporations that want to get a slice of the newly proven market.
Look at what SpaceX accomplished. Look where OpenAI is given the time it’s been operating. Look at Tesla rate of production increase.
It’s going to be very hard to compete with Tesla at this point. So much ressources, so much bright engineers, all the knowledge in manufacturing, all the training on vision, etc.
But before Tesla had more limited resources, now it’s almost a non issue.
Is that what one gets paid 7 figures for - to go out in public and claim they "know" things like this with a straight face?
Not only is there no 'T' in FAANG, but the industry/product is completely different.
Or maybe the author just wanted to make a joke about coffee. Who knows.
He graduated in 2016, worked at Google in Bay Area, and now is joining a startup at a VP level.
I graduated in 2008, obtained a PhD in 2014 in a no name EU university, worked in odd companies for a while and joined FAANG 4 years ago as a mid level developer, where I am still ATM.
Looking at this disparity I wonder what could be possible explanations:
* OP is a beast and has grown very quickly in a short time.
* I'm particularly inept and I'm growing very slowly.
* Working in the right conditions (e.g. Bay Area, Big Tech, right team) can greatly accelerate your growth.
* Startups have a big title inflation.
Why does valley culture makes it seem like everything is possible and anything innovative can happen soon? The innovation in AI really seems like it is being made on a thin line of engineering and compute. It doesn't happen overnight. It requires some people working through and through to pull it off. These days it requires collective contribution.
This perfectly echoes my own thoughts. The advances being trumpeted in AI are functions of hardware advances that allow us to have massively overparameterised models, models which essentially 'make the map the size of the territory'[0], which is why they only succeed at a narrow class of interpolation problems. And even then nothing useful. That's why we're still being sold the "computer wins at board game" trope of the 90s, and yet somehow also being told that we're right on the verge of AGI.
(OK, it's not only that. There's also a healthy amount of p-hacking, and a 'clever Hans effect' where the developer likely-unconsciously intervenes to assure the right answer via all the shadowy 'configuration' knobs ('oversampling', 'regularisation', 'feedforward', etc). I always say: if you develop a real AI, come show me a demo where it answers a hard question whose answer you - all of us - don't already know.)
[0] Or far larger, actually. Google the 'lottery ticket hypothesis'.
I take it you have not seen the recent Dall-E 2 results? Clearly that model is not just working on a narrow space.
See https://openai.com/dall-e-2/ and the many awe-inspiring examples on Twitter
But, if you look at your smartphone, virtually every popular application the average person uses--Gmail, Uber, Instagram, TikTok, Siri/Google Assistant, Netflix, your camera, and more--all owe huge pieces of their functionality to ML that's only become feasible in the last decade because of the research you're referencing.
The way people hype AGI/AI/ML whatever undervalues the actual effort behind these remarkable feat. There is so much effort being made to make this work. Deep learning works when it is engineered properly. So it is just another tool in the toolbox!
Look at how graphics community is approaching deep learning. They already had sampling methods but with MLPs (NeRFs), they are using it as glorified database. So it's engineering!
I want to underscore that AI/ML/DL research requires ground breaking innovation not only in algorithms but also in hardware and software engineering.
Machine learning / neural nets also (like I said) get to claim credit for a hell of a lot of things which are just products of colossal advances in hardware – simply of its becoming possible to run statistical methods over very very large '1:1 scale' sample sets – and not due to a specific statistical technique (NN) which is not remotely new and has been heavily researched for about 40-50 years now.
For that, we need artificial comprehension, which we do not. Artificial comprehension, the ability to generalize systems to their base components and then virtually operate those base concepts to define what is possible, to virtual recreate physical working system, virtually improve them, and with those improvements being physically realizable is probably what will finally create AGI. We need a Calculus for pure ideas, not just numbers.
PhD -> low-level dev -> FAANG mid-level is nothing to scoff at so you're doing pretty well.
Competing within a giant company for perf ratings feels like school and I'm over it. But the other parts of the job are great.
It's a startup... titles in a 50 people organization don't compare to 50,000 people organization titles.
I'm sure you can go and be a VP at a startup too, if that's what you want to do. Just go and network at Incubator, Investor, & Entrepreneur events/meetups/organizations, and come up with an idea & customers, then execute and try to get customers on board... rinse and repeat.
As you identified, location is the next big factor. If you are still in Europe, my advice is to leave or to start your own company there. If you are working for primarily US based companies in Europe there will always be a limit to the level of exposure you get to leadership and to how fast you can rise up the hierarchy.
Finally, don't discount Eric's profile. Through some combination of his public profile and professional work, he's established a reputation and following. That is just as important as any hard engineering work in securing a senior/leadership role.
I'm glad someone validates my belief that "location matters". I moved to the US 2 months ago after 2 years in Covid VISA limbo working remotely for a US team. Settling down has been extremely painful so far but I hope it's worth the effort in the long run.
Unfortunately the tech world is not really a meritocracy.
When I look at my own circle of technical people the most incredible ones from a pure technical ability are divided between working at FAANG making 500k+ and working a relatively unknown companies making ~200K or less. One of the most mindbendingly brilliant people I know is working in relative obscurity, known very well only among other people that are top in the field, but their resume looks very ordinary compared to their behind the scenes contributions to major projects.
Managing a career in tech is largely independent from technical skills and abilities. I have met a shocking number of people making lots of money at prestigious institutions that are "meh" as far as technical ability goes (of course there's some great ones as well), and have met plenty of brilliant people working relative obscurity.
The success is largely a function of both background (Brown does beat a "no name EU university") and personal desire to have a prestigious career. There is a lot of self promotion going on in this piece, in fact the OP has already convinced you that they might be just a wildly better person than you. If they can convince you they are this amazing, then they also can convince the leadership team at a start up. But do recognize that their skill demonstrated so far is only in convincing you of this.
We all have different trajectories and choices. This comment makes it seem like if you aren't a technical wizard then you might as well be useless. This is not reality
This is the answer. I have grown more in ~7 years* of random SFBA startups than I did in the previous 13 years of career in Europe. Just because the kind of startup that's a dime a dozen over here is a once in a lifetime opportunity back home.
To put this contrast into numbers: In 2021, during the pandemic while "SFBA is dying" was the mem, the Bay Area raised as much startup investment as all of Europe.
*I wasn't as career aggressive as I could've been, mostly for visa-related reasons.
1) The PhD takes a huge hit on your opportunity cost.
2008-2014 is 6 years of time; for me, it was the delta between starting my career as a junior engineer and becoming a tech lead at a hot unicorn which let me pivot to a CTO role at a small startup.
2) Academic credentialism has real effects.
This guy did a CS degree at an Ivy in the US. He has been set up for commercial success in the US tech industry through a halo effect you cannot also access unless you gained access to that institutional grooming at the same age. By choosing to do that PHD in EU (and a no name one at that), you forfeited that access.
In my experience, while the effect of this goes down over time, it has extremely strong launch + early compounding effects.
3) Risk tolerance can work to your benefit or against it.
You are working at a FAANG which is the safest and most cash lucrative option. In all likelihood, you have a great WLB and now a great blue chip brand on your resumé. However, the cost of this is that you're generally not going to get access to projects or culture that, by virtue of your participation, set you on an extremely steep growth path.
To get access to that, IMO, there's no real alternative to achieving strong outcomes working at a startup. Of course, that can be hard to do -- how do you figure out which ones are future winners, and how do you get them to let you come on board? I have no great answer rather than early career trial and error (accepting some of it will work out poorly and uncomfortably so).
I wouldn't say that "OP is a beast" per se, but it's much more likely that they have been groomed (working in the right conditions) in ways that you may not have. And yes, startups titles are not comparable to big company titles. It's apples and oranges.
The company he joined is a Series A startup, so absolutely an early stage company where whether you're VP/CXO, you're functionally going to be doing a player/coach role at most with tons of strategy baked in. But I wouldn't call that inflation, per sé. Sure, it's not the equivalent of being an experienced people leader and executive at a big corporation manning a giant organization at its helm. But you are often times in charge with significantly more responsibility and do not have bureaucratic friction and slow pace to hide behind. Doing a startup is just different. It's insanely risky, overall has poor risk adjusted rewards, and often is a magnet for shady characters. But if you can filter out the wheat from the chaff, you get access to the best career opportunities available, bar none.
Right place, right time + talent + willingness to take risk.
I’d take some things here with a grain of salt, like “low 7 figures compensation (staff level)” at FAAN (can eliminate G because they’re not likely to hire him back at L+1 immediately after he quits). ML is still somewhat hot, but 7 figures is an outlier for staff-level comp.
AAPL: ~$450k https://www.levels.fyi/company/Apple/salaries/Software-Engin...
AMZN: ~$600k https://www.levels.fyi/company/Amazon/salaries/Software-Engi...
FB: ~$600k https://www.levels.fyi/company/Facebook/salaries/Software-En...
Nobody there is reporting $1M+ offers for staff level. While I’m sure it’s happened, it’s pretty far outside the staff pay band (excluding equity gains during the 2020-2021 market run-up, which are sadly behind us) and would be a truly exceptional offer even in the current climate. That, plus the fact that it sounds like he didn’t get many formal offers (“I did not initiate the formal HR interview process with most of them”) and wasn’t pitting offers against each other, makes me skeptical.
Basically, yeah, small/not-yet-massive startups have insane overinflation in titles. Had plenty of former college classmates who became "VPs" or "staff engineers" at super small startups a couple years out of college. Getting plenty of recruiter messages on linkedin myself for "staff engineer" positions at random startups, despite me not even being a senior at a FAANG yet, and only being about 4.5 years out of college.
Another thing is, no matter how smart or hard working you are, being in the right place at the right time is extremely important. It won't help much if you lack skills, but being in the right place at the right time is like a force multiplier on your skills and the work you do. Which is partially why most of the big opportunities are still heavily concentrated in a few geographic spots (despite there being no tangible technical need for that).
Don't beat yourself up over it, titles don't mean that much. You are able to start a one-man-shop LLC and call yourself a VP, a director, or whatever else you want. The real question is, with that title, are they being compensated as much as you are? If they decide to quit and get a job at a "regular" tech company after, will that VP title translate into anything more than an L4/L5? Just some food for thought.
It's clear from the blog post that the author is in the same boat. They lament CEOs not having time to do research but took a VP position. An actual VP doesn't have time do research so they're clearly not an actual VP. So they're likely a tech lead with an inflated title.
and.. of course startups have title inflation ;)