501 karma · joined May 14, 2017
The only thing that worries me with SQL is when having to write UDFs for, say, computing a Z-score. But maybe it's just because I have never done it? Do you have any good resources about this?
Now that we have buy in the pressure isn't so bad. I have a startup background and this is one of the reasons why I was hired. I am not too worried about having to convince my manager, she's great and is the one who started criticizing the legacy (she arrived a few months ago) and asked me to dive in and give my opinion.
She agreed to let me work 30% of my time on this, and the rest on delivering direct value to the clients. The related tickets will be part of the sprint. I will do as you say now, and create items that address the concerns.
> If it is only used by data science team then you can rip apart the solution and work towards a more logical infrastructure where all you are doing is cleansing, normalizing and deriving features and these become your central feature repository which your team can pull and build models.
It is only used by data scientists. What do you mean by a feature repository? How would you organize it so people can push new features? This sounds very interesting.
> you are talking about denormalization as something you are planning to do, that should have been the starting point of any HDFS/SPARK/Parquet based solution.
It is something that we have to do, but the table have been dumped as is in S3 and every project rebuilds the whole derived dataset regularly. Since these operations are very brittle (a lot of manual work and even transformations performed in notebooks), this is something people dread doing. I am trying to secure this at the moment by writing Makefiles that remove human intervention, but at the end of the day I would like to avoid people spend hours waiting for new data when they need it.
> I can suggest tools for explorations, data quality check etc.
I would appreciate it. Put simply, we get data about the evolution of the stock of clients, transactions with their clients, product descriptions, etc. that is dumped into S3 (I scheduled a chat with people upstream to see what happens). We have 3,4 projects for each client. What currently happens is every team writes the same code to build features in their separate repositories, this code is re-executed every time new data arrives (weekly). These features are then used in prediction models.
Besides the brittleness of the process, I found that people are reluctant to analyse the data because it takes an unreasonable amount of time.
I talk about adding infrastructure in my original post, but I'm very well aware that my time is currently better spent consolidating the existing as much as I can so the clients can get correct results faster.
Do you enforce consistent code formatting in your codebase (using CI, maybe)? If so, why wouldn’t you do that with commits?
A good git log prevents unnecessary headaches and saves a lot of time. Typical situation: you just found a couple of lines whose purpose is not completely obvious. If the author is still there, you both waste time discussing. If she/he left the company and you have to figure it out by yourself. Now imagine using git log to identify when these lines where introduced/changed and getting an explanation as to why.
So yeah, writing good commit messages can feel like a waste of time. Like writing good code can feel like a waste of time. You should not be doing it to please your colleague, but rather to contribute to a maintainable codebase. It’s worth the extra minutes here and there.
Bonus point: less documentation to write.
- First, see if you can find other employees in the same situation. Being a group helped us. - Document everything. Commit on git first thing in the morning, and before leaving. Write down every conversation you had right after you had them—we tend to forget very quickly. Dump your calendar and emails regularly. You can find scripts to dump Slack too. If your boss has written anything abusive them, it is also a good idea to take screenshots. - I guess moving is not an option for your family? Otherwise start looking in other markets. - Freelancing is a good option otherwise. I know it sucks to do this on top of your job, but you only to save enough to get yourself out of this.
Once you’re out, and only when you’re out, find a good lawyer. Do not hesitate to sue. It’s not greed, it’s reparation for the abuse. You deserve it.
Good luck to you.
I developed this when I was in academia and it transposed very well (better?) in the software industry. If you do that regularly it also helps course-correct when needed.
Sure, you can bootstrap something really quickly in Python and Nim, but when it comes to creating beefy APIs, I find that Go offers a much shorter development cycle. Sure it is slightly more verbose, but when your code compiles it usually works (which is partly due to the way Go deals with errors). If you need to handle a large number of requests, scaling with Go is a breeze and a nightmare in Python.
I have a background in data science, so have a lot of experience with Python. I’d always wanted to do more backend stuff: push my models in productions, do the ETL work myself... But quickly got frustrated with Python.
Go and Erlang have solved a lot of my pain points. Erlang is a bit of a strange beast for people who come from procedural languages, but it allowed me to build ETL pipelines that are more robust than anything I could do with Python in less than half the number of lines. Even for heavy computations, I now find myself spinning Python nodes to do the heavy lifting with Pytorch and co, and doing all the orchestration with Erlang. It also made me a way better programmer.
I use Go for computation-heavy APIs that don’t require a lot of data massaging. It’s a very efficient, weirdly comforting (!) language.
Now I think I got the bug, I really want to learn Clojure and Nim this year :)
Cons:
- You’ll have to learn about a lot of different people at the same time. In early stages startup you get to meet people one after another and earn their trust one after another. This is different.
- You probably have a different vision than that of the previous CEO, and you’ll have to go the extra mile to align everyone.
- You inherit a culture you didn’t shape. You’ll probably have to change that, and it will take time for people to adjust. Just lead by example. - All of this takes time, and meanwhile the company needs to survive :)
Pros:
- You’ll ramp up your people skills like never before. Gaining the respect of an entire team as a CEO, and (sometimes) repairing damage will take a lot of courage and empathy.
- Obviously, if you turn this around you’ll get major valley cred. And, never underestimate this, a huge satisfaction.
If you decide to go that way, good luck, and you already have my respect!
`alias did="vim +'normal Go' +'r!date' ~/did.txt"`
Then type `did` in your terminal and write what you just did.
More generally, you can get a great boost of productivity by mastering your text editor. I have a preference for Vim, but have nothing against emacs.
What disturbs me the most are colleagues who said it was my fault since I hadn’t shared my grief on social media, and algorithms could not learn (I work in data « science »). As if targeted advertising was a law of Nature and there was nothing we could do about it. Just deal with it. And blame yourself for not showcasing your pain in front of hundreds of people.
The issue with algorithms is not that these events are statistically unlikely (unfortunately). They just have memory, enough memory to be accurate. So the day happy messages are replaced with the word « death », it hasn’t had the time to learn baby was not there anymore. There may be ways to improve that, but I bet you’d be trading accuracy for the remaining 95% people. Guess what version wins at a business meeting?
I’m not blaming colleagues who write these algorithms. They didn’t know this could happen, they didn’t think about it before this happened to me. We’re all math and CS geeks priding ourselves in [insert your metric here] score above the 90% mark, and that’s all we optimize for. I’ve heard a joke once that the difference between science majors and humanities major was that scientists build the atomic bomb, humanities major explain to them why it’s not a good idea to use it. Targeted advertising is not as bad as the atomic bomb, of course, but you get the idea.