149 karma · joined May 10, 2013
There’s no way for me to know or even check what are the possible errors this function can return? Sure, sometimes a comment in the library might be illuminating, but sometimes not.
I agree that errors as values that I can handle at the call site rarely feel useful.
Some of the Is, As ergonomics have improved, but damn if coding agents don’t love to just fmt.error any code it writes. Thus hiding the useful information about the error in a string.
At my new company, the use of stored procedures unchecked has really hurt part of the companies ability to build new features so I'm surprised to see what seems like sound advice, "don't use stored procedures", called out as a cargo cult.
Yes! Gibson recognized they were getting cut out of the vintage market and started making not only the reissues (RIs), but also the limited edition copy-of-famous-person's-guitar. What gets me is that Reissues these days are priced so close to vintage instruments. It's so hard to justify the purchase.
It keeps that part of my brain occupied, but not focus, while I work on the task at hand.
This video from a year ago goes into more engineering depth than promotional videos: https://www.youtube.com/watch?v=S5t6aYhj6pU
the non-obvious challenges of working in this space are that you have to deal with real world constraints that you just don't have to think about in 100% cloud based software. hardware goes down! internet connections go out! electricity goes out! how much processing can you do in store vs out of store?
that's just the tech in the store. what about humans? humans are now in your programming state moving shit around. in-store associates miss-stock items. kids do kid shit. people don't place things back exactly where they found them.
Then you have interesting distributed problems: how do you handle late data? what should be a massively parallel problem is really a graph of interactions that have to be resolved in just the right order so you can generate an accurate receipt.
and you know what's crazy? the vast majority of the time it works! exactly how they say it does: with machine learning models.
bonkers.
That being said a company as big as Amazon will always have edge cases.
Plugins - software for audio programs - are available but audio engineers are famously persnickety.
F major is red. E-minor is a desert sand color.
Also, checkout typing.io which has lessons based on open source code.
UPS drivers pee in bottles while out delivering, too.
In another life I worked facilities at a UPS distribution center and my team would have to pick up those bottles :-/
When yah gotta go, yah gotta go.
I use this mapping in my current setup and find it really useful.
I use a modified version of the default symbol layer that's more programmer friendly (I think) and I put the toggle for this layer under my right index finger. I feel like I'm more accurate now because it's not just my right pinky doing all the symbol work.
https://configure.ergodox-ez.com/ergodox-ez/layouts/aN90b/la...
That may be obvious for others, but it wasn't for me.
Should take about an hour with testing to get a Pyspark script together to read in a DynamoDB table and write it out to S3.
https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programm...
You’ll then need to crawl the S3 data to add it to your Glue catalog and then you can query it with Athena.
There are no indexes in the traditional MySQL / Postgres sense of the word. You can, however, layout the data to make your querying more efficient. See: https://docs.aws.amazon.com/athena/latest/ug/partitions.html
However, I read a harrowing / awe-inspring blog post about someone doing just that. So...¯\_(ツ)_/¯
Now I've learned to read the Best Practices section of any AWS service before I start implementing. It saves a lot of heartache
* The default timeout is ~ 48 hours and you pay per Data Processing Unit (DPU) that you've provisioned the Job. * Currently it supports Python and Scala. As far as I'm aware you can't run Java jobs directly, but you can upload JAR libraries and use them in your code.
Re serverless vs dedicated / ephemeral clusters: Like with any serverless runtime environment you are trading convenience (across a few dimensions) for flexibility.
The Glue environment runs in a few limited runtimes and uses a specific version of Spark that you have no control over updating. Given that it's pretty quick to author a job, you can set the required DPU and Glue handles that, and you don't have to worry about sizing the cluster for your data size. For me most of my jobs fit within those constraints.
At some point on the cost curve it may make sense for you to move all of your jobs from Glue into a dedicated cluster on EMR. You may also get there sooner if you need to use specific frameworks or libraries.
Thumbs up on the Glue Dev endpoint. It's been killer. I had trouble setting up a Notebook (I wanted to get fancy with Docker) and I usually use the Python repl link that's provided.
I'm working on a follow up post that removes the Data Pipeline -> Lambda and uses the new Glue DynamoDB integration.
Looking forward to your blog post!
As I mentioned in another comment I, personally, wouldn't advocate for new teams or startups to use a service oriented architecture right out the gate. It's too easy to end up with a distributed monolith (circular dependencies between services, services uptimes that are tightly coupled). Engineers also tend to underestimate the build tooling, observability excellence and discipline you need to make it seamless.