755 karma · joined June 21, 2018
Before: https://github.com/savarin/ledger/blob/67d6e236296e4787e8924...
After: https://github.com/savarin/ledger/blob/506584542d1c1c2751e2d...
https://ezzeriesa.notion.site/Writing-clean-code-with-ChatGP...
This is detailed in the Python code. https://github.com/savarin/bitpacker/blob/239d68dcd3ec5db67e...
Yes using the back row works too! I was trying to see if we can get an improvement on the 18 additional bits needed from the post (or 14 additional bits by taking advantage of knight and bishop ordering).
These discussions have been great, very much enjoying seeing the incremental improvements!
Edit: I did another pass, item 4 in the notes did mention this.
> [4] For en passant we need the pawn to remain on its home file. Hence we exclude the pawn from this step if it can be captured en passant. Captures can appear on any file.
If we can guarantee that the pawn stays in its own file then I see a path for improvement (by having the actual position vs back row usage as a 1-bit toggle), but this is not broadly the case due to movement across files on pawn captures.
Let me ponder a bit more on the pawn locations. We do need the pawn ordering since the we've encoded the string representing the promotions in sorted order. In the case of en passant we know exactly where the pawn would be, and while we can use the back row, I haven't quite figured out how to use this to encode reliably.
We did start integrating dbt towards the end of my time in the role. Our data stack was built in 2018, so a fair bit of time before data infra-as-a-service became a thing. The idea is dbt would help our internal consumers to more easily self serve. That said I did see complaints about dbt pricing recently; as they say there’s no free lunch.
Re: ORMs, I respectfully disagree. I’ve come across many teams that treat their Python/Rust/Go codebase with ownership and craft, I have not seen the same be said about SQL queries. It’s almost like a 'tragedy of the commons’ problem - columns keep getting added, logic gets patched, more CTEs to abstract things out but in the end adds to the obfuscation.
ORMs don’t fix everything but it does help constraint the ‘degrees of freedom’ and help keeps logic repeatable and consistent, and generally better than writing your own string-manipulation functions. An idea I had I continued (I wrote the post early last year) was to use static analysis tools like Meta’s UPM to allow refactoring of tables / DAGs (keep interfaces the same but ‘flatter’ DAGs, less duplicate transforms).
Interestingly enough, I currently work on ML and impressed to see how much modeling can be done in the cloud compared to my earlier stint in the space (which had a dedicated engineering team focused on features and inference). On the flipside I similarly see an explosion of SQL strings, some parts handled with care more than others.
I’ve not looked into a data mesh but a friend did mention pushing his org to embrace it - self note to follow up to see how that's going. Looks like there are a couple of ‘dimensions’ to it; my broader take is that keeping things sensible is both a technical and organizational challenge.
I look forward to future blog posts on ‘how we refactored our SQL queries’, maybe there’s a startup idea there somewhere.
Could you share more specific details? Happy to look over / revise where needed.
More broadly is the issue of the gap of what you think the role is, and what the role actually is when you join. There are definitely cases where this is accidental. The best way I can think of to close the gap is to maybe do a short-term contract, but may be challenging to do under time constraints etc.
On the former, we had custom tooling at the time of writing but now use GCP Dataflow (which I can say I'm rather partial to).
The table was ~250mm rows per day. The alternative would be doing it in memory; Python unit tests do feel a bit more natural but for a table that size we decided it was best to do it on disk / let the DBMS deal with it. Perhaps Spark is an idea but then there's the trade-off of customers losing context.
It also helped us 'refactor with confidence' - we can happily change the query to incorporate new use cases while knowing the core logic is still sound.
Love it.
https://tailscale.com/blog/database-for-2022