195 karma · joined November 29, 2019
A couple of services looked like they had a memory leak. Memory was continuously increasing over time. Thanks to Python 3.14, we were able to use memray to understand what was going on. Those services were recreating HTTP clients (aiohttp) for every inbound request, and memory allocated by the downstream SSL lib was growing faster than it was being released.
We ended up rolling back to 3.13, which fixed the issue. I'll try again with 3.14.5.
most tasks I can do better and faster with composer 2
a fellow engineer reported a bug on a code I had written a few months back. I used his report as prompt for composer 2, gpt-5.4-high and claude-4.6-opus-max-thinking. composer found the issue spot on. gpt found another possible vector a couple of minutes later, but a way less likely one and one that would eventually self heal (thus not actually reproducing what we observed on production). claude had barely started when the other two had finished
also, i don't have a budget per se. but it is expected that i over deliver if i'm over spending
I've been trying to use Claude Code seriously for over a month, but every time I do it, I get the impression that it would take me less work to do with Cursor.
I'm on the enterprise plan, so it can get pricey. This is why I used to stick mostly to auto mode.
Now Composer 2 has taken over as my default model. It is not as intelligent as OpenAI's or Anthropic's flagship models, but I feel it has as good as or better intuition. With way better pricing. It can get stuck in more complex tasks though.
Being able to get in the loop, stop and instruct or change models makes all the difference. And that is why I've stayed in the editor mode until now. Let's see if 3.0 changes that.
But I think that is beside the point.
Individuals are not fungible, but team members are - or at least can be, depending on how you structure your teams.
And as your org grows, you want predictability on a team level. Skipping a bunch of reasoning steps, this means having somewhat fungible team members, to give you redundancy.
The engineering parallel here is the tradeoff between resilience and efficiency. You can make a system more reliable by adding redundancy. You make a system more efficient by removing redundancy.
The author also mentions a leader adjudicator, which means there is probably some sort of coordination to pick a leader. This raises the question of how a leader is picked, and if leadership changes based on how hot a key is in a given AZ.
This blog series is a great read. Every day Marc drops excellent content and leaves room for questions, which he ends up answering on the following days. Hope more details come next.
> We’ve learned from building and operating large-scale systems for nearly two decades that coordination and locking get in the way of scalability, latency, and reliability for systems of all sizes. In fact, avoiding unnecessary coordination is the fundamental enabler for scaling in distributed systems
I hope there is a follow-up since the points the author only glossed over are important to understanding the architecture and trade-offs. I would like to know about the cross-adjudicator coordination protocol and how the journal works.
From the information available, it seems that DSQL should be pretty fast as long as you keep writes local. Once you add active-active replication and start writing to the same key in different regions the coordination costs should slow the system down significantly (or not - but if that is the case I want to know how they managed to do it).
- How the OS knows it can clean up an inode after a hard link is deleted? The post mentioned inodes don't see hard links
- What does it mean to have a dead/dangling soft link?
I somehow started to find they kind of beautiful when I worked at a company that only had pint-sized American glasses at their office. Now most cups in my house have this design. They are dirty cheap and very easy to replace.
- Soft Skills Engineering: very humorous and light take on the people side of SWE
- StaffEng: interviews with engineers with staff+ roles
- Lenny's: if you are into product, growth and startups in general
It was already mentioned by a sibling comment, but podcasts from the Changelog are also very good.
I was expecting a larger grace period from announcement to price change though. You got it right when you started charging for stopped volumes - you overcommunicated and gave users estimates for a few months before you started charging. Smoothing the increase over the next 4 months was nice.
A couple of nits as well: - AFAIK prices increased in all but two regions. The announcement could be clearer or more direct regarding this. - The pricing page could have a clearer from/to. I have to click on the example toggle to know how much machines cost before the change.
Making sure interviewers show genuine interest, and are open to different candidate background is one of the most difficult things to guarantee when you have many people interviewing.
> guiding technical direction, design and implementation of Rust component libraries, SDKs, and contributing to the building of global scale services in Rust.
Maybe the article misquoted it. Maybe the posting was updated after the article. The only time C# shows up is on the required qualifications section.
I suggest taking a look at the replica settings (https://litestream.io/reference/config/#replica-settings) and making sure the default values work for you.
Using both the default retention (24h) and snapshot (24h) intervals will not give you 24h of data history. This happens because any snapshots older than 24h will be deleted. Granted, you'll have another snapshot ready, but it will only contain a short history of your DBs changes.
Also, test restore times. If you need to restore in a hurry, having to download 24h of wal logs can take a while.
I started using it during a time when it was unmaintained (between v0.3.9 and v0.3.10). There were a few bugs I hit, e.g. memory usage spikes. Shout out to hifi, who created patches and was active at the Slack community. I relied on his fork for a while, nice to see it integrated on main.
Probably because of this.
> but it does mean that the wal file may grow indefinitely if the checkpointer never gets a chance to finish without a writer appending to the wal file. There are also circumstances in which long-running readers may prevent a checkpointer from checkpointing the entire wal file - also causing the wal file to grow indefinitely in a busy system.
> Wal2 mode does not have this problem. In wal2 mode, wal files do not grow indefinitely even if the checkpointer never has a chance to finish uninterrupted.
I don't get how wal2 fixes the long-running reader problem though. Maybe they were just referring to the former problem?
Most of the situations where we needed to drastically scale up were known ahead of time as well (e.g. campaign from customer), and we would preallocate instances or even more clusters.
I may be forcing my memory, but if I'm not mistaken, our auto scaling was setup in a way that the system could handle sudden load increases of ~50% without noticeable disruption. Spikes bigger than this could lead to increased latency and/or error rate.