HNHacker News
TopNewBestAskShowJobs

pavel_pt

24 karma · joined January 27, 2011

submissionscomments
pavel_pt··on Today is when the Amazon brain drain sent AWS down the spout
Business hours for the team receiving the alarm; many issues can wait to be resolved during your own waking hours if they are not impacting customers.
pavel_pt··on Today is when the Amazon brain drain sent AWS down the spout
I worked at AWS (EC2 specifically), and the comment is accurate.

Engineers own their alarms, which they set up themselves during working hours. An engineer on call carries a "pager" for a given system they own as part of a small team. If your own alert rules get tripped, you will be automatically paged regardless of time of day. There are a variety of mechanisms to prioritize and delay issues until business hours, and suppress alarms based on various conditions - e.g. the health of your own dependencies.

End user tickets can not page engineers but fellow internal teams can. Generally escalation and paging additional help in the event that one can not handle the situation is encouraged and many tenured/senior engineers are very keen to help, even at weird hours.

pavel_pt··on The Heart of Innovation: Why Most Startups Fail
Fantastically insightful; surprised this didn't gain more traction here. Curious if the gook contains deeper insights still.
pavel_pt··on Amazon Ends Support for Chime
I think things like the Slack conference call integration depend on this. In fact that’s the only use outside of the official client I’m aware of.
pavel_pt··on Initial details about why CrowdStrike's CSAgent.sys crashed
If that’s all it takes an attacker, you’re doing AWS wrong.
pavel_pt··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
I hope @sewen will expand on this but from the blog post he wrote to announce Restate to the world back in August '23:

> Stateful Functions (in Apache Flink): Our thoughts started a while back, and our early experiments created StateFun. These thoughts and ideas then grew to be much much more now, resulting in Restate. Of course, you can still recognize some of the StateFun roots in Restate.

The full post is at: https://restate.dev/blog/why-we-built-restate/

pavel_pt··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
Restate also stores a deployment version along with other invocation metadata. FaaS platforms like AWS Lambda make it very easy to retain old versions of your code, and Restate will complete a started invocation with the handlers that it started with. This way, you can "drain" older executions while new incoming requests are routed to the latest version.

You still have to ensure that all versions of handler code that may potentially be activated are fully compatible with all persisted state they may be expected to access, but that's not much different from handling rolling deployments in a large system.

pavel_pt··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
Disclaimer: I work on Restate together with @p10jkle.

You can absolutely do something similar with a RDBMS.

I tend to think of building services in state machines: every important step is tracked somewhere safe, and causes a state transition through the state machine. If doing this by hand, you would reach out to a DBMS and explicitly checkpoint your state whenever something important happens.

To achieve idempotency, you'd end up peppering your code with prepare-commit type steps where you first read the stored state and decide, at each logical step, whether you're resuming a prior partial execution or starting fresh. This gets old very quickly and so most code ends up relying on maybe a single idempotency check at the start, and caller retries. You would also need an external task queue or a sweeper of some sort to pick up and redrive partially-completed executions.

The beauty of a complete purpose-built system like Restate is that it gives you a durable journal service that's designed for the task of tracking executions, and also provides you with an SDK that makes it very easy to achieve the "chain of idempotent blocks" effect without hand-rolling a giant state machine yourself.

You don't have to use Restate to persist data, though you can - and you get the benefit of having the state changes automatically commit with the same isolation properties as part of the journaling process. But you could easily orchestrate writes into external stores such as RDBMS, K-V, queues with the same guaranteed-progress semantics as the rest of your Restate service. Its execution semantics make this easier and more pleasant as you get retries out of the box.

Finally, it's worth mentioning that we expose a PostgreSQL protocol-compatible SQL query endpoint. This allows you to query any state you do choose to store in Restate alongside service metadata, i.e. reflect on active invocations.

pavel_pt··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
We don't have specific plans for our next SDK to build, but Python definitely comes up often - thank you for the input!
pavel_pt··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
Appreciate the feedback! What kind of support do you wish for, if there was one thing you would prioritize?
pavel_pt··on Code that sleeps for a month: Solving durable execution's immutability problem
I think you do it just like you handle compatibility between services – you never remove parameters; you only ever add new optional ones if you have to. This way a message from the past will be compatible with a future handler, same as if you have a caller that depends on you which is using an outdated client/API definition.

But you are right; it's very hard to reason about testing such systems, since you may have accumulated state which causes your handler logic to behave differently. The problem exists in service architectures in general though, it's just very hard to miss with intentionally delayed processing.

pavel_pt··on Solving durable execution's immutability problem
How do you imagine detection of conflicts working?
pavel_pt··on Solving durable execution's immutability problem
+1 on views, for those things than need direct DB access.

One effective pattern I’ve seen in large DB deployments is to separate the write schema from the read schema. That is, treat what is allowed to be written into the DB separately from the shapes of any views that exist. The views themselves are a tightly coupled client of the DB log - by constraining writers, you can migrate/rebuild views, then point services to read from the migrated views, and retire old views.

This allows you to keep accepting writes - you never have to shut down the write path. If you’re introducing new shapes or the DB, you’d prepare a new view, the “widen” the write schema, and begin accepting writes in the new shape, and only then re-point clients to read from the new view.

To drop elements of your read schema, you do the dance in reverse. First, constrain your writes. Then, build new views that don’t require the elements you removed. Gradually, update application code to work on the new, reduced views. When you’re done reading from the original views, you can drop them.

This is inherently much less efficient than online or offline DB migrations. But it’s a sane strategy for wrangling very large systems with very low risk.

Versioned workflows are in practice distributed entities that interact with their peers, themselves defined by versioned interfaces. By tracking the versions of handlers touched by any given execution, we can imagine a similar experience for deployed code - including garbage-collecting unused versions.

pavel_pt··on Solving durable execution's immutability problem
Are there any aspects in particular that would be of interest?
pavel_pt··on Solving durable execution's immutability problem
Not just in durable execution, but in control systems in general! Anything that operates on persisted state in a distributed environment has similar issues. Durable execution frameworks make them more apparent - and may also help us chart a path for how to deal with updates more rigorously.
pavel_pt··on AWS CloudWatch Events
For customers with big infrastructure and/or a high rate of modifying resources, the ability to be notified of events occurring within AWS instead of than polling various describe APIs should be a huge win.
pavel_pt··on RAML – RESTful API modeling language
Ditto, used RAML on a previous project and everyone involved really enjoyed it. It manages to capture just enough about how thinks work without getting overbearing. I particularly like the fact that it allows for examples to be specified.

I haven't built anything with Swagger but I never clicked with it the way I instantly did with RAML. It's a pity - there seems to be a lot more industry and open source support behind Swagger than there is for RAML which is mostly backed by MuleSoft.

pavel_pt··on RAML – RESTful API modeling language
The good parts of WSDL, I presume? ;-) I spent way too long staring at raw WSDL, I almost think certain profiles are actually OK. But SOAP on the whole is a world of pain.
pavel_pt··on Why should you hire a polyglot programmer?
One should hire people who care, and who continuously broaden their own horizons. When it comes to developers, these inevitably happen to also be polyglot programmers in my experience.