HNHacker News
TopNewBestAskShowJobs

p10jkle

393 karma · joined October 23, 2017

submissionscomments
p10jkle··on How London cracked mobile phone coverage on the Underground
There already is signal on many underground lines, and it’s pretty rare that people are playing things out loud in my experience?
p10jkle··on Trump says Venezuela’s Maduro captured after strikes
https://en.wikipedia.org/wiki/Sovereign

> The roles of a sovereign vary from monarch, ruler or head of state to head of municipal government or head of a chivalric order. As a result, the word sovereignty has more recently also come to mean independence or autonomy.

p10jkle··on Ofcom fines 4chan £20K and counting for violating UK's Online Safety Act
London is the third most visited city and Heathrow is the second most popular airport by international visitors. The prospect of being arrested upon arrival there might be a little annoying.
p10jkle··on Building a modern durable execution engine from first principles
Hey, I work on Restate. There are lots of differences throughout the architecture and the developer experience, but the one most relevant to this article is that Restate is itself a self-contained distributed stream-processing engine, which it uses to offer extremely low latency durable execution with strong consistency across AZs/regions. Other products tend to layer on top of other stores, which will inherit the good things and the bad things about those stores when it comes to throughput/latency/multi-region/consistency.

We are putting a lot of work into high throughput, low latency, distributed use cases, hence some of the decisions in this article. We felt that this necessitated a new database.

p10jkle··on Show HN: Factorio Learning Environment – Agents Build Factories
Wow, fascinating. I wonder if in a few years every in-game opponent will just be an LLM with access to a game-controlling API like the one you've created.

Did you find there are particular types of tasks that the models struggle with? Or does difficulty mostly just scale with the number of items they need to place?

p10jkle··on The Anatomy of a Durable Execution Stack from First Principles
As discussed in the article, we have built our own storage engine from the ground up, which we did because we believe it will achieve better performance by taking advantage of the features of the system (streaming data, single writer etc) instead of shoehorning it into a DBMS. So, our performance goals are very high throughput (100s of thousands of actions per second, scaling horizontally), with very low latencies (like, 40ms p90 under load for a 3 step workflow)
p10jkle··on Every System is a Log: Avoiding coordination in distributed applications
I feel this happening to me too... depressing
p10jkle··on Why don't you move abroad?
The twitter office is on Market Street, so this is totally plausible
p10jkle··on Waverley, the last seagoing paddle steamer
I saw this going under Tower Bridge last week! Made me late for the dentist... :)
p10jkle··on Distributed transactions in Go: Read before you try
This is also a good use case for durable execution, see eg https://restate.dev
p10jkle··on Infinite checkboxes in 71 lines of backend code
Full source here: https://github.com/jackkleeman/restate-checkbox/tree/main

But the fun bit is the virtual object: https://github.com/jackkleeman/restate-checkbox/blob/main/ob...

p10jkle··on Launch HN: Hatchet (YC W24) – Open-source task queue, now with a cloud version
Maybe let them have their launch? Mitchell said it best:

https://x.com/mitchellh/status/1759626842817069290?s=46&t=57...

p10jkle··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
We haven’t built any client side encryption tools yet. I don’t think it would be particularly difficult to do an MVP. If it’s very important to your use case, come chat to us in Discord? https://discord.com/invite/skW3AZ6uGd
p10jkle··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
Ah I see what you mean. In this case the handler should complete with a terminal error - we weren't able to finish the task in time. Of course, many types of errors and timeouts are valid application-level results, not transient infrastructure issues. And sadly, tight timeouts push transient issues into application-level issues, and this is unavoidable, I think
p10jkle··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
Thanks! I'm not familiar with Jobrunr, but we can definitely help with orchestrating async tasks (as well as sync rpc calls), especially if its important that they run to completion
p10jkle··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
Hey! I managed to get a POC running on Cloudflare workers, I had to make some small changes to the SDK eg to remove the http2 import, convert the Cloudflare request type into the Lambda request type, and add some methods to the Buffer type. I suspect similar things would be needed on Deno platforms. We have it on our todo list (scheduled within weeks not months) to make it possible to import a version of the library that just works out of the box on these platforms. I think if we had someone with a use case asking for it, we would happily build that even sooner - maybe come chat in our discord? https://discord.gg/skW3AZ6uGd

Once http2 stuff is removed, there's nothing particularly odd that our library does that shouldn't work in all platforms, but I'm sure there will be some papercuts until we are actively testing against these targets

p10jkle··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
Yeah, this is pretty much exactly how we propose its done (restate services are inherently versioned, you can register new code as a new version and old invocations will go to the old version).

The only caveat being that we generally recommend that you keep it to just a few minutes, and use delayed calls and our state primitives to have effects that span longer than that. Eg, to poll repeatedly a handler can delayed-call itself over and over, and to wait for a human, we have awakeables (https://docs.restate.dev/develop/ts/awakeables/)

More discussion: https://restate.dev/blog/code-that-sleeps-for-a-month/

p10jkle··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
Yeah, definitely. We would like to have modes of operation where Restate puts its state only in S3. In that world, it could potentially run for short periods, and sleep when there's no work to do.

Cloud only has an early access free tier right now. We intend to make Cloud into a highly multitenant offering, which will make the cost of a user that isn't doing anything with their cluster effectively 0. In that world, we can do really cost effective consumption pricing for low-volume serverless use cases. Absolutely this requires trust, and some users will always want to self host, and we want to make that as easy and cost effective as possible. Its worth noting that we should be able to support client side encryption for journal entries, in time - in which case, you don't have to trust us nearly as much.

p10jkle··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
Super interesting question! If we were inventing modern tech from scratch, I think there's space for this, definitely. Our goal though is that people can use their primitives in the systems they have already, which means Java, Go, Python, TS support are all table stakes
p10jkle··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
It's always been a lively topic within Restate. The conversation goes a bit like this

> Let users write code how they want, its our job to make it work!

> Yes, but it's simply not safe to do this!

I think we need to offer our users a lot of stuff to get it right:

1. Tools so they know when a deploy puts in-flight invocations at risk, or maybe even in their editor, showing what invocations exist at each line of a handler

2. Nudge towards delayed call patterns whereever we can

3. Escape hatches if they absolutely have to change a long-running handler - ways to branch their code on the running version, clever cancellation tricks, 'restart as a new call' operation

Sadly no silver bullet. Delayed calls get you a lot of the way though :p

p10jkle··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
A visualisation/dashboard is a top priority! Distributed architecture (to support multiple nodes for HA and horizontal scaling) is being actively worked on and will land in the coming months
p10jkle··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
I think its a cognito default - will take a look!
p10jkle··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
Yes, definitely, but we can also cover response time bound tasks! Not just async. Typical p90 of a 3-step workflow is 50ms. Our goal is to run on every RPC, anywhere you need reliability
p10jkle··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
I think you would need to validate the response from the first call before determining it to be a success?
p10jkle··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
https://news.ycombinator.com/item?id=40659968 Absolutely, sorry if im not tight enough with my language. Maybe should be described as 'operation idempotency' vs 'handler idempotency'. IMO, an entire handler re-executing is much harder to reason about and test for than a particular operation re-executing individually, with nothing else changing between executions
p10jkle··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
Absolutely, individual atomic side effects need to be idempotent. We can't solve the fundamental distributed system problem there (eg an HTTP 500 - did it actually get executed) However, the string of operations doesn't need to be idempotent - lets say your handler does 3 tasks A B C, and the machine dies at C. Only C will be re-executed. A and B need to be atomically idempotent, but once we move on, we don't start again

Critical point - its much easier to think about and test for the re-execution of C in a vacuum, than to test for A B C all re-executing in sequence, with a variable number of those having already executed before

p10jkle··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
Great question! https://news.ycombinator.com/item?id=40659687
p10jkle··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
> I was of course just thinking about the "front" of the execution, when you're sleeping for 2 days and you want to switch out a future step. Switching out logic that has already been committed is a harder problem. That's a goo point.

You make a good point - this is the idea behind 'delayed calls' which are really one of my favourite things about Restate. Don't save all the intermediary state - just serialise the service name, the handler name, and the arguments, and store that for a month or whatever. That is a very tractable problem - ie just request object versioning

p10jkle··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
> Could you elaborate on that?

What I mean is that executing a software artifact from, lets say, a month ago, just to get month-old business logic, is extremely dangerous because of non-business-logic elements. Maybe it uses the old DB connection string, or a library with a CVE. Its a 'hack' to address old code versions in order to get the business logic that a request originally executed on - a hack that I feel should be used for minutes, not eve hours.

p10jkle··on Show HN: Restate – Low-latency durable workflows for JavaScript/Java, in Rust
I don't personally believe this immutability property should be used for handlers that run for more than say 5 minutes. Any longer than that, I'd suggest the use of delayed calls, which explicitly will serialise the handler arguments instead of saving the whole journal. I agree executing code that is even just an hour old is unacceptable in almost all cases.

Obviously you can still sleep for a month, but I really see no way to make such a handler safely updatable without editing the code to branch on versions, which can become a mess really quick (but good for getting out of a jam!)

Page 1 of 4Next →