HNHacker News
TopNewBestAskShowJobs

bitwalker

511 karma · joined May 10, 2013

Senior Compiler Engineer at Miden
submissionscomments
bitwalker··on A different and often better way to downsample your Prometheus metrics
Prometheus without any supporting tooling isn't really designed for long term storage as I understand it, however it is built to support long term storage and querying via its remote read/write protocol. Prometheus will write data to remote storage, and can delegate queries to that storage, rather than using its own local storage as it does by default.

Of the various tools that expose the remote read/write APIs, I like the looks of Promscale/TimescaleDB the most so far, but other options like Thanos might make more sense if you need to collect metrics from a bunch of Prometheuses. That said, maybe you can still use Promscale/TimescaleDB with Thanos as the storage backend, I can't recall the details on its requirements though, so it might not be suitable for that case. For my own use cases though, Promscale is a great solution.

bitwalker··on PostgreSQL 14 on Kubernetes
I think this is where the operator pattern really shines. By defining a custom resource that contains the cluster configuration, the operator can detect certain types of changes, such as upgrading Postgres to a new major version, and automate that change the same way that you'd do it manually. Of course, if such an operator doesn't already exist, its on you to build it, but with popular databases like Postgres, they are generally already out there in some form. That said, I'm not sure if existing Postgres operators handle your specific example.
bitwalker··on Ask HN: Is it worth learning Elixir, from a jobs perspective?
I've been working with Elixir full-time since ~2014, and while I'm probably a bit of an outlier due to my open source libraries and speaking at conferences, I've so far found that there are always opportunities available if I look around. I also get the periodic recruiting emails as well, so there is certainly interest out there. I think if you've got the ability to demonstrate a good programming background (particulary in other functional languages), you won't have much trouble, Elixir is a pretty easy language to start being productive with in a hurry. I know of at least one organization that built up a whole team from people that were essentially brand new with the language and had success with it.

How long will it take to be good enough to get a job? If you are already a senior-level developer, with experience in a functional language, you can probably get hired without any experience at all. Otherwise, you probably need at least a couple months of getting familiar. Publishing an open source library as a means of demonstrating your skill level, and that fills some kind of niche, and allows you to get a feel for the conventions and tooling, is well worth the investment in time IMO. Someone really passionate about finding a job with Elixir could probably get up to speed in just a couple weeks, enough to be productive enough to contribute as part of a team - but that would be basically spending all day every day building something, reading a book like _Elixir In Action_, and actively asking questions on the ElixirForum, IRC, or the Elixir Slack channel.

I think its very doable, but you'll always be at the mercy of who is looking at the moment, and what they are looking for.

bitwalker··on Reviews of Android TV launcher after Google added ads to the homescreen
Which problem? Because there aren't ads on the home screen, or even in the Apple TV app (other than promos for TV shows, but that's like...why you're there).
bitwalker··on SQLAlchemy 1.4
I believe Ecto is largely inspired by LINQ from C#, but I wouldn't be surprised if SQLAlchemy was an inspiration as well.
bitwalker··on Launch HN: Lunatic (YC W21) – An Erlang Inspired WebAssembly Platform
Meetings are scheduled here, along with their planned agendas: https://github.com/WebAssembly/meetings/tree/master/stack/20...
bitwalker··on GameStop Is Rage Against the Financial Machine
They already cashed out $13MM, and the remaining stock were bought at like $14/share. Even if it crashes, they still called this thing like a year out and got rich off of it.
bitwalker··on Why Not Rust?
I imagine the advent of more functionality being available in `const fn` would probably reduce a lot of the `macro_rules!` and proc macro use - at least it has for my projects.

Definitely would like to see a more in-depth guide to patterns for structuring large code bases. I've finally dialed things in pretty well in my own head, but it took quite a while to get there, with a lot of lessons learned the hard way.

bitwalker··on Elixir Is Erlang, not Ruby
> sending a message to a remote node is just a special case of eval. instead of arbitrary code you're evaling `pid ! msg`. and what is spawning a remote process if not remote code eval?

They are not equivalent at all, sending a message is sending data, evaluation is execution of arbitrary code. BEAM does not implement send/2 using eval. Spawning a process on a remote node only involves eval if you spawn a fun, but spawning an MFA is not eval, it’s executing code already defined on that node.

> as for ETS, you can query any data structure with arbitrary functions. that's exactly what i mean when i say there's limited query capabilities. all you can really do is read the keys and values and pass them to functions

You misunderstood, you can _query_ with arbitrary functions, not read some data and then traverse it like a regular data structure (obviously you can do that too).

> my experience and the experience of others is that elixir and erlang are not significantly more efficient than other languages and do not lead to a reduction in the total number of nodes you need to run.

I’m not sure what your experience is with Erlang or Elixir, but you seem to have some significant misconceptions about their implementation and capabilities. I’ve been working with both professionally for 5 years and casually for almost double that, and my take is significantly more nuanced than that. Neither are a silver bullet or magic, but they excel in the domains where concurrency and fault tolerance are the dominant priorities, and they are both very productive languages to work in. They have their weak points, as all languages do, language design is fundamentally about trade offs, and these two are no different.

If all you are building are stateless HTTP APIs, then yes, there are loads of equally capable languages for that, but Elixir is certainly pleasant and capable for the task, so it’s not really meaningful to make that statement. Using that as the baseline for evaluating languages isn’t particularly useful either - it’s essentially the bare minimum requirement of any general purpose language.

bitwalker··on Elixir Is Erlang, not Ruby
For the use case you are describing, none of my points are important really - an HTTP request that hits a database, then pushes something onto a queue for background processing doesn't exhibit any problems from a process bottleneck point of view on that end of things. You still need to have some logic to deal with backpressure from the queue, but that is a language agnostic concern.

Where you could hit a bottleneck might be in the background processing though, take for example the following scenario:

- A pool of N background job worker processes each pull an item off a queue, and spawn a process to perform the task in isolation - A singleton process S provides exclusive access to some resource - Each task calls some code which needs to interact with the resource controlled by S.

The problem with the above is that all of that concurrency/parallelism is nullified by the fact that the tasks are all going to block on S to do their work, the bottleneck of the design.

To be clear, you should always gather telemetry first, but lets assume that you've gathered that and you can clearly see that this bottleneck is an issue (the process mailbox has frequently got many messages waiting to be received, the average time to completion for jobs is increasing). To solve this depends on why the resource is held by S in the first place.

If its because the resource is not thread safe and requires exclusive access, then unless you can find a way to avoid needing the resource in every task, there isn't much you can do, but this should be fairly uncommon in practice.

If S exists because you needed to store some shared state somewhere, and someone told you that an Agent or GenServer was the way to go, then you could move that data to ETS and make it publically accessible so that functions which operate on that data read it from ETS directly rather than call the process. Now you've removed that bottleneck entirely.

If S exists because it needs to protect access to some data, but not all of it, and most tasks don't need to access the protected data, then you can move the parts that do not need to be protected into ETS, and keep the rest in the process. This might reduce the amount of contention on that singleton process by a huge amount, but if even half the processes no longer need to block on accessing it, then you've regained at least that much concurrency in the task processing code.

---

The example above is something I've seen numerous times, but the important pattern to note is that you have some task that you've tried to parallelize by spawning multiple processes, but that task itself depends on something that is not, or cannot be done concurrently/in parallel.

Any time this pattern arises, you need to either find a way to enable concurrency in that dependency, or you should avoid doing the task in parallel in the first place. This is ultimately true of any parallelizable task - its only parallelizable if all of the tasks dependencies are themselves parallelizable, otherwise you end up bottlenecked on those dependencies and you've gained little to no benefit.

Where it becomes a bigger problem is when you consider the system at a higher level. Bottlenecks reduce throughput, which may end up, via backpressure, causing errors on the client due to overload, or depending on the domain, data being dropped because it can't be handled in time (e.g. soft real-time systems).

I don't have any code examples that really encompass all of this in one place, if you are interested in something specific, I can try to throw something together for you. Or if you have specific questions I can point you to some resources I've used to help understand some of these concepts.

bitwalker··on Elixir Is Erlang, not Ruby
> distribution, for example, is a much lauded feature of elixir/erlang but if you look into the implementation it's really just a persistent tcp connection with a function that evals code it's sent on the other end...

I mean, this is just straight up incorrect. Yes the underlying transport is TCP, but using remote evaluation is definitely _not_ the common case. Messages sent between nodes are handled by the virtual machine just like messages sent locally, that is the main benefit of distributed Erlang - referential transparency. Yes, you _can_ evaluate code on a remote node, which can come in handy for troubleshooting or orchestration, but it is certainly not the default mode of operation.

> there's no security model

I mean, there is, but it isn't a rich one. If one node in the cluster is compromised, the cluster is compromised, but the distribution channel is very unlikely to be the means by which the initial compromise happens if you've taken even the most basic precautions with its configuration. It would be nice to be able to tightly control what a given node will allow to be sent to it from other nodes (i.e. disallow remote eval, only allow messaging to specific processes), and I don't think there are any fundamental blockers, its just not been considered a significant enough issue to draw contribution on that front.

> the persistent connections won't scale past a modest cluster size

I mean, there is already at least one alternative in the community for doing distribution with large clusters, Partisan in particular is what I'm thinking of.

> these are both very crude key/value stores with only very limited query capabilities

What? You can literally query ETS with an arbitrary function, you are limited only by your ability to write a function to express what you want to query.

You shouldn't use them in place of a database, but they are hardly crude or primitive.

> elixir/erlang are excellent for software that runs on appliance style hardware where you can't simply add machines to a cluster. it is, in fact, what erlang was designed to do. what this ignores though is that this is a terrible model for a service exposed over the internet that can run on any arbitrary machine in any data center you want

I think you are misconstruing the point of "doing more with less" - the point isn't that you only need to run a single node, but that the _total number of nodes_ you need to run are a fraction of those for other platforms. There are plenty of stories of companies replacing large clusters with a couple Erlang/Elixir nodes. Scaling them is also trivial, since scaling horizontally past 2 nodes doesn't require any fundamental refactoring. Switching from something designed to run standalone in parallel with a bunch of nodes versus distributed _does_ require different architectural choices, and could require significant refactoring, but making that jump would require significant changes in any language, as it is a fundamentally different approach.

> elixir/erlang's features that increase it's reliability on a single machine are a cost you pay not an added benefit. the message passing actor model erlang built it's supervision tree features around are a set of restrictions that are imposed so you can build more reliable stateful services on machines that don't have access to more conventional approaches to reliability (like being stateless and pushing state out to purpose built reliable stores)

I'm not sure how you arrived at the idea that you can't build stateless servers with Erlang/Elixir, you obviously can, there are no restrictions in place that prevent that. Supervisors are certainly not imposing any constraints that would make that more difficult.

The benefits of supervision are entirely about _handling failure_, i.e. resiliency and recovery. Supervision allows you to handle failure by restarting the components of the system affected by a fault from a clean slate, while letting the rest of the system continue to do useful work. This applies to stateless systems as much as stateful ones, though the benefits are more significant to stateful systems.

> the idea that these features are appropriate for a totally standard http api running in aws or digital ocean or whatever backed by a postgres database and a memcache/redis cluster is not really born out by reality however. if it were surely other languages would have incorporated these features by now? they've been around for 30 years and the complexity (particularly of distribution and ets) is low enough you could probably implement them in a weekend

The reason why these features don't make an appearance in other languages (which they do to a certain extent, e.g. Akka/Quasar for the JVM which provide actors, Pony which features an actor model, libraries like Actix for Rust which try to provide similar functionality as Erlang) is that without the language being built around them from the ground up, they lose their effectiveness. Supervision works best when the entire system is supervised, and supervision without processes/actors/green threads provides no meaningful unit of execution around which to structure the supervision tree. Supervision itself is built on fundamental features provided by the BEAM virtual machine (namely links/monitors, and the fact that exceptions are implemented in such a way that unhandled exceptions get translated into process exits and thus can be handled like any other exit). The entire virtual machine and language is essentially designed around making processes, messaging, and error handling cohesive and efficient. Could other languages provide some of this? Probably, though it certainly isn't something that could be done in a weekend. No language can provide it at the same level of integration and quality without essentially being designed around it from the start though, and ultimately that's why we aren't seeing it added to languages after the fact.

bitwalker··on Elixir Is Erlang, not Ruby
My opinion is that this depends entirely on the cost relative to the overall task, and how likely cache hits are to occur. If cache hits are very likely and the task occurs frequently, I'd strongly consider storing it in ETS. If cache hits are unlikely, then it depends purely on how expensive the task is, but generally there isn't a lot of benefit to caching things that are infrequently accessed.

I wouldn't cache database queries unless the query is expensive, or the results rarely change but are frequently accessed.

Generally though, whether to store something in ETS or not is situational - your best bet is actually measuring things and moving stuff into ETS later when you've identified the areas where it will actually make a meaningful difference.

> This part throws me off because I remember hearing various things in Phoenix work in a distributed fashion without needing Redis.

This is true, but it depends on what kind of consistency model you need for that distributed state. The data you are referring to (I believe) is for Phoenix Presence, and is perfectly fine with an eventually consistent model. If you need stronger guarantees than that, you'll need a different solution than the one used by Phoenix - and for most things that require strong consistency, its better to rely on the database to provide that for you, rather than reinvent the wheel yourself. There are exceptions to that rule, but for most situations, it just doesn't make sense to avoid hitting the database if you already have one. For use cases that would normally use ETS, but can't due to distribution, Mnesia is an option, but it has its own set of caveats (as does any distributed data store), so its important to evaluate them against the requirements your system has.

bitwalker··on Elixir Is Erlang, not Ruby
Process bottlenecks are a design problem, not a language or syntax problem; and are mitigated largely by a few points that can be factored in during design or PR review:

- Be wary of places where you have N:1 process dependencies, where N is large and the number of messages exchanged between each member of N and the single process are frequent/numerous. Since each process can only handle received messages sequentially, there is little point in spawning a lot of tasks in parallel if each process has to talk to the same upstream process to do anything

- Set up telemetry that samples the number of messages sitting in the process mailbox; if a process is becoming a bottleneck, it is going to be frequently overloaded and have a lot of messages in its mailbox. If you have the telemetry, you can see when this starts to happen, and take steps to deal with it before it starts causing problems for you. Likewise, its probably useful in general to have telemetry on how long each unit of work takes in server-like processes, so you can get a sense of throughput and factor that data into your design.

- Avoid sending large messages between processes, instead spawn a process to hold the data and then send a function to that process which operates on the data and returns only the result; or store the data in ETS if you have a lot of concurrent consumers. It can also be helpful to denormalize the data when you store it in ETS so you can access specific parts of it without copying the entire object out of ETS on every access. The goal here is to make messaging cheap and avoid copying lots of data around.

- Take steps to ensure process dependencies in your design are structured as trees, i.e. avoid dependency graphs that can contain cycles. It is all too easy for a change to introduce the possibility of deadlock if you play fast and loose with what processes can talk to each other. If your process dependencies mirror your supervisor tree, then you can protect against this by only allowing dependencies between branches in one direction (usually toward the parts of the tree that were started earlier in the supervisor tree)

I think the problem is that Elixir is still relatively young, and due to the language evolution and the lack of established documented doctrine from the Erlang community, there is a lot of techniques, tips, design patterns, etc., that are being rediscovered; likewise there are a lot of seemingly good ideas that turn out to be not so great in practice, but are encountered on the road to the truly sound patterns. So you get a lot of people writing about the lessons they are learning, and because of the gaps in knowledge, the result is that the information may be missing things, or providing a more complex solution when there is a simpler one, etc. Ultimately this is an important process, and now that Elixir has largely stabilized, this will only improve (and its is already pretty good, certainly far better than when I first started with the language years ago).

bitwalker··on Named Parameters in C++20
I'd expect scalar replacement of aggregates (SROA) would eliminate the use of a struct like this in some cases, but there are definitely limits, especially if not compiling with LTO enabled, since optimizations across compilation units will be limited. Honestly doesn't seem worth the cost.
bitwalker··on Enigma: Erlang VM Implementation in Rust
Absolutely! I strongly believe we should be building alternative implementations of the Erlang compiler, the runtime, the virtual machine - all of the major components. It provides a lot of useful data that, if nothing else, can be used to improve the original implementations in the BEAM. Enigma is a great example of how reimplementing such a core piece of that infrastructure can provide better learning opportunities for the community, as well as a way to clean the slate and start with a fresh look at the problem with modern tools.

If nothing else I'd really like to see projects like ours demonstrate the value to the core Erlang/OTP team in addressing the lack of documentation in some areas - ideally in the form of one or more specifications. Erlang deserves a specification at this point - it is very stable, and a spec would at the very least provide additional structure for future evolution. Core Erlang had a specification, but it is very much out of date at this point - considering how widely it is used as an IR for BEAM languages in general, it's disappointing it hasn't been kept up to date.

bitwalker··on Enigma: Erlang VM Implementation in Rust
I suspect there exists some internal documentation at Ericsson that just hasn't been cleaned up and added to the source repository - but it's entirely possible that the core Erlang/OTP team solely relies on passing the knowledge on from engineer to engineer.

I'm also not saying that the C code is unmaintainable. It's definitely a bear to dive into, but by spending enough time with it, it starts to unfold in front of you. The main issue I have, is that none of the specification/design documentation exists as part of the source repository. Maybe it doesn't exist at all, but in that case, I'd really hope that some of those core engineers would have taken the time to write some of that stuff down. In any case, none of it is readily available AFAIK.

bitwalker··on Enigma: Erlang VM Implementation in Rust
100%, when I started building Lumen, I spent an enormous amount of time working out how various parts of the BEAM runtime were implemented, and it was (and still is sometimes) a grind. That macro-heavy C code is just such a bear to read. I understand why it was written that way, but working with it is just unpleasant.

Having Enigma, and other implementations like it, provides a huge value in terms of understanding how it all fits together - understanding the Enigma implementation and then going and trying to make sense of the BEAM would probably be a way better path than trying to dive into the BEAM straight away.

bitwalker··on Enigma: Erlang VM Implementation in Rust
There are really two major pieces of the BEAM, the parts implemented in Erlang (e.g. the compiler, OTP), and the parts implemented in C (the VM, or emulator as it is called in the codebase). Virtually all of the C code is undocumented in any meaningful way, short of some internal documentation on a handful of topics, as well as a those parts of the code which have thorough comments that explain some tricky aspect of the implementation.

From my own experience, those parts that are commented or documented tend to clarify some specific design constraints (for example, why processes have multiple locks on different parts, and why they are locked in a specific order, or the rationale of the carrier design); but you never really get a clear picture of why things overall are architected the way they are overall, what designs were considered and discarded due to some deficiency, what tradeoffs were made, etc. I think much of the actual content like that which may exist, is either buried in the minds of the original engineers, or in some internal documentation at Ericsson that has never been released. My suspicion is that you'd need to dig through mountains of emails and such to piece together a more complete picture of how things where put together over time.

It's also the fact that the BEAM just has a lot of really complex pieces built in to it after all this time. Everything from binary pattern matching and construction, to garbage collection and memory management, ETS, Mnesia, etc. Each one of those things is not only non-trivial, but have evolved significantly over time, through the hands of many engineers. It also doesn't help that large portions of the C implementation are written in an extremely macro heavy style, which makes it quite hard to read without knowing what all the macros do and how they play together.

Projects like Enigma, or Lumen, have a lot to give back to the community in the form of documenting how these pieces are built. Unfortunately, the lack of a specification for the Erlang language and its runtime, means it is very much a grind to work out how things are currently implemented, and why.

bitwalker··on Enigma: Erlang VM Implementation in Rust
I think we're coming at things from sufficiently different angles that it's not really necessary to merge our projects, other than perhaps some common parts that could benefit from that kind of sharing. Enigma is really a faithful reimplementation of the VM in Rust; while Lumen is an ahead-of-time compiler, which requires a very different approach.

I think the goal of Enigma in making the BEAM architecture easier to understand, and providing a great learning platform for getting involved in working on the BEAM itself, or just on Enigma as an alternative is a great idea, and something I know I wish I had been able to have on hand when I was trying to understand the deep inner workings of the BEAM implementation. I if Enigma was only ever that, it would still be worth the effort spent on it.

Lumen does have parts that could likely be shared with projects like Enigma - namely the high-level IR we use in the frontend (EIR, Erlang Intermediate Representation). That IR could be used in any Rust-based VM/compiler targeting Erlang, or an Erlang-derived language, and work there would directly benefit downstream consumers of the IR in terms of better optimization, etc. We've also put work into our term representation, and various parts of the runtime, like reproducing the core parts of the BEAM garbage collector, etc. While some of those things are intertwined with the compiler, much of it could be easily extracted and used elsewhere.

bitwalker··on Lumen – Elixir and Erlang in the Browser
Sure, I can't imagine any particular reason why that wouldn't work. That said, we don't have distribution implemented yet, since we are focused on a single node in the browser initially; so it depends on which comes first - SharedArrayBuffer being stabilized again in the browsers so we can enable multiple schedulers, or getting distribution implemented thoroughly enough to support the approach you mentioned as a workaround.
bitwalker··on Lumen – Elixir and Erlang in the Browser
Yes, Lumen has a scheduler and does preemption the same way the BEAM does.
bitwalker··on Lumen – Elixir and Erlang in the Browser
Initially we’re not planning to support multiple schedulers via WebWorkers, but the runtime has support for multiple schedulers generally speaking. Since the compiler can target other platforms, we didn’t want to artificially limit all targets due to any particular target limitations. As far as Wasm is concerned though, once browser support for SharedArrayBuffer returns, then we’d be looking to take advantage of multiple threads.
bitwalker··on Lumen – Elixir and Erlang in the Browser
That’s not how WebAssembly works, at all, it is handled directly by browser JIT compilers.
bitwalker··on Lumen – Elixir and Erlang in the Browser
There is also a Python library called Elixir - naming is hard.
bitwalker··on Lisping at JPL (2002)
Yeah, I learned Logo in elementary school and it was something that stuck with me until I actually started programming later on in school. I particularly remember having to write a program to navigate a maze, it was a blast!
bitwalker··on Functional Imperative Programming with Elixir (2018)
It just isn't needed, most things can be expressed using nothing more than recursion; more complex things can be expressed using `for` comprehensions, or by composing the various `Stream` APIs.

I think `while` can more succinctly express certain patterns, but that window is pretty narrow, so adding an additional keyword to the language that has so much overlap with other options just doesn't make sense.

Elixir is a little unique in that it has no mutability, so fundamentally a `while` loop isn't really a thing that makes sense - you can't mutate anything in the containing scope, so you are left with basically nothing useful you can do. The macro I build in this post is a clever way of making it work, but only with the bindings given to the macro, and it isn't actually mutating them.

For other FP languages that are immutable, they likely don't have a `while`, or provide some kind of machinery that emulates it (like Haskell's monadic loops). For FP languages with opt-in mutability, they might have `while`, just depends on the language.

bitwalker··on Functional Imperative Programming with Elixir (2018)
> Do you use this while-macro in production code or just test code?

I use the `while` macro in the Distillery tests, since it made expressing certain tests much less verbose - namely those dealing with building up a release, spinning it up in a new node, and then verifying conditions on the running node meet expectations. In general though, I don't. I haven't needed it elsewhere (well, I have one other library I have considered using it in, but haven't yet). I would be hesitant to introduce it without a clear need, just because it would likely catch other developers off guard unless clearly documented.

To be clear though, I would support something like it in the language, but also understand why it isn't there; there just isn't a clear enough need, and as you've demonstrated, `Stream` can more or less be used to cover the general cases.

> P.S. thanks for all of your Elixir tools!

Thanks for the kind words!

bitwalker··on “What Alan Kay Got Wrong About Objects”
Apologies, Professor, I missed your reply!

SharedArrayBuffer makes communication between schedulers essentially equivalent to non-browser environments, but support for it is still iffy at best due to Spectre. Luckily, that is a temporary problem, not a fundamental issue (SharedArrayBuffer being disabled that is, not Spectre).

In a single-threaded setting, there isn't any significant overhead that I'm aware of between a browser and non-browser environment, aside from any overhead potentially incurred by executing on the WebAssembly VM rather than directly on the host machine. Of course, running single-threaded is less than ideal for other reasons.

bitwalker··on Functional Imperative Programming with Elixir (2018)
Well, I pointed out at the beginning that one can use recursion, and the implementation of the `while` construct itself uses recursion as well (in fact, you can't implement `while` without it in Elixir).

In any case, expressing the equivalent of a `while` loop with predicates and timeouts is quite syntactically noisy in Elixir - it is certainly doable, but much less clear than the imperative equivalent.

bitwalker··on Functional Imperative Programming with Elixir (2018)
Sure, the implementation of the macro is complicated, but the actual `while` construct that results is just as expressive as that of any language with a native `while`.

But that's all beside the point, this was just an exercise, and was useful in a project where the reduction in complexity it provided was handy

← PreviousPage 3 of 7Next →