HNHacker News
TopNewBestAskShowJobs

ibgeek

167 karma · joined March 7, 2008

submissionscomments
ibgeek··on Andy Pavlo joins ClickHouse to establish ClickHouse Labs
Will you maintain a connection to CMU? If so, what will you continue doing and at what percent of effort?
ibgeek··on Canvas is down as ShinyHunters threatens to leak schools’ data
Didn't realize that. Thanks for the info!
ibgeek··on Canvas is down as ShinyHunters threatens to leak schools’ data
Moodle is an open-source LMS that can be self-hosted.

https://moodle.org/

ibgeek··on Granite 4.1: IBM's 8B Model Matching 32B MoE
They did:

https://huggingface.co/collections/ibm-granite/granite-embed...

311M and 97M versions.

ibgeek··on ARM AGI CPU: Specs and SKUs
Yes, but they function as sister companies right now rather than one company.
ibgeek··on ARM AGI CPU: Specs and SKUs
Agreed. The ARM AGI CPU supports a newer version of the vectorized instructions and has matrix math extensions that the AmpereOne M doesn’t. Also has almost twice the memory bandwidth. One paper at least, the AGI CPU seems like a better choice for AI workloads. Ampere is really pushing the AI workload use cases for the AmpereOne M, so this really makes their lives a lot harder.
ibgeek··on Nvidia Launches Vera CPU, Purpose-Built for Agentic AI
I own one of these systems. My interpretation is the Ampere systems are targeted at lower cost scale out. The Ampere Altra CPUs are limited to DDR4. The raw single core performance doesn’t match Intel or AMD offerings. You get a lot of cores for a lower hardware cost and at lower energy usage.

The Nvidia CPUs are designed for a very specific use case. They are designed for high performance with less concern about cost control.

The newer AmpereOne CPUs use DDR5 with the AmpereOne M supporting even higher memory bandwidth. Even then, I doubt the AmpereOne CPUs will match the performance of the Nvidia Rubin CPUs. But the Ampere processors are available for general use. I am guessing that Nvidia is only going to sell the complete rack system and only to high-volume customers.

ibgeek··on Launch HN: IonRouter (YC W26) – High-throughput, low-cost inference
Since you are very focused on specific Nvidia hardware, I wonder if Nvidia would either buy you out to benefit from your tech or implement their own version without your involvement. Seems risky to me as a potential customer.
ibgeek··on D Programming Language
I used to think of D as the same category as C# and Java, but I realized that it has two important differences. (I am much more experienced with Java/JVM than C#/.Net, so this may not all apply.)

1. Very load overhead calling of native libraries. Wrapping native libraries from Java using JNI requires quite a bit of complex code, configuring the build system, and the overhead of the calls. So, most projects only use libraries written in a JVM-language -- the integration is not nearly as widespread as seen in the Python world. The Foreign Function and Memory (FFM) API is supposed to make this a lot easier and faster. We'll see if projects start to integrate native libraries more frequently. My understanding is that foreign function calls in Go are also expensive.

2. Doesn't require a VM. Java and C# require a VM. D (like Go) generate native binaries.

As such, D is a really great choice when you need to write glue code around native libraries. D makes it easy, the calls are low overhead, and there isn't much need for data marshaling and un-marshaling because the data type representations are consistent. D has lower cognitive overhead, more guardrails (which are useful when quickly prototyping code), and a faster / more convenient compile-debug loop, especially wrt to C++ templates versus D generics.

ibgeek··on Ask HN: What are the recommender systems papers from 2024-2025?
The ACM Recommender Systems conference is one of the leading venues in the field. You might check out what papers were accepted for the 2024 and 2025 conferences:

https://recsys.acm.org/

ibgeek··on Rust's Block Pattern
This seems like a great way to group semantically-related statements, reduce variable leakage, and reduce the potential to silently introduce additional dependencies on variables. Seems lighter weight (especially from a cognitive load perspective) than lambdas. Appropriate for when there is a single user of the block -- avoids polluting the namespace with additional functions. Can be easily turned into a separate function once there are multiple users.
ibgeek··on The universal weight subspace hypothesis
They are analyzing models trained on classification tasks. At the end of the day, classification is about (a) engineering features that separate the classes and (b) finding a way to represent the boundary. It's not surprising to me that they would find these models can be described using a small number of dimensions and that they would observe similar structure across classification problems. The number of dimensions needed is basically a function of the number of classes. Embeddings in 1 dimension can linearly separate 2 classes, 2 dimensions can linearly separate 4 classes, 3 dimensions can linearly separate 8 classes, etc.
ibgeek··on MinIO is now in maintenance-mode
Time to fork and bring back removed features. :). An advantage of it being AGPL licensed.
ibgeek··on Arduino published updated terms and conditions: no longer an open commons
Maybe two different things here: SBCs that run Linux versus microcontrollers (MCUs).

MCUs are lower power, have less overhead, and can perform hard real-time tasks. Most of what Arduino focuses on are MCUs. The equivalent is the Raspberry Pi Pico.

In my experience, the key thing is the library ecosystem for the C++ runtime environment. There are a large number of Arduino and third-party high-level libraries provided through their package management system that make it really easy to use sensors and other hardware without needing to write intermediate level code that uses SPI or I2C. And it all integrates and works together. The Pico C/C++ SDK is lower level and doesn’t have a good library / package management story, so you have to read vendor data sheets to figure out how to communicate with hardware and then write your own libraries.

It’s much more common for less experienced users to use MicroPython. It has a package management and library ecosystem. But it’s also harder to write anything of any complexity that fits within the small RAM available without calling gc.collect() in every other line.

ibgeek··on Feature Extraction with KNN
I'm not sure if I'm understanding correctly, but it reminds me of the kernel trick. The distances between the training samples and a target sample are computed, the distances are scaled through a kernel function, and the scaled distances are used as features.

https://en.wikipedia.org/wiki/Kernel_method

ibgeek··on Show HN: DBOS Java – Postgres-Backed Durable Workflows
I really wish you guys would change the name since the product has moved so far away from the goals and concepts in the original publication. :). I love the product and what you are doing -- it's definitely needed and valuable.
ibgeek··on Bcachefs Goes to "Externally Maintained"
This isn’t BTRFS
ibgeek··on Ask HN: What are the biggest PITAs about managing VMs and containers?
Thanks! I’ll take a look at quadlet.

I find that I tend to package one-off tasks as containers as well. For example, create database tables and users. Compose supports these sort of things. Ansible actually makes it easy to use and block on container tasks that you don’t detach.

I’m not interested in running kubernetes, even locally.

ibgeek··on Ask HN: What are the biggest PITAs about managing VMs and containers?
Ok one more to add that is a kind-of an abuse of containers: Some compute cluster solutions (like those used for HPC) are using containers to manage software installations on the clusters. They are trying to unify containers with the standard Unix environment, however, so that users still see their home directory (mounted in the container) and other paths so that running applications in the container is the same experience as running it directly on the host OS. This is just a TERRIBLE solution. I much prefer Environment Modules or something like Python's virtual environments (if it worked for arbitrary software installs) as a solution.

https://en.wikipedia.org/wiki/Environment_Modules_(software)

ibgeek··on Ask HN: What are the biggest PITAs about managing VMs and containers?
One of the goals of containers are to unify the development and deployment environments. I hate developing and testing code in containers, so I develop and test code outside them and then package and test it again in a container.

Containerized apps need a lot of special boilerplate to determine how much CPU and memory they are allowed to use. It’s a lot easier to control resource limits with virtual machines because the application in the system resources are all dedicated to the application.

Orchestration of multiple containers for dev environments is just short of feature complete. With Compose, it’s hard to bring down specific services and their dependencies so you can then rebuild and rerun. I end up writing Ansible playbooks to start and stop components that are designed to be executed in particular sequences. Ansible makes it hard to detach a container, wait a specified time, and see if it’s running. Compose just needs to be updated to support management of shutting down and restarting containers, so I can move away from Ansible.

Services like Kafka that query the host name and broadcast it are difficult to containerize since the host name inside the container doesn’t match the external host name. Requires manual overrides which are hard to specify at run time because the orchestrators don’t make it easy to pass in the host name to the container. (This is more of a Kafka issue, though.)

ibgeek··on Should you ditch Spark for DuckDB or Polars?
Good write up. The only real bias I can detect is that the author seems to conflate their (lack of) familiarity with ease of use. I bet if they spent a few months using DuckDB and Polars on a daily basis, they might find some of the tasks just as easy or easier to implement.
ibgeek··on 8 months of OCaml after 8 years of Haskell in production (2023)
The Meta post is particularly interesting. Thanks for sharing!
ibgeek··on 8 months of OCaml after 8 years of Haskell in production (2023)
I think the title is misleading. This isn't really about either language in production environments. As other commenters mentioned, a post about production would cover topics like whether there were any tooling / dependency updates that broke a build, whether they encountered any noticeable bugs in production caused by libraries / run time, and how efficiently the run times handle high load (e.g., with GC).

This is more about syntax differences. Even then, I'd be curious how well both languages accommodate themselves to teams and long term projects. In both cases, you will have multiple people working on parts of the code base. Are people able to read and modify code they haven't written -- for example, when fixing bugs? When incorporating new sub components, how well did the type systems prevent errors due to refactoring? It would be interesting to know if Haskell prevents a number of practical problems that occurred with OCaml or if, in practice, there was no difference for the types of bugs they encountered.

This blog post feels more like someone is comparing basic language features found in reviews for new users rather than sharing deep experience and gotchas that only come from long-term use.

ibgeek··on Thinking in Actors – Challenging your software modelling to be simpler
It's not clear from the article whether actors offer significant benefits (or disadvantages) for data modeling versus the traditional OO paradigm. The article reads more like an introduction that describes the problem and teases a solution rather a complete article that offers a solution and evaluation of it.
ibgeek··on Thinking in Actors – Challenging your software modelling to be simpler
The article seems to be smashing together two (seemingly) unrelated topics and doesn't offer much in the way of a solution. What alternative design does the author propose? Is it possible to solve the problem with traditional object-oriented design techniques? It's not clear that the issues presented require or substantially benefit from the actor model without seeing a best in-class OO example.
ibgeek··on pg_flo – Stream, transform, and re-route PostgreSQL data in real-time
This is very cool!
ibgeek··on Why do random forests work? They are self-regularizing adaptive smoothers
Biological trees don’t make predictions. Second or third sentence contains the phrase “randomized tree ensembles not only make predictions.”
ibgeek··on Launch HN: Fortress (YC S24) – Database platform for multi-tenant SaaS
Multi-tenant stuff is very interesting to me.

Do you provide any per-tenant resource limits or prioritization (storage, memory, network [rates plus total], CPU)? Anything to limit the impact of noisy neighbors?

Do you provide per-tenant accounting (for billing) capabilities?

ibgeek··on US colleges are cutting majors and slashing programs after years of delays
Most universities publish a high-level breakdown of their expenses through the government IPEDS database. For teaching-focused institutions like mine, a vast majority of the money is spent on salaries. We are committed to keeping class sizes to 20 students or less and hiring experts to keep teaching quality high. That means that there haven’t been any increases in efficiency other than extracting more labor from faculty. This also seems to be true in other areas as well. Startups solve business process needs by purchasing services (e.g., Bamboo HR, expense reporting software, DocuSign, etc.). Universities still have people doing all of these things manually.
ibgeek··on Show HN: DBOS – Transactional Serverless for TypeScript Apps
That helps. Thanks!

Giving the programmer control over the checkpoints make sense. Also, limiting it to “workflows” (not necessarily entire programs) and requiring that these be deterministic makes sense.

I did some working with Folding@Home in grad school. All of the simulations would save their state to disk every N iterations to allow restarts. The state was relatively small (lots of compute, not a lot of data).

The DBOS papers focused on implementing core OS kernel functionality on top of a distributed relational database. Trying to make a true distributed OS.

This product seems like a pretty different direction (reliable execution and enhanced observability). Do your future plans tie back to the original project goals or are you going to keep going in a different direction?

Page 1 of 3Next →