HNHacker News
TopNewBestAskShowJobs

CHY872

1,379 karma · joined April 8, 2014

submissionscomments
CHY872··on 'Attention is all you need' coauthor says he's 'sick' of transformers
In computer vision transformers have basically taken over most perception fields. If you look at paperswithcode benchmarks it’s common to find like 10/10 recent winners being transformer based against common CV problems. Note, I’m not talking about VLMs here, just small ViTs with a few million parameters. YOLOs and other CNNs are still hanging around for detection but it’s only a matter of time.
CHY872··on The Speed of VITs and CNNs
The article basically argues: You would expect to get similarly good results with subsampling in practice. E.g. no need to process at 1920x1080 when you can do 960x540. Separately, you can break down many problems into smaller tiles and get similar quality results without the compute overheads of a high res ViT.
CHY872··on Signal to leave Sweden if backdoor law passes
Archive formats are hard to make reproducible because there are lots of ways of making different yet equivalent archives. So it’s not surprising to me that someone would fail at this hurdle and find it frustrating to resolve. Nix defined their own format for this to avoid this exact problem.
CHY872··on Why does FM sound better than AM?
It’s not immediately clear that Shannon’s theorem is a good point of comparison here, since it’s only recently that coding schemes have really approached the Shannon limits, and FM and AM do not use these.

Even if one does assume a Shannon-perfect coding scheme, as the noise ratio gets greater the benefits of spreading a signal across a higher bandwidth fades. Furthermore, most coding schemes hit their maximum inefficiency as the signal to noise ratio decreases and messages start to be too garbled to be well decoded.

I’d additionally note that folks get near the Shannon noise limit _through_ ‘magic noise rejection’ (aka turbo and ldpc codes). It’s therefore not obvious that FM isn’t gaining clarity due to a noise rejection mechanic. The ‘capture effect’ is well described as an interference reducing mechanism.

Empirically, radio manufacturers who do produce sophisticated long range radio usually advertise a longer range when spreading available power across a narrower rather than wider bandwidth.

CHY872··on Ask HN: What's an appropriate compensation counter offer in London 2024?
No, you (generally) pay income tax on the current value when you receive it, and then capital gains on any change in value.

Essentially, if I pay you £15k for some services, you pay income tax on that £15k. If I buy you a car for £15k, taxman still wants £15k. Same with equity.

How it can work is that you can be granted shares and pay the taxes at time of granting (which for a founder is zero), but might not be nice for an employee.

You can also give an employee options, and this can get complicated (single or double vest).

But in general, for any equity instrument in a stock plan, you get charged income tax at one stage, then capital gains at a later stage, and you can trade off when you want to trigger each. In this case, employer has gone for a model where income tax is deferred as late as possible.

CHY872··on Ask HN: What's an appropriate compensation counter offer in London 2024?
Firstly, work out how much your 0.2% equity is likely to convert into. You'll probably pay income tax and employers NI on them (and in one year!) and so you'll likely end up paying 55% tax on them. How much do the founders think the company is worth right now? If it's close to £500M, that's a big part of your comp. If it's close to £50M, it's not.

Next, work out where you want to be in terms of comp in a few years, rather than thinking of how to optimise the cash right now. For example, I'd stop worrying about your tax-free allowance gradually disappearing, and instead try to work out how to get it to all be gone. In 5 years, the person who makes £120k is £40k better off than the person who makes £100k, after tax. That... sounds worth it. And it's usually easier for the person who's getting paid £120k to get paid £130k than it is for the person who's getting paid £100k. This is to say, having high tax brackets is a benefit, not a curse.

And then, it's probably worth noting - this isn't directly their money, especially if they're looking to sell. It's probably worth having the conversation of like, 'what would I need to do in order to justify £100k/year?'. Or, alternatively, negotiating on the vesting of your stock, since that's effectively free. If they think the company is going to be sold in the next few years, that's a relatively small giveaway for you. If the company's grown a lot in four years, it's unlikely a significant increase in stock is on the table.

Don't overestimate how long it'd take a good new person to catch up. I've rarely seen a role where a new person can't be effective within 6 months.

CHY872··on Batteries as a Military Enabler
Probably a few factors.

1. Scaling. You want to reap the rewards of someone else investing billions, and while billions of ICE engines are built every year, most of them are much bigger than 50cc. 2. Tolerances, as you say. I know for example that jet engines have low efficiency at small sizes due to efficiency being driven by the gaps between certain rotating parts, which are relatively larger. 3. Certain parts that need miniaturisation are more expensive on smaller engines. For example, a 50cc would typically have a carburettor, a bigger engine fuel injection. A fuel injector would be significantly larger per unit. 4. Some parts are just harder to miniaturise. For example, small turbochargers have to work harder and at much higher RPMs to achieve the same boost due to area scaling quadratically with diameter.

CHY872··on Batteries as a Military Enabler
Battery efficiency is _broadly_ linear; 10kg will give you 100x the capacity of 0.1kg of batteries. This isn't the case with generators.

Under 5kg of batteries, engines basically can't compete. The smallest viable petrol engines (you want engines made in large quantities) are around 4kg and need fuel, and so you end up in a situation in which 4kg of engine and 1kg of fuel is as useful as 5kg of batteries (for motors of that size), but every subsequent 1kg of fuel is then also as useful as 5kg of batteries. But you can't really drop this down much, as a 1kg engine is much less useful than 1kg of batteries, and a drone with 5kg of batteries is really a very large drone.

Drones are additionally typically very small and light, with mass at an absolute premium. A typical quadcopter will weigh under a kilogram and have maybe 200g of batteries for 20-30 minutes endurance. A drone with a 2.5m wingspan will typically have room for maybe 1-2kg of extra payload, and an engine will not fit into the battery slot.

Furthermore, they are extremely sensitive to weight balance issues.

This is to say, once you're in the world where you want chemical fuel, you might as well design a drone for it, rather than trying to retrofit. The mass of retrofitting will mess up your prior drone design, and petrol changes the dynamics of what you're trying to do enough that you might as well just do it all differently.

Essentially, if you're making a 1-10kg drone, you want battery. If you're making a 10-20kg drone, you might want battery, you might want petrol. Above 20kg, you probably want petrol.

CHY872··on LLM-generated code must not be committed without prior written approval by core
'Please translate this AVX-256 algorithm to Arm NEON' is absolutely a prompt I'd give ChatGPT, and similar prompts have revealed new and useful intrinsics. Of course, results to be checked.

Assurance is a complex topic, and any safety critical device should have a carefully thought through architecture and rigorous testing program which minimises the risk of incidents. It therefore seems scarcely relevant here, beyond the fact that a well defined delivery system should be able to handle multiple human errors during implementation without leading to crucial failure modes occuring.

CHY872··on LLM-generated code must not be committed without prior written approval by core
Yes, if you take a problem you don't understand, ask a GPT to write a solution, do nothing to check the solution, trust it blindly, and then use the solution for some safety critical problem, you're playing with fire. But there's a spectrum, and please don't assume I'm at that end of it.

Lots of the time you understand the problem, but the problem is repetitive. Parsing a weird file format might well be that. Beyond that, you have solutions that are easily checked. For example, if I ask ChatGPT to optimise an algorithm for a certain CPU cache, I can easily read whether it did that. And then, there are parts of a software job that are crucial and subtle, and parts that are not.

As a practitioner, traditionally that leads to a shift in the focus and speed with which you approach a task - some pieces of code are 100 lines that took you 2 weeks to get to and were hard fought, some are 2000 lines which you wrote in a day.

Lastly, so much of solid software is being able to understand a probably unfamiliar domain, and ChatGPT can be a great buddy in terms of gaining problem context, finding the limits of your own understanding.

I don't use co-pilot like things, but I've found ChatGPT to be a massive enabler in terms of being able to be productive in unfamiliar problem-spaces.

CHY872··on LLM-generated code must not be committed without prior written approval by core
'My customer has given me the documentation to arcane antiquated format X (insert pdf, but it includes 24-bit integers, hex encoded data, semantically signficant whitespace). Here is a sample of the message format, and this struct should represent the contents. Please write me a Python parser which takes an input file in the format and provides the output. In particular, given input X, the output should be equivalent to Y. There should be a set of unit tests for important functionality, which should explain to a reader what is complicated about the format'. is something that ChatGPT4 will just drop out a solution to that works in a few seconds.

If your job is to make an accurate parser for it, probably you want to hand code it. If your job is to make sense of the data the customer has provided you with, this is merely an impediment to your actual job, and ChatGPT has you covered. Yes, there'll be mistakes. But ChatGPT can do in a few seconds what'd take you hours.

CHY872··on Java virtual threads hit with pinning issue
The specific edge case which is most annoying here is that locking a Java object with synchronized converts all the non-blocking calls that Loom does further down the stack to be blocking ones, which then lock up your carrier threads. So you get this action as a distance thing.

E.g.

    for (int i = 0; i < 500; i++) {
        newVirtualThread(() -> synchronized (new Object()) {
             Thread.sleep(100_000);
        });
    }
will blow up your JVM (modulo some compensation mechanisms that work definitely kinda) and that's odd.
CHY872··on Java virtual threads hit with pinning issue
I think this is a bit more subtle and nasty than bog standard priority inversion. Specifically, in present virtual threads, if you take out a lock with synchronized, if you then do a normally-non-blocking asynchronous API call while you have the lock, you block one of your very few OS threads because the virtual thread can't now be migrated off the OS thread.

IMO the compound thing is what makes it be nasty. E.g. you have a function `doSomething` which does some RPC, and that's all nicely non-blocking. But someone called map.computeIfAbsent(x, k -> doSomething(k)), and that uses synchronized on the inside so now your non-blocking API calls all magically became blocking, no further action required.

CHY872··on Java virtual threads hit with pinning issue
Nah, it's not normal or really documented. Normally when languages have async, they end up with dual primitives - the synchronous one and the asynchronous one, and everyone learns that if they do the blocking operation on the async thread, they deserve their deadlock. Generally you then end up with two variants of the API, with the standard function colouring problem where you can call lockBlocking() from a non-async context, lockAsync() from an async, and woe betide you if you get mixed up. Ideally your language can help you avoid these issues sometimes (e.g. with Rust you can't hold a non-async lock over await points).

Java instead did a big thing where they made all the primitives work fine in both cases with one API, something which is honestly really hard and really reflects the level of thought put into modern Java features. But there's a long tail of stuff that's still getting cleaned up (e.g. various weird I/O apis), and honestly I think it _is_ weird that an actual language keyword made it into the long tail.

It's extra fun that it's actually quite hard to hit performance problems as a result of this, as the JVM will actually detect that it's getting to this problem and boot up threads to compensate.

CHY872··on Java virtual threads hit with pinning issue
I think this one's weird because the language puts so much effort into making it hard to hit this that you won't notice until it's a gigantic problem. For me, this was enough of a problem that I wrote a Java agent which patches the method which blocks the carrier thread to just throw (ByteBuddy makes it easy!).

Like, it makes sense that when you synchronized on an object, there's room for contention. It seems weird that synchronizing on an _uncontended_ object can cause contention. But that's what this is. The behaviour is that you have target numCores carrier threads, and if someone synchronizes on an object and then does a non-blocking sleep, it's now blocking, because synchronizing upgraded the non-blocking I/O to blocking.

So basically when you hit this issue, it's because not only has the bad thing been happening, it's also been happening badly enough that all the compensation mechanisms have failed.

It's just weird that a whole language level keyword behaves this badly.

CHY872··on Java virtual threads hit with pinning issue
I've been coding in Java for a bit longer than that, and in my memory it's always been known that ReentrantLock is _faster_, but synchronized has better language level support (as in, the language won't let you forget to unlock, whereas with locks you need to remember your finally statement). And then, unlike ReentrantLock, synchronized uses no extra memory, which is good for things that end up being uncontended, and then in both cases, the performance doesn't matter if the lock is uncontended, which most of mine were.

There of course are things like ErrorProne that statically check that you unlocked your lock, but there's still bugs possible.

CHY872··on Meta's new LLM-based test generator
bazhenov/tango does something like this for performance tests, basically to counter system behaviour you run the old and new implementation at the same time.
CHY872··on Microsoft seeks Rust developers to rewrite core C# code
C# has a garbage collector for specifically tracking memory, but lifetimes are more broadly useful.

For example, Rust lifetimes (this is also the case in C++ afaik) can be used to suitably scope the lifetimes of mutexes, to have temporary folders which are deleted when they go out of scope, to require that a connection pool is destroyed _after_ the last connection inside it is returned, etc, etc.

Mostly, garbage collected language do a bad job of cleaning up objects which refer to resources held elsewhere. Java had persistent issues with direct ByteBuffers (which were wrappers around malloc (but not free!)). Locks are easily held too long. File handles are easily left open. And depending on your GC settings, that file descriptor that's holding a 10GB file around may not get cleaned up for hours.

Refcounted languages can be somewhat better, but they don't avoid the bug, they just mitigate the effects.

CHY872··on If you can't reproduce the model then it's not open-source
And, that's obviously fun, because with LLMs, you have the LLM itself which cost hundreds of thousands in compute to train, but given you have the weights it's eminently fine-tunable. So it's actually not really like Linux - rather it's closer to something like a car, where you had no hope of making it in the first place but now you have it, maybe you can modify it.
CHY872··on SD4J – Stable Diffusion pipeline in Java using ONNX Runtime
Java 5 and Java 8 were both very big - generics in 5, lambdas in 8. 6 and 7 were iterative in comparison.

There were very many important changes in the meantime over that timeframe, but generics and lambdas fundamentally changed how you use the language - Java 4 is not the same language as 5, same between 7 and 8. This is not the case for the 6 and 7 releases.

CHY872··on The first commercial carbon-sucking facility in the US opens in California
There’s just an economic bar where below that point, it’ll be what happens. Similar thing with solar and electric cars. Already cheaper to install solar than keep running coal. Give it a couple years and a government will be offering a rebate on cars only if the car can double as grid storage when parked.
CHY872··on The first commercial carbon-sucking facility in the US opens in California
I while ago I rand some relatively unscientific numbers, and learned that even swamps and forests are relatively inefficient users of space as it pertained to carbon capture. For example, a well managed bamboo plantation will yield 25T/hectare/year of bamboo, and this includes quite a bit of labour for the management. Meanwhile, the same area, covered in solar panels, will get you to around 500MWh per year of produced capacity.

The world emissions per capita are presently around 5T per year, per capita. The world has 5 billion hectares (out of 13B total) of agricultural land in general, overall (and generally if land _could_ be used for agriculture it _is_ used for agriculture), so generally, there's not a tonne of space to do lots of carbon sequestering quickly with agriculture, especially as you'd need to move the carbon you've created somewhere else.

The moral is that if you can ever get sequestering carbon down such that you can sequester 1 tonne of carbon for 20MWh power, you're close to a major winner, at least in all the deserts where you clearly can't grow bamboo but can grow solar, because at that point your cost of operations is just cost of infra. That startup thinks it can get down to around 1MWh/tonne, which is comparatively awesome!

My conclusion (before getting to something super scientific) was that if you want to rely on trees and swamps for your carbon capture, you basically end up with massive geopolitical issues because you need to cover most of the world in trees and swamps, but most of the world's land is already used to grow food. Meanwhile, carbon capture can work in areas where land is not (as) valuable.

If they are able to get to $50/tonne, that implies 1MWh/tonne, so that's about 500 tonnes/year per hectare. That'd mean that if you covered arizona in solar panels, you'd be sequestering 1/5 of human carbon output. Whereas if you grew bamboo, you'd cover the contiguous US for the same output.

Do I believe them? No. But even 10x worse is cheap enough to change the world.

CHY872··on OpenBSD: Removing syscall(2) from libc and kernel
OpenBSD’s position is far from unique. It’s shared with MacOS (and iOS).
CHY872··on Uber migrates microservices to multi-cloud platform running Kubernetes and Mesos
The size of the savings you can negotiate is a function of how locked in you are. A big customer will get discounts well beyond reserved instances, in response for (usually) committing to increase their expenditure above where it is.

The better your competing offer, the better the negotiating position. And while Amazon hasn’t gotten more expensive per se, it’s certainly not gotten cheaper.

CHY872··on Progress on No-GIL CPython
In fairness, concurrent data structures generally don’t represent what gp is talking about. Yes, very, very difficult to write a solid concurrent ring buffer. Writing Java’s concurrent hashmap - hard. Using it to implement a simple in memory kv store: not that hard.

But ensuring some parallelism while maintaining thread safety is straightforward in many contexts - an uncontended mutex is close to zero overhead. Languages like Rust make it harder to struggle with some of the thornier issues (data races which cannot be simulated by stepping threads).

E.g. in Java a typical system looks like a thread per request with some shared underlying data structures like caches or connection pools. Relatively easy to use these safely, or to guard some shared object with a synchronized.

Likewise with parallelism - a lot of problems just boil down to ‘do a few map reduces’ and the parallelism is pretty trivial.

Obviously, concurrent systems are fiendish to reason through - but there are a lot of cases where the complexity can be side stepped. Doesn’t seem to stop people writing scary code on the daily though.

CHY872··on The Changing Relationship Between Memory, I/O, and Network Bandwidth
right - why won't 'networking w/ 800GbE' be different (in terms of weird programming model, crazy interesting papers, etc)?
CHY872··on The Changing Relationship Between Memory, I/O, and Network Bandwidth
The other element which appears unsaid is that in a typical datacentre, your bisection bandwidth is typically << (num computers * network bandwidth per computer). Or in other words, even if your computation is bandwidth and not latency starved, you're not obviously going to be able to do a gigantic data shuffle quickly unless it's within a rack. This is to say, I'm not sure _just how much_ this would affect your overall target system architecture at present. Once you're switching at a couple of terabits things are likely quite different in those terms.

The other element which is a bit scary is that it's fairly rare these days for mass market companies to index deeply on tech which isn't available in public clouds, so until AWS supports this sort of thing, it's unlikely that many folks will target it. You can kind of see this with the Optane PDIMMs - they looked absolutely fantastic, but given you couldn't get them on any AWS instance there wasn't much point actually trying to use them outside of very specific applications - as a software engineer this hardware lets me build my software very differently and in a simpler way, but how can I possibly risk architecting based on that if it then cannot support a customer's cloud migration?

Obviously v different in HPC contexts.

And it should always be said - latency is very different between these. Memory latency still measured in nanoseconds, PCIe latency still measured in 10s of microseconds, about 3 orders of magnitude difference.

CHY872··on Remote Attestation TLS (RA-TLS)
These APIs can be combined quite nicely with 'secure enclave' processing technologies like Intel SGX. The idea is that (at the very least) a processor can attest that the process communicating with your server is an unmodified copy of your binary. A further version might be that the data associated with that process remains encrypted and other processes are unable to read it. Apparently it mostly works! But as is usual with modern processors, there are many side channels.

This has some cool use cases! For example, Microsoft SQL server can already use SGX to implement additional data security. A user can run sql queries including pushdown filters on data the administrator of the server can never access, because certain columns of the data is encrypted and never held unencrypted in memory. If you're at a tech company and have ever worried about rogue administrators accessing the data of your users, these enclaves are great for that (in theory)!

Right now, people have to build their own transport layers, which interact with the attestation APIs. These folks are trying to build something that is as easy to set up as TLS.

A problem with all this tech is that to the extent it can be used to make business problems easier to solve, it makes it much harder to introspect what running software is doing, which from a software freedom perspective tends to raise hackles. My hope would be that in general, this is more used like 'corporate TLS interception' insofar as your personal device does not do it, but I'd expect that mobile device vendors use it before too long.

CHY872··on Alpine Linux does not make the news
My experience with musl has been that while it’ll most probably work, you’re severely at risk of a 30%+ perf degradation, and that means the juice is really not worth the squeeze.

glibc shouldn’t be statically compiled in because it’s lgpl and so immediately infects your code if you do.

The zig linker is quite nice here because it lets you pick what glibc you want to be compatible with.

CHY872··on Yelp rebuilds corrupted Cassandra cluster using its data streaming architecture
I have >5 years of experience using Cassandra in production, involving thousands of clusters storing petabytes of data. My conclusion from that time is that Cassandra is simply not robust enough to be a general purpose database (the team are working on it but they're coming from a really rough starting place) - there are lots of ways to cause data corruption, and Cassandra does enough dynamic repairing that it can be hard to catch this before your backups are dropped due to time windowing. Unfortunately, the juice may still be worth the squeeze - Cassandra's storage model lends itself very nicely to disaster recovery workflows in a way which something like Oracle or FoundationDB does not (and it's Cassandra so you'll need it!), while the ability to horizontally scale gets you out of so many operations issues. If you've got a schema which works well in Cassandra, you've probably solved a lot of the issues you might have.

Example of fairly standard Cassandra bug (don't know if present on latest release, certainly was a year or two ago): When you add a new node to the cluster, it 'bootstraps', where it copies ~1/n the data from other nodes. When you are done bootstrapping, it's copied a bunch of data from other nodes, but the other nodes still contain that data. You then run 'cleanups' on the other nodes to remove the (now stale and unusable) data so as to get your disk space back.

If you accidentally run a cleanup on the new node as it is being bootstrapped, it will succeed, you will delete all the data that's been copied over so far, and Cassandra will _not_ terminate the bootstrap. Everything will be green, but your new node will suddenly be using 0 disk space. When the bootstrap finishes, possibly days later, your cluster will be immediately corrupted due to violated replication guarantees - but only on data that hasn't been read or written over that period, because if it was written it'll be re-replicated, and if it was read Cassandra will silently repair at this time. Repairs resolve the issue, but if you've made this mistake due to scripting, if you get unlucky it's possible to just delete all replicas of some data between repairs.

Example of other Cassandra bug (again, might be outdated): Cassandra nodes identify themselves on startups with IPs, and the owned token ranges are not persisted, they're streamed from other nodes in the cluster. If you've deployed your Cassandra in K8s and you reboot multiple nodes in one go and they swap IPs upon reboot, you may now find yourself in a split brain situation in which nodes magically forget they own certain data ranges and think they own each others data (or maybe it's that the nodes still think they own the right ranges but other nodes think they own the wrong ranges). Wasn't close enough to fully debug that one.

It's a mess. Would seek to avoid problem spaces where I might need to use it again, though if by chance ended up in a space where it made sense, probably wouldn't avoid the tech.

Page 1 of 18Next →