HNHacker News
TopNewBestAskShowJobs

kkielhofner

4,538 karma · joined February 12, 2013

Atomic Canyon: https://atomic-canyon.com.

E-mail: kris [@] $DOMAIN

submissionscomments
kkielhofner··on Understanding the BM25 full text search algorithm
I've used it for hybrid search and it works quite well.

Overall I'm really happy to see Typesense mentioned here.

A lot of the smaller scale RAG projects, etc you see around would be well served by Typesense but it seems to be relatively unknown for whatever reasons. It's probably one of the easiest solutions to deploy, has reasonable defaults, good docs, easy clustering, etc while still be very capable, performant, and powerful if you need to dig in further.

kkielhofner··on Netflix buffering issues: Boxing fans complain about Jake Paul vs. Mike Tyson
> I'm an engineering manager

How are you involved in the hiring process?

> Our engineers are fucking morons. And this guy was the dumbest of the bunch.

Very indicative of a toxic culture you seem to have been pulled in to and likely have contributed to by this point given your language and broad generalizations.

Describing a wide group of people you're also responsible for as "fucking morons" says more about you than them.

kkielhofner··on Serving 15 terabytes of 4K video for $2.18
They will.

- It's a cornerstone of their brand.

- R2 users are at least paying for something (storage).

- Their network is massively overbuilt to be able to absorb DDoS attacks.

- They offer free bandwidth with their CDN - including completely free users. These resources need to get fetched from origin and/or cached. R2 doesn't have to fetch from origin which eliminates the bandwidth required for the fetch.

Large providers typically pay very little/nothing for bandwidth other than equipment costs (ports, etc). As a large provider they have free peering to most of the "last mile" ISP/eyeball networks in the world. This benefits all parties because these ISPs don't have to pay transit providers and neither does Cloudflare and it's faster. Same goes for all of the big clouds.

People who think AWS/GCP/Azure/etc bandwidth pricing of $0.12/GB (or whatever) is fair/reasonable have no idea what bandwidth actually costs these operators. As noted it's effectively nothing and the big clouds capitalize on this ignorance by charging insane markups for bandwidth.

kkielhofner··on AMD outsells Intel in the datacenter space
ICC, IPP, QAT, etc are definitely an edge.

In AI world they have OpenVINO, Intel Neural Compressor, and a slew of other implementations that typically offer dramatic performance improvements.

Like we see with AMD trying to compete with Nvidia software matters - a lot.

kkielhofner··on Colorado scrambles to change voting-system passwords after accidental leak
> The US voting machines are just waiting to be hacked, just a matter of when, not if.

The US election system is very distributed and fragmented - there is virtually no standardization.

Even in the tightest margins for something like President you'd need to have seriously good data to figure out which random municipality voting system(s) you'd need to target to actually affect the outcome.

kkielhofner··on Embeddings are underrated
> Do you mind sharing why you chose SPLADE-esque sparse embeddings?

I can provide what I can provide publicly. The first thing we ever do is develop benchmarks given the uniqueness of the nuclear energy space and our application. In this case it's FermiBench[0].

When working with operating nuclear power plants there are some fairly unique challenges:

1. Document collections tend to be in the billions of pages. When you have regulatory requirements to extensively document EVERYTHING and plants that have been operating for several decades you end up with a lot of data...

2. There are very strict security requirements - generally speaking everything is on-prem and hard air-gapped. We don't have the luxury of cloud elasticity. Sparse embeddings are very efficient especially in terms of RAM and storage. Especially important when factoring in budgetary requirements. We're already dropping in eight H100s (minimum) so it starts to creep up fast...

3. Existing document/record management systems in the nuclear space are keyword search based if they have search at all. This has led to substantial user conditioning - they're not exactly used to what we'd call "semantic search". Sparse embeddings in combination with other techniques bridge that well.

4. Interpretability. It's nice to be able to peek at the embedding and be able to get something out of it at a glance.

So it's basically a combination of efficiency, performance, and meeting users where they are. Our Fermi model series is still v1 but we've found performance (in every sense of the word) to be very good based on benchmarking and initial user testing.

I should also add that some aspects of this (like pretrained BERT) are fairly compute-intense to train. Fortunately we work with the Department of Energy Oak Ridge National Laboratory and developed all of this on Frontier[1] (for free).

[0] - https://huggingface.co/datasets/atomic-canyon/FermiBench

[1] - https://en.wikipedia.org/wiki/Frontier_(supercomputer)

kkielhofner··on Embeddings are underrated
> word Wood dominated the embedding values, but these were supposed to go into 2 different categories

When faced with a similar challenge we developed a custom tokenizer, pretrained BERT base model[0], and finally a SPLADE-esque sparse embedding model[1] on top of that.

[0] - https://huggingface.co/atomic-canyon/fermi-bert-1024

[1] - https://huggingface.co/atomic-canyon/fermi-1024

kkielhofner··on Embeddings are underrated
> we don't want to hurt performance on other real-world tasks just to do well on MTEB

Nice!

Fortunately MTEB lets you sort by model parameter size because using 7B parameter LLMs for embeddings is just... Yuck.

kkielhofner··on Embeddings are underrated
LLMs have nearly completely sucked the oxygen out of the room when it comes to machine learning or "AI".

I'm shocked at the number of startups, etc you see trying to do RAG, etc that basically have no idea what they are, how they actually work, etc.

The "R" in RAG stands for retrieval - as in the entire field of information retrieval. But let's ignore that and skip right to the "G" (generative)...

Garbage in, garbage out people!

kkielhofner··on Embeddings are underrated
My startup (Atomic Canyon) developed embedding models for the nuclear energy space[0].

Let's just say that if you think off-the-shelf embedding models are going to work well with this kind of highly specialized content you're going to have a rough time.

[0] - https://huggingface.co/atomic-canyon/fermi-1024

kkielhofner··on Embeddings are underrated
> they're not a complete replacement for simpler methods like BM25

There are embedding approaches that balance "semantic understanding" with BM25-ish.

They're still pretty obscure outside of the information retrieval space but sparse embeddings[0] are the "most" widely used.

[0] - https://zilliz.com/learn/sparse-and-dense-embeddings

kkielhofner··on AMD Will Need Another Decade to Try to Pass Nvidia
Jensen has said for years that 30% of their R&D spend is on software. Needless to say as they continue to crush it financially this number continues to completely race past AMD.

Turns out people don’t actually want GPUs, they want solutions that happen to run best on GPUs. Nvidia understands that, AMD doesn’t.

Lisa Su keeps talking about “chips chips chips” and MAYBE “Oh btw here’s a minor ROCm update”. Meanwhile, Nvidia continues to masterfully execute deeper and wider into overall solutions and ecosystems - a substantial portion of which is software.

Nvidia is at the point where they’re eating the entire stack. They do a lot of work on their own models and then package them up nice and tight for you with NIM and Nvidia AI Enterprise. On top of stuff like Metropolis, RIVA, countless things. They even have a ton of frameworks to ingest/handle data, finetune/train, and then deploy via NIM.

Enterprise customers can be 100% Nvidia for a solution. When Nvidia is the #1-#2 most valuable company in the world “no one ever got fired for buying Nvidia” hits hard.

The people who say “AMD and Nvidia are equal - it’s all PyTorch anyway” have no view of the larger picture.

With x86_64, day one you could take a drive out of an Intel system, put it in an AMD system, and it would boot and run perfectly. You can still do that today unless you build something for REALLY specific/obscure CPU instructions.

Needless to say that’s not the case with GPUs and a lot of people that make the AMD vs Intel comparison don’t seem to understand that.

kkielhofner··on Strava was used to locate the most powerful people
Shouldn't be much of a surprise, this made news back in 2018 when the same was realized with soldiers and secret military bases:

https://www.theguardian.com/world/2018/jan/28/fitness-tracki...

kkielhofner··on OpenAI builds first chip with Broadcom and TSMC, scales back foundry ambition
For reference seven trillion dollars is 25% of US GDP.

Yeah, that's um, wild.

kkielhofner··on People Are Sick and Tired of All Their Subscriptions
Living in a climate (Wisconsin) that has extreme highs and lows my understanding is this is typically intended to smooth-out a gas bill (as one example) moving from $10/mo in the summer to $400/mo in the winter.

It’s a budgeting thing.

kkielhofner··on GenAI is set to create a mountainous increase in e-waste
As one example take a peek at /r/LocalLLaMA[0] (I suspect you know). These people are snapping up anything and everything they can get their hands on at a reasonable price.

To your point on the P40, it's an eight year old card but fortunately Nvidia has a history of long term support (especially for "datacenter" GPUs). The Pascal series is still fully supported by the latest Nvidia driver and CUDA releases, and projects like llama.cpp are still fairly regularly adding performance optimizations for even Maxwell series GPUs!

Current V100/A100/H100/etc hardware families are not going to end up as e-waste anytime soon. In fact, compare used pricing (and demand) of GPUs to CPUs, RAM, disk, motherboards, etc from eight years ago... That hardware ends up in the trash/at e-waste recyclers much, much sooner (even with /r/homelab).

[0] - https://old.reddit.com/r/LocalLLaMA/

kkielhofner··on Dramatic drop in marijuana use among U.S. youth over a decade
> If you get wasted on anything, or do anything silly, or act weird, someone will pull out a phone and video you.

Anecdotally this seems to be the key impact. With social media almost everyone now has a "brand" and that brand is typically not supported with a post of you out of it, sloppy, etc.

Along those lines, there also seems to be MUCH more emphasis on health - granted superficial health (looking good) but health nonetheless. Fortunately the standards for "ideal beauty" for women especially have shifted from the 90s/2000s no-such-thing-as-too-thin dangerous and extremely unhealthy to a physique that is well-muscled and actually healthy (while being inclusive of different body types).

When I'm at the gym and the high school/college kids show up I just can't believe their level of physical fitness and development. Self-selecting given it's the gym but when I was in high school (class of 2002) the most fit kid on the football, basketball, track, volleyball, etc teams would look out of shape next to what appears to be the "average" gym-goer of this generation. The numbers also seem to be quite a bit higher - there are A LOT of these kids hitting it really hard in the gym.

Needless to say this clearly obsessive-level focus and work is not supported by using drugs like marijuana and alcohol. If nothing else having a lot of followers is much more important and "cool".

If anything I'm more interested in usage statistics of steroids and other performance-enhancing drugs. Some of the physiques, performance, etc I see just don't seem possible to achieve naturally at 16-25.

kkielhofner··on Geico repatriates work from the cloud, continues ambitious infra overhaul
> That so many people working in software don’t have deep hardware expertise or are not familiar with data centers plays to that hand.

I like to remind myself that AWS is 20 years old. That's an entire generation of people from devs to C-Suite that likely don't know anything else. For many of these people all they know about hardware is their laptop. All they know about bandwidth is what they pay their local ISP. All they know about storage is (maybe) USB flash drives and what Apple charges for 256GB vs 512GB.

This is not a criticism. More of a reality check to myself and others that at this point not a lot of people outside of bigger cloud and ISPs know what an Autonomous System is or what buying transit and peering costs. Nor have they bought a cabinet of bare metal or talked to a co-lo provider.

> Not a criticism, just an observation from my experiences.

Exactly!

kkielhofner··on Understanding Round Robin DNS
TTL isn't universally respected. Consider the following path:

Your machine -> Local router -> Configured upstream DNS Server (ISP/CF/Quad8/etc) -> ? -> Authoritative DNS Server

Any one of those layers can override/mess with/cache in a variety of ways including TTL. This is why Cloudflare and a variety of other providers use IP anycast. They accepted DNS for what it is and worked around it.

Not only is the IP always the IP, the "global" BGP routing table actually universally and consistently updates much faster than DNS. Then whatever routers, machines, etc downstream from that don't matter.

kkielhofner··on Geico repatriates work from the cloud, continues ambitious infra overhaul
> Dell/EMC says "Hey, here is drive replacement." We do it, 2 hours later, the volume is knocked offline. Apparently, there was mismatch between backplane version, drive version and through some weird edge case, it knocked the volume offline. Yes, they fixed it, no it wasn't pretty since a bunch of applications had to be recovered.

Anecdotal (as is my position). I can theoretically understand this happening but not only have I never seen it, such an issue would need to be escalated. That's a "this is unacceptable" high-level phone call. A call you more than likely have a chance of someone in actual authority answering because IME unless you have SERIOUS spend with big cloud you'll be lucky to make it a rung or two up sales/support.

Plus backups and redundancies that should prevent even the failure of a chassis/storage/etc from being a significant critical issue.

> their failures tend to be you twiddling your thumbs vs hair on fire on phone with the vendor trying to get it resolved

As a Founder/CTO I have the opposite take - put me and my team in a position to /do something/ vs sitting around waiting for AWS to come back whenever it decides to and while they obscure comms, don't update the fake status dashboards, etc. Meanwhile you're telling your customer "Ummm, we don't know - Amazon has a problem. When it comes back I guess it's back".

Coming from a background of telecom, healthcare, and nuclear energy I can't believe that even flies.

kkielhofner··on Geico repatriates work from the cloud, continues ambitious infra overhaul
> Also, anyone in this industry long enough has been around for "Oh, we will just replace that broken piece of hardware" that ended up "WHY IS EVERYTHING ON FIRE?" because versions didn't match up, hardware was rejected

I've been doing this for 25 years and I'm not sure what this means. Dell isn't going to come back to you and say "sorry but we can't fix this". With the warranty SLA worst case scenario they'll just replace the entire machine if they have to although I don't remember ever seeing it come to that.

> just plain "Actually, THAT failure mode isn't redundant."

When it comes down to it similar issues exist with clouds - regions, availability zones, etc. Big clouds have had multiple widespread outages just this year[0].

From that reference you can see that MS and Amazon themselves struggle to design, build, and run solutions for their own products in their own clouds.

It's always interesting to see marquee household name companies/products/solutions go down when US-East (or whatever) is having a bad day again.

Cloud can be a lot of things but a silver bullet for reliability and uptime isn't one of them.

[0] - https://www.forbes.com/sites/emilsayegh/2024/07/31/microsoft...

kkielhofner··on Geico repatriates work from the cloud, continues ambitious infra overhaul
> They negotiate network and power contracts at a scale that exceeds any typical Fortune 500 company.

..and then mark it up. AWS overall has 38% operating margin[0]. Depending on your application this can hit you really hard (cloud egress bandwidth being an especially obscene offender).

> I'm skeptical that running your own data center will end up a cost saver in the long run.

It's not cloud -or- your own Azure-scale datacenter. There are any number of approaches in between including hybrid to offload stuff like CDN, storage, edge services, etc to cloud but the fact remains many companies can run the entire business from a few beefy machines in co-location facilities. Most companies, solutions, etc are not actually Google, Snapchat, Geico, etc scale and never will be.

Throw in some minor accounting tricks like leasing (with or without Section 179) and these kinds of "creative" approaches are often impossible to beat from a pricing/performance and even uptime standpoint. That's certainly been my experience.

[0] - https://www.theinformation.com/articles/why-aws-fat-margins-...

kkielhofner··on Geico repatriates work from the cloud, continues ambitious infra overhaul
> With a part time person you will see downtime when a machine fails

If a hardware failure causes downtime you're doing it wrong. Additionally, big cloud scaring people from hardware with marketing and FUD has been very effective. Modern hardware is insanely reliable and performant - I don't think I've seen a datacenter/enterprise NVMe drive fail yet. It's not 2005 with spinning disks and power supplies blowing up left and right anymore.

> With 1 person that person will sometimes be on vacation when a zero day takes you down. With 2 people 1 will be on vacation when the second gets sick. You end up needing at least 5 people before you have enough people that you have redundancy for humans issues and the ability to train people in whatever is the latest needed.

Hardware vendors (Dell, etc) have highly-discounted warranty services. In the event of a hardware failure you open a ticket and they dispatch someone directly to the facility (often within hours by SLA) and it gets handled.

Same thing for shipping HW directly to co-lo and they rack/cable/bootstrap for a nominal fee, remote hands for weird edge-cases, etc.

A lot of takes here and elsewhere seem to be either big-cloud or Meta-level datacenter. I have operated POPs in a dozen co-location ("datacenter") facilities (a cabinet or two each) no one on staff ever stepped foot in with hardware we owned (and/or financed) that no one ever saw or touched. We operated this with two people looking after it as part of their broader roles and responsibilities and frankly they didn't have much to do.

There is an entire industry that provides any number of highly flexible and cost-effective approaches for everything in between.

kkielhofner··on Fearless SSH: Short-lived certificates bring Zero Trust to infrastructure
It’s a joke from a famous moment in HN history:

https://news.ycombinator.com/item?id=9224

kkielhofner··on Ask HN: I have 24 core server with 1TB of DDR4 RAM, what should I run?
> what model can i run on 1TB

With 1TB of RAM you can run nearly anything available (405B essentially being the largest ATM). Llama 405B in FP8 precision fits in H100x8 which is 640GB VRAM. Quantization is a very deep and involved well (far too much for an HN comment).

I'm aware it "works" but I don't bother with CPU, GGUF, even llama.cpp so I can't really speak to it. They're just not even remotely usable for my applications.

> tokens per second

Sloooowwww. With 405B it could very well be seconds per token but this is where a lot of system factors come in. You can find benchmarks out there but you'll see stuff like a very high spec AMD EPYC bare metal system with very fast DDR4/5, tons of memory channels, etc doing low single-digit tokens per second with 70B.

> ill get a GPU too but not sure how much VRAM I need for the best value for your buck

Most of my experience is top-end GPU so I can't really speak to this. You may want to pop in at https://www.reddit.com/r/LocalLLaMA/ - there is much more expertise there for this range of hardware (CPU and/or more VRAM limited GPU configs).

kkielhofner··on LLMD: A Large Language Model for Interpreting Longitudinal Medical Records
Have you seen/heard of Abridge[0]? Long story short their secret sauce comes in two main forms:

1. Accurate speech rec, diarization, etc to record a clinician-patient encounter. No notes, no scribes, no "physician staring at Epic when they should be looking at and talking to you".

2. Parsing of transcripts to correctly and accurately populate the patient EHR record - including various structured fields, etc.

Needless to say you're in this space so I don't have to tell you - every Epic/Cerner install is basically a snowflake so there's a lot going on here, especially at scale.

[0] - https://www.abridge.com/

kkielhofner··on Running an open source app: Usage, costs and community donations
Yes Cloudflare and all of that but they’ll do it for free.

Then you get to determine gains you may get from caching and other potential optimizations from one of the best eyeball connected providers in the world. Oh plus the ability to fend off the largest DDoS attacks ever seen.

Cloudflare tunnels enable you to do all of this through an encrypted tunnel without exposing the machine/services to the internet at all. Cloudflare will still MITM all traffic but so does Hetzner (obviously). At least with the tunnel the connection is persistent so you don’t incur TLS handshaking, etc CPU overhead with each client connection.

Bonus points - you can move hosting providers without any hassle, configure hosting provider redundancy (Hetzner + whoever), all of that stuff.

kkielhofner··on LLMD: A Large Language Model for Interpreting Longitudinal Medical Records
An often-ignored/forgotten/unknown fact about utilizing LLMs is that you really need to develop your own benchmark for your specific application/use-case. It’s step 1.

“This model scores higher on MMLU” or some other off-the-shelf benchmark may (likely?) have essentially nothing to do with performance on a given specific use-case, especially when it’s highly specialized.

They can give you a general idea of the capabilities of a model but if you don’t have a benchmark for what you’re trying to do in the end you’re flying blind.

kkielhofner··on One of Florida's most lethal python hunters
They get paid per verified kill:

https://www.sfwmd.gov/our-work/python-program

No idea how it actually works but shooting them where they are has to be tough to verify.

kkielhofner··on Ask HN: I have 24 core server with 1TB of DDR4 RAM, what should I run?
llama.cpp and others can run purely on CPU[0]. Even production grade serving frameworks like vLLM[1].

There are a variety of other LLM inference implementations that can run on CPU as well.

[0] - https://github.com/ggerganov/llama.cpp?tab=readme-ov-file#su...

[1] - https://docs.vllm.ai/en/v0.6.1/getting_started/cpu-installat...

Page 1 of 34Next →