HNHacker News
TopNewBestAskShowJobs

xfalcox

1,038 karma · joined November 7, 2013

[ my public key: https://keybase.io/falcofantastic; my proof: https://keybase.io/falcofantastic/sigs/_PKYsKf2wmCyt834lEh6N4POje9RoICd3Ta7qezTzJE ]
submissionscomments
xfalcox··on Decision models like Jev don't beat LLM-as-a-judge or traditional classifiers
Doesn't the article covers the speed part by showing that Qwen 3.6 35A3B has lower latency and same accuracy?
xfalcox··on vLLM v0.28.0
Yeah, DeepSeek 4 Flash on vLLM has been an adventure indeed. It finally stabilized for me on 2 x H200 using a commit a few days before 0.28, so this release should be good for you.
xfalcox··on Qwen3.8 27B scores 52 on Artificial Analysis
Have you tried running it on a single 5090? Dual 5090 require https://github.com/aikitoria/open-gpu-kernel-modules for higher perf. Are you using TP?
xfalcox··on Designing emoji for the way we communicate today
I'm wondering the same! How that article has no links is beyond me.
xfalcox··on Show HN: Getting GLM 5.2 running on my slow computer
Question to the OP, have you tested this on a machine where the entire model and context fit in RAM ?
xfalcox··on Show HN: Getting GLM 5.2 running on my slow computer
README covers that

https://github.com/JustVugg/colibri#ssd-wear-warning

xfalcox··on Use your Nvidia GPU's VRAM as swap space on Linux
Given my dev machine has 32GB of RAM and 32GB of VRAM that sits mostly idle when I'm not running AI models, this is not that bad of an idea.
xfalcox··on Google releases Gemma 4 open models
Comparing a model you can downloads weights for with an API-only model doesn't make much sense.
xfalcox··on Claude wrote a full FreeBSD remote kernel RCE with root shell
Our CEO did that at our company and found 33 CVEs. Rails also did that and found 7 or 8.
xfalcox··on Show HN: RatatuiRuby wraps Rust Ratatui as a RubyGem – TUIs with the joy of Ruby
I just made a new installer for Discourse on CharmRuby, now I gotta check this out and see if porting is feasible. Hopefully this reduces the app size, that is quite large with CharmRuby
xfalcox··on Tell HN: The Google Tenor GIF API has been shut down
That is a great fit for the GIF integration in Discourse.

I was able to quickly add support for it at https://github.com/discourse/discourse-gifs/pull/107

Love to see WEBP support. Do you plan on adding support for AVIF?

Also, this is used by many Discourse sites, we should talk.

xfalcox··on Sergey Brin's Unretirement
First time I was in San Francisco and someone introduced themselves like that, going even beyond, was indeed a super weird experience being a brazilian.
xfalcox··on Z-Image: Powerful and highly efficient image generation model with 6B parameters
We have vLLM for running text LLMs in production. What is the equivalent for this model?
xfalcox··on 28M Hacker News comments as vector embedding search dataset
I am partial to https://huggingface.co/Qwen/Qwen3-Embedding-0.6B nowadays.

Open weights, multilingual, 32k context.

xfalcox··on Honda: 2 years of ml vs 1 month of prompting - heres what we learned
It's the Amazon own model. I'm baffled someone would pick it, even more that someone would test Llama 4 for a task in an age where Sonnet 4.5 is already out, so in the last 45 days.

Looks like they were limited by AWS Bedrock options.

xfalcox··on The Case Against PGVector
> what does the rag for uploaded files do in discourse?

You can upload files that will act as RAG files for an AI bot. The bot can also have access to forum content, plus the ability to run tools in our sandboxed JS environment, making it possible for Discourse to host AI bots.

> also, when i run a discourse search does it really do both a regular keyword search and a vector search? how do you combine results?

Yes, it does both. In the full page search it does keyword first, then vector asynchronously, which can be toggled by the user in the UI. It's auto toggled when keyword has zero results now. Results are combined using reciprocal rank fusion.

In the quick header search we simply append vector search to keyword search results when keyword returns less than 4 results.

> does all discourse instances have those features? for example, internals.rust-lang.org, do they use pgvector?

Yes, all use PGvector. In our hosting all instances default to having the vector features enabled, we run embeddings using https://github.com/huggingface/text-embeddings-inference

xfalcox··on The Case Against PGVector
We host thousands of forums but each one has its own database, which means we get a sort of free sharding of the data where each instance has less than a million topics on average.

I can totally see that at a trillion scale for a single shard you want a specialized dedicated service, but that is also true for most things in tech when you get to the extreme scale .

xfalcox··on The Case Against PGVector
I was taken back when I saw what was basically zero recall loss in the real world task of finding related topics, by doing the same thing you described where we over capture with binary embeddings, and only use the full (or half) precision on the subset.

Making the storage cost of the index 32 times smaller is the difference of being able to offer this at scale without worrying too much about the overhead.

xfalcox··on The Case Against PGVector
In Discourse embeddings power:

- Related Topics, a list of topics to read next, which uses embeddings of the current topic as the key to search for similar ones

- Suggesting tags and categories when composing a new topic

- Augmented search

- RAG for uploaded files

xfalcox··on The Case Against PGVector
Also worth mentioning that we use quantization extensively:

- halfvec (16bit float) for storage - bit (binary vectors) for indexes

Which makes the storage cost and on-going performance good enough that we could enable this in all our hosting.

xfalcox··on The Case Against PGVector
> Nobody’s actually run this in production

We do at Discourse, in thousands of databases, and it's leveraged in most of the billions of page views we serve.

> Pre- vs. Post-Filtering (or: why you need to become a query planner expert)

This was fixed in version 0.8.0 via Iterative Scans (https://github.com/pgvector/pgvector?tab=readme-ov-file#iter...)

> Just use a real vector database

If you are running a single service that may be an easier sell, but it's not a silver bullet.

xfalcox··on Embedding Text Documents with Qwen3
Depends on your needs. You surely don't want 32k long chunks for doing the standard RAG pipeline, that's for sure.

My use case is basically a recommendation engine, where retrieve a list of similar forum topics based on the current read one. As with dynamic user generated content, it can vary from 10 to 100k tokens. Ideally I would generate embeddings from an LLM generated summary, but that would increase inference costs considerably at the scale I'm applying it.

Having a larger possible context out of the box just made a simple swap of embeddeding models increase quality of recommendations greatly.

xfalcox··on Embedding Text Documents with Qwen3
Just migrated all embeddings to this same model a few weeks ago in my company, and it's a game changer. Having 32k context is a 64x increase when compared with our previous used model. Plus being natively multilingual and producing very standard 1024 long arrays made it a seamless transition even with millions of embeddings across thousands of databases.

I do recommend using https://github.com/huggingface/text-embeddings-inference for fast inference.

xfalcox··on Rerank-2.5 and rerank-2.5-lite: instruction-following rerankers
Having a public tokenizer is quite useful, specially for embeddings. It allows you to do the chunking locally without going to the internet.
xfalcox··on GPT-OSS vs. Qwen3 and a detailed look how things evolved since GPT-2
Qwen 3 is not slow by any metrics.

Which model, inference software and hardware are you running it on?

The 30BA3B variant flies on any GPU.

xfalcox··on Workhorse LLMs: Why Open Source Models Dominate Closed Source for Batch Tasks
You'd be surprised how often people in enterprise can be left waiting months to get an API key approved for an LLM provider.
xfalcox··on Storefront Web Components
This looks like a great fit for allowing people to monetize their Discourse forums, by having partners stores and plugging those instead of ads.

Will build a quick poc integration. How can I contact you with feedback?

xfalcox··on Show HN: A backend agnostic Ruby framework for building reactive desktop apps
This looks super cool, exactly what I've been wanting to create some useful widgets! Thanks for sharing!
xfalcox··on I analyzed chord progressions in 680k songs
I guess one aspect missing here is weighting more popular songs on that analysis.

I assume that the analysis is simply counting every song chords, so a unknown band you've never heard about has the same impact as The Ramones.

I'd like to see the same graph weighted by band popularity using either YouTube or Spotify data.

xfalcox··on Cohere Launches Embed 4
No downloadable open weights ?

Looks like I'll stay on [bge-m3](https://huggingface.co/BAAI/bge-m3)

Page 1 of 11Next →