HNHacker News
TopNewBestAskShowJobs

trengrj

1,241 karma · joined August 14, 2013

applied research @ weaviate https://x.com/trengrj
submissionscomments
trengrj··on I’m joining OpenAI
This is very true but I think there is an incredibly long tail of people who "can't survive a terminal session" and I actually question if a terminal ui will win out long term.
trengrj··on Cloudflare outage on November 18, 2025 post mortem
Classic combination of errors:

Having the feature table pivoted (with 200 feature1, feature2, etc columns) meant they had to do meta queries to system.columns to get all the feature columns which made the query sensitive to permissioning changes (especially duplicate databases).

A Crowdstrike style config update that affects all nodes but obviously isn't tested in any QA or staged rollout strategy beforehand (the application panicking straight away with this new file basically proves this).

Finally an error with bot management config files should probably disable bot management vs crash the core proxy.

I'm interested here why they even decided to name Clickhouse as this error could have been caused by any other database. I can see though the replicas updating causing flip / flopping of results would have been really frustrating for incident responders.

trengrj··on Muvera: Making multi-vector retrieval as fast as single-vector search
We added Muvera to Weaviate recently https://weaviate.io/blog/muvera and also have a nice podcast on it https://www.youtube.com/watch?v=nSW5g1H4zoU.

When looking at multi-vector / ColBERT style approaches, the embedding per token approach can massively increase costs. You might go from a single 768 dimension vector to 128 x 130 = 16,640 dimensions. Even with better results from a multi-vector model this can make it unfeasible for many use-cases.

Muvera, converts the multiple vectors into a single fixed dimension (usually net smaller) vector that can be used by any ANN index. As you now have a single vector you can use all your existing ANN algorithms and stack other quantization techniques for memory savings. In my opinion it is a much better approach than PLAID because it doesn't require specific index structures or clustering assumptions and can achieve lower latency.

trengrj··on Australian Online Safety Amendment (Social Media Minimum Age) Bill 2024
As someone living in Australia this law is ridiculous and basically outlaws a chunk of my childhood where I used forums, irc, and blogs.

As a parent I do share concerns for short duration video content like Tiktok, Reels, and Youtube shorts etc, but I think any sensible regulation there would be better suited to everyone.

Age verification has such terrible consequences to anonymity that it shouldn't be an option. From the Explanatory Memorandum it looks like the current eSafety Commissioner was involved which unfortunately explains a lot.

trengrj··on OpenAI is set to lose $5B this year
I'd argue that the current business model is clear (pay subscription for chat or api) but the valuations are unclear.

If open models continue to keep up and industry as a whole just keeps improving on benchmarks fairly slowly and evenly (a big assumption but what has been happening over past year), then you would assume valuations would eventually drop. Saying that there is still a lot of easy value left on the table (great voice integration etc) and a lot of room for innovations to be incorporated in other industries.

Additionally each foundation LLM provider has a non-zero chance of a large step change in performance. If they can keep the value of this step change it would have huge upsides but this is also hard to reason about (what is the valuation of private companies and the share market if AGI exists - does this even matter anymore).

trengrj··on Do we think about vector storage wrong?
Hey @rvrs, I work on Weaviate and we are doing some improvements around increasing write throughput:

1. gRPC. Using gRPC to write vectors has had a really nice performance boost. It is released in Weaviate core but here is still some work on do on the clients. Feel free to get in contact if you would like to try it out.

2. Parameter tuning. lowering `efConstruction` can speed up imports.

3. We are also working on async indexing https://github.com/weaviate/weaviate/issues/3463 which will further speed things up.

In comparison with pgvector, Weaviate has more flexible query options such as hybrid search and quantization to save memory on larger datasets.

trengrj··on On Hybrid Search with Qdrant
Adding a cross-encoder to your app means including pytorch/transformers and a model as a dependency. For people using OpenAI or Cohere embedding apis and lightweight infrastructure this can be a big pain.
trengrj··on On Hybrid Search with Qdrant
I work at Weaviate, a few comments on why we implemented hybrid search [1].

- Using two separate systems for traditional BM25 and vector search and keeping them in-sync is pretty difficult from an operations perspective. A combined system is much easier to manage and will have better end-to-end latency.

- For combining scores, a linear combination like this article suggests is not recommended and instead rank fusion https://rodgerbenham.github.io/bc17-adcs.pdf (where you care what each method ranks first rather than the absolute score) is used.

- The point of adding both search methods is for dealing with what researchers term "out of domain data". This is for datasets the model producing the vectors was not trained on. Research from Google https://arxiv.org/abs/2201.10582 suggests hybrid search with rank fusion helps in this case by around 20.4%. For "in domain" data, the model (usually transformer based) will out perform BM25.

- Using a cross encoder [2] is a good component to add to improve relevance. It will just though rerank the final results, so if the initial search returns 100 garbage results the cross encoder won't be able to help.

[1] https://weaviate.io/blog/hybrid-search-explained [2] https://www.sbert.net/examples/applications/cross-encoder/RE...

trengrj··on Tell HN: Gitlab Premium pricing increases incoming $19 to $29
There are two reasons that self-hosted is expensive, both having no relation to real costs.

1. If you are self-hosting, you have now segmented yourself into the customer category with the government, military, and large corporates concerned about security. This segment will pay more and so will be charged more. In an enterprise software company the vast majority of revenue comes from the larger customers vs many small customers, so these are the users that pricing gets optimised for.

2. Gitlab wants to be seen in the market as a SaaS company as subscription revenue is far preferable to licence revenue from a churn perspective, and also from a general trend perspective where self-hosted, hard to maintain solutions get replaced by cloud solutions.

trengrj··on Postman is limiting local collection runner to 25 runs for basic plans
Does it work well with org-babel?
trengrj··on Elastic, Loki and SigNoz – A Perf Benchmark of Open-Source Logging Platforms
I'm not a fan of competitors creating benchmarks like this as when faced with any tuning decision, they will usually pick the one the makes their competitors slower. But anyway lets take a look at how they tuned Elasticsearch.

Disclaimer I used to work at Elastic!

- Used Logstash instead of Beats for simple task of reading syslog json data. Beats (https://www.elastic.co/guide/en/beats/filebeat/current/fileb...) would have performed better especially around resource usage.

- Set very low Logstash heap of 256mb https://github.com/SigNoz/logs-benchmark/blob/0b2451e6108d8f...

- Added grok processor https://github.com/SigNoz/logs-benchmark/blob/0b2451e6108d8f... Dissect is faster here

- No index template configuration This would cause higher disk usage than needed due to duplicate mappings. Again a Logstash vs Beats thing. For this test more primary shards and a larger refresh interval would also improve things.

- Graph complaining Elasticsearch using 60% available memory. This is as configured, they could use less with not much impact to performance.

- Document counts do not match.. This is probably due to using syslog with random generated data vs creating a test dataset on disk and reading the same data into all platforms.

- Aggregation queries were not provided in repo https://github.com/SigNoz/logs-benchmark so cannot validate.

I'm actually surprised Elastic did so well in this benchmark given the misconfiguration.

trengrj··on Ask HN: What is your home networking setup?
My setup is the following:

1. Ubiquiti ER-X

2. Ubiquiti AP Lite x 2 (upstairs and downstairs)

3. Inbuilt Cat6 cabling in house

The benefit of this approach is I get to use prosumer hardware but at reasonable cost (total < $350 AUD). For the AP's I have just setup via pairing on the mobile app and use the same SSID and passwords which allows for easy roaming.

I'm contemplating upgrading to an Ubiquiti dream machine pro to replace the ER-X for more ports and ability to have video recording & security cameras but really happy with current setup from a wifi performance and stability perspective.

trengrj··on A shark mystery millions of years in the making
Great white sharks are so afraid of killer whales that they will leave a feeding area to them for over a year. The killer whales often kill a shark to only eat their liver.

Evolution of cetaceans started around 50 million years ago https://en.m.wikipedia.org/wiki/Evolution_of_cetaceans. You could probably make a solid argument that sharks being displaced as apex predators in the sea would have caused the die out of many species. Perhaps a crossover occurred around 19 million years ago.

trengrj··on QNAP ships NAS backup software with hidden credentials
If you want a small NAS in a similar form factor I'd recommend Helios64 5-bay NAS https://kobol.io/. It is an Arm64 board runs mainline Armbian. Also comes with 2.5Gbit networking and a built in UPS battery.

I don't understand why people who care about security and have linux knowledge would use Synology/QNAP. They are both proprietary, often exposed to the internet, and packed full of so many features that they are consistently full of vulnerabilities (SynoLocker/QLocker etc).

trengrj··on Response from Greg Kroah-Hartman to University of Minnesota Researchers
His supervisor published the hypocrite paper with the introduced vulnerabilities so they are far more linked than just both being at UMN.
trengrj··on Pulsar vs. Kafka
With Pulsar vs Kafka, I don't see a huge argument between either one functionality wise as they have so much in common (distributed log, Java based, avoid copying memory, use Zookeeper). Because Kafka is more supported and well-known it seems Pulsar needs to be an order of magnitude more performant to capture developer mindshare.

I see the same with Spark vs Flink in that similarities outweigh differences. I wonder if this is some sort of emergent pattern in open source software.

trengrj··on Installing FreeNAS on my QNAP TS-459
Rather than waste time using one of the proprietary NAS systems like QNAP and Synology, why not support a system that runs and supports unmodified Linux and publishes full schematics like the upcoming Kobol64 https://kobol.io/?
trengrj··on Western Digital Lawsuit for Shipping Slower SMR Hard Drives Including WD Red NAS
https://www.youtube.com/watch?v=nGcwVYbkePs is very informative around the reasons for SMR and the problems around them.
trengrj··on Ask HN: What website, from your early days on the net, do you miss?
43 Things https://en.wikipedia.org/wiki/43_Things
trengrj··on Moving Away from Gmail
I'm really looking forward to what DHH can do with https://hey.com/. My aim with an email provider is that I trust it. A non-bootstrapped company with a founder outspoken against privacy issues seems a good place to try.
trengrj··on To see off piranha this fish has strong armour
By ”general English” I assume you are meaning American English?
trengrj··on Building a home NAS on a shoestring budget with the Rock64 SBC (2018)
I shudder at the idea of using USB connectors for important data. I have recently purchased a Kobol NAS https://kobol.io/. It is amazing because it has the following features in a good price point:

1. Open source

2. 2GB ECC ram

3. 4 x SATA

4. Hackable GPIOs

trengrj··on In Chrome 71: Signed HTTP Exchanges
If I'm reading this right, does this mean that current Google AMP powered pages will now be able to impersonate the url of the original publisher site in Chrome provided the publisher performs a key exchange?
trengrj··on Thelio – System76
> Why do they keep marketing Ubuntu as a different OS? The is confusing to people who don't realize that it's literally just a customized Debian.

Why do they keep marketing Debian as a different OS? It is literally just Linux and some GNU system components cobbled together.

trengrj··on The Single Board Computer Database
The best I have found so far is Helios https://kobol.io/helios4/. Unfortunately I missed the second shipment.
trengrj··on The Single Board Computer Database
Quantity of SATA ports would be great as well. It is far too difficult to find an small ARM board with 4 SATA ports.
trengrj··on Apple Supplier List – Top 200 [pdf]
Is it telling that supermicro is not on this list?
trengrj··on Ask HN: Data diff tool for tabular data?
If you store your data in a slowly changing dimensional format you will be able to create a table where you can query the data as it was at any point of time.

The easiest way to do this is to use an ETL tool as hand coding this can be a bit hairy. You can use an open source etl tool like Talend for this as it has inbuilt Postgres or MySQL scd modules.

trengrj··on Five Eyes' Statement of Principles on Access to Evidence and Encryption
If you look at Australia on a map, the premise of a NBN for everyone was definitely going to be difficult but I'd argue that the project hasn't failed. I have NBN at home and it works well. I think we will look back in 10 years with a much softer view on the project.
trengrj··on Cloud9Trader – Simple, powerful platform for algorithmic trading
You unfortunately might hit some trademark issues with AWS's Cloud9 https://en.wikipedia.org/wiki/Cloud9_IDE.
Page 1 of 7Next →