HNHacker News
TopNewBestAskShowJobs

matt4711

214 karma · joined January 28, 2011

submissionscomments
matt4711··on The car industry A/B tested selling a car with and without CarPlay
I made the same choice. Also did not even consider the Blazer knowing they are identical.
matt4711··on Keenable SELECT: an agent that searches the web in SQL
You need ways to sift through large amounts of unstructured data and do data analysis after. We created these functions to be able to use LLMs to parse unstructured data from the into structured data which we then process with a SQL like engine to map/group/filter etc.
matt4711··on Keenable SELECT: an agent that searches the web in SQL
Thanks for the feedback on the cancel/abort/edit requests. we are working on this already.
matt4711··on Keenable SELECT: an agent that searches the web in SQL
Hey, I'm on the team working on this. In terms of accuracy the best way to gauge the usefulness of these reports is to search for some you are deeply familiar with and then estimate.

In general I find these reports have some gaps but for most topics they provide a reasonable picture of what is going on. I personally prefer them over deep research reports coming out of chatgpt etc. as they are just walls of text that I skip to go to the result table.

For example as somebody interested in coffee this is pretty comprehensive based on my previous research work: https://select.keenable.ai/r/dual-boiler-prosumer-espresso-m...

matt4711··on Keenable SELECT: an agent that searches the web in SQL
Hey! I’m on the team working on this. The semantic web comparison is pretty close to how we think about this.

One key difference is that we’re using LLMs to create structured data on the fly based on the query. That means we don't need to relying on webmasters to explicitly annotate pages with metadata which in my view was one of the main reasons why the semantic web did not work out.

matt4711··on Needle: The benchmark your search engine can't memorize
> Besides the clear AI smell, this nonsensical claim also plainly contradicts the methodology's key evaluation claim that the quality of an engine's results should be measured against how much it overlaps with the reranked aggregate of the other engines. The benchmark thus seemingly values an engine's ability to "answer unanswerable questions" at zero.

The engine itself is part of the reranked aggregate so if it finds something useful and everybody else does not it gets full credit.

matt4711··on Needle: The benchmark your search engine can't memorize
> Yeah? Care to cite anything for that?

This automatically optimizing for clicks using ML is the main way google and other "human" focused search engines have been improving for 20 years. Not sure what citation is needed here.

matt4711··on Needle: The benchmark your search engine can't memorize
The judgements are available on huggingface so training is possible. But given that we evaluate on new queries daily there would need to be some generalization happening for this to show up in the benchmark.
matt4711··on Needle: The benchmark your search engine can't memorize
It is hard to be fair I agree. We tried to be open about what we do here: github.com/keenableai/needle

The queries from what I can tell are not trivial. The actual github repo of the benchmark has a judgement/query browser where you can inspect the different query streams: https://keenableai.github.io/needle/

matt4711··on Needle: The benchmark your search engine can't memorize
One of the authors here. We have been seeing lots of benchmaxxing and leakage in standard web search benchmarks such as BrowseComp.

We developed this live benchmark with daily/hourly sampled fresh queries matching real agentic search traffic to estimate actual search performance of different AI search providers.

matt4711··on Show HN: Keenable – A different web search API for AI agents
If you look at our benchmarks at https://keenableai.github.io/needle/ we are competitive in quality to exa (the market leader) but much cheaper and lower latency.

This provides interesting tradeoffs where models can call search much more frequently which we believe is one thing that is not happening today.

Our benchmarks (see https://keenableai.github.io/needle/card-lowshared/ ) also suggest that some players in the market, such as Parallel and Brave, may be using the same underlying search index.

matt4711··on Amazon backs power plant that may become top source of US climate pollution
I remember a couple of years ago Amazon added a new leadership principle:

Success and Scale Bring Broad Responsibility

We started in a garage, but we’re not there anymore. We are big, we impact the world, and we are far from perfect. We must be humble and thoughtful about even the secondary effects of our actions. Our local communities, planet, and future generations need us to be better every day. We must begin each day with a determination to make better, do better, and be better for our customers, our employees, our partners, and the world at large. And we must end every day knowing we can do even more tomorrow. Leaders create more than they consume and always leave things better than how they found them.

these actions are exactly the opposite of what they claim to try to strive towards.

matt4711··on As AI eats the web, the internet’s collective memory is disappearing
There will be new ways and incentives for content creators to be compensated. Many AI search startups are already talking about this or have created programs that help incentivize content creation.
matt4711··on Infini-Gram: Scaling unbounded n-gram language models to a trillion tokens
A paper [1] we wrote in 2015 (cited by the authors) uses some more sophisticated data structures (compressed suffix trees) and Kneser–Ney smoothing to get the same "unlimited" context. I imagine with better smoothing and the same larger corpus sizes as the authors use this could improve on some of the results the authors provide.

Back then neural LMs were just beginning to emerge and we briefly experimented with using one of these unlimited N-gram models for pre-training but never got any results.

[1] https://aclanthology.org/D15-1288.pdf

matt4711··on Texas Power Companies Are Remotely Raising Temps on Residents' Smart Thermostats
LADWP (LA power provider) has a similar opt-in program for Nest owners with the following conditions:

How the program works

* We’ll adjust your connected thermostat(s) when summer electricity demand is at its highest to help decrease stress on the grid

* Prior to an event, we may lower your thermostat(s) by a few degrees to help maintain comfort during the actual event

* if the temperature in your home feels uncomfortable, you can opt out of an event at any time by adjusting your thermostat.

* You’ll receive a $60 e-gift card at the end of the season for participating.

Signing up nets you another $125 gift card.

matt4711··on ExamSoft's remote bar exam sparks privacy and facial recognition concerns
I find the requirement of using a laptop without being allowed to use an external monitor to view potentially very long documents (half screen!) for hours on a tiny screen to be ridiculous.
matt4711··on Bio-Inspired Hashing for Unsupervised Similarity Search (With John Hopfield)
Looking at the paper, comparing methods that use k bits per hash to a method that uses k*log(m) bits per hash seems unfair and misleading.
matt4711··on Why Chinese Is So Damn Hard (1992)
Funny enough "learning characters/words in the context of the vocabulary they are in" is exactly what NLP machine learning models use to learn "rich" word/text representations based on the "distribution hypothesis" which states that the choice of words in the same context share a common meaning.
matt4711··on Academic Torrents – Making 27TB of research data available
Like the requirement that you have to delete tweets in datasets that have been deleted on twitter?
matt4711··on Academic Torrents – Making 27TB of research data available
I'm pretty sure all the twitter datasets violate the twitter TOCs.
matt4711··on 5G standard is ready: Rel-15 success spans 3GPP groups
From the first link: "Comparing OFDM to LTE today we find a better scalability to a much lower latency (an order of magnitude lower round-trip time [RTT] than LTE today) in OFDM."

Doesn't LTE already have quite good latency properties?

matt4711··on A Shrinking Pie? The IPv4 Transfer Market in 2017
Hetzner is involved in many of the non-related entity transfers.
matt4711··on Nvidia’s New Policy Limits GeForce Data Center Usage
I thought the NVIDIA drivers for the more fancy cards (TITAN etc) are the same as for the gforce cards. Wouldn't this restriction apply to those cards as well? Doesn't make much sense to me...
matt4711··on BitFunnel: Revisiting Signatures for Search [pdf]
> I find it hard to believe this. Their main index is certainly not all-RAM (there must be some flash and maybe even disk), and the throughput would just not be enough for something like BitFunnel.

From looking at the github repo it does look like the system runs entirely in main memory.

matt4711··on BitFunnel: Revisiting Signatures for Search [pdf]
I was at the SIGIR'17 presentation of this paper (won best paper award btw) and have some comments in general:

- They mentioned (from what I remember) that they now use BitFunnel as they core of the complete Bing search engine not just the fresh parts.

- When I read the paper and looking at the code, it looks like their index doesn't include frequency information whereas your PEF code does. It is unclear what was counted in the experiments.

- If you look at the code, they are actually doing much more complicated stuff than just regular bloom filters by "bin packing" the hash positions for each term to reduce false positive rates (see https://github.com/BitFunnel/BitFunnel/issues/278 ). I'm nor sure if it is "fair" to compare a system developed by 10+ engineers over many years to a "phd student" code base developed over short period of time. I think the PEF code is excellent but I'm more talking about that engineering efforts can have a large impact on performance.

- I'm fairly sure you are right regarding the lack of URL-sorting. However, this can have another cause. If you consider Figure 4 in the paper which shows how "higher ranking rows" group documents together to allow faster intersection. URL sorting causes clusters in document-ids. Say, in the example in Fig. 4 there might be a cluster for that specific term for documents 0,1,2,3. This would mean the "higher ranking" row approach becomes worse (more false positives) when clustering occurs in the collection. So while URL-sorting helps PEF, it will most likely make BitFunnel worse.

matt4711··on MTuner is a C/C++ memory profiler and memory leak finder
Is there a binary version of this tool that can be downloaded somewhere? Seems a bit of a pain to install.
matt4711··on Show HN: Fast suffix arrays in Python
There are much faster SA construction algorithms than skew (check out divsufsort). The O(n) algorithms using induced sorting are also likely much faster than this work. The constants of recent O(n) algorithms are very low.

here a link for some state-of-the-art benchmarks of non-parallel SA construction algorithms: https://github.com/y-256/libdivsufsort/blob/wiki/SACA_Benchm...

there also now exist linear speedup parallel SA construction algorithms.

matt4711··on Wikipedia’s Switch to HTTPS Has Successfully Fought Government Censorship
I think china blocks zh.wikipedia.org but all other languages are not blocked.
matt4711··on Transform Data by Example [video]
There is a paper describing such a method (not sure if that is what was implemented):

"Zhongjun Jin, Michael R. Anderson, Michael J. Cafarella, H. V. Jagadish: Foofah: Transforming Data By Example. SIGMOD Conference 2017: 683-698"

matt4711··on Space-Efficient Construction of Compressed Indexes in Deterministic Linear Time
In general the techniques in zstd/lzma are much faster than using compressed text indexes.

As a somewhat expert in the field I doubt that this paper has any practical implications.

Page 1 of 4Next →