214 karma · joined January 28, 2011
In general I find these reports have some gaps but for most topics they provide a reasonable picture of what is going on. I personally prefer them over deep research reports coming out of chatgpt etc. as they are just walls of text that I skip to go to the result table.
For example as somebody interested in coffee this is pretty comprehensive based on my previous research work: https://select.keenable.ai/r/dual-boiler-prosumer-espresso-m...
One key difference is that we’re using LLMs to create structured data on the fly based on the query. That means we don't need to relying on webmasters to explicitly annotate pages with metadata which in my view was one of the main reasons why the semantic web did not work out.
The engine itself is part of the reranked aggregate so if it finds something useful and everybody else does not it gets full credit.
This automatically optimizing for clicks using ML is the main way google and other "human" focused search engines have been improving for 20 years. Not sure what citation is needed here.
The queries from what I can tell are not trivial. The actual github repo of the benchmark has a judgement/query browser where you can inspect the different query streams: https://keenableai.github.io/needle/
We developed this live benchmark with daily/hourly sampled fresh queries matching real agentic search traffic to estimate actual search performance of different AI search providers.
This provides interesting tradeoffs where models can call search much more frequently which we believe is one thing that is not happening today.
Our benchmarks (see https://keenableai.github.io/needle/card-lowshared/ ) also suggest that some players in the market, such as Parallel and Brave, may be using the same underlying search index.
Success and Scale Bring Broad Responsibility
We started in a garage, but we’re not there anymore. We are big, we impact the world, and we are far from perfect. We must be humble and thoughtful about even the secondary effects of our actions. Our local communities, planet, and future generations need us to be better every day. We must begin each day with a determination to make better, do better, and be better for our customers, our employees, our partners, and the world at large. And we must end every day knowing we can do even more tomorrow. Leaders create more than they consume and always leave things better than how they found them.
these actions are exactly the opposite of what they claim to try to strive towards.
Back then neural LMs were just beginning to emerge and we briefly experimented with using one of these unlimited N-gram models for pre-training but never got any results.
How the program works
* We’ll adjust your connected thermostat(s) when summer electricity demand is at its highest to help decrease stress on the grid
* Prior to an event, we may lower your thermostat(s) by a few degrees to help maintain comfort during the actual event
* if the temperature in your home feels uncomfortable, you can opt out of an event at any time by adjusting your thermostat.
* You’ll receive a $60 e-gift card at the end of the season for participating.
Signing up nets you another $125 gift card.
Doesn't LTE already have quite good latency properties?
From looking at the github repo it does look like the system runs entirely in main memory.
- They mentioned (from what I remember) that they now use BitFunnel as they core of the complete Bing search engine not just the fresh parts.
- When I read the paper and looking at the code, it looks like their index doesn't include frequency information whereas your PEF code does. It is unclear what was counted in the experiments.
- If you look at the code, they are actually doing much more complicated stuff than just regular bloom filters by "bin packing" the hash positions for each term to reduce false positive rates (see https://github.com/BitFunnel/BitFunnel/issues/278 ). I'm nor sure if it is "fair" to compare a system developed by 10+ engineers over many years to a "phd student" code base developed over short period of time. I think the PEF code is excellent but I'm more talking about that engineering efforts can have a large impact on performance.
- I'm fairly sure you are right regarding the lack of URL-sorting. However, this can have another cause. If you consider Figure 4 in the paper which shows how "higher ranking rows" group documents together to allow faster intersection. URL sorting causes clusters in document-ids. Say, in the example in Fig. 4 there might be a cluster for that specific term for documents 0,1,2,3. This would mean the "higher ranking" row approach becomes worse (more false positives) when clustering occurs in the collection. So while URL-sorting helps PEF, it will most likely make BitFunnel worse.
here a link for some state-of-the-art benchmarks of non-parallel SA construction algorithms: https://github.com/y-256/libdivsufsort/blob/wiki/SACA_Benchm...
there also now exist linear speedup parallel SA construction algorithms.
"Zhongjun Jin, Michael R. Anderson, Michael J. Cafarella, H. V. Jagadish: Foofah: Transforming Data By Example. SIGMOD Conference 2017: 683-698"
As a somewhat expert in the field I doubt that this paper has any practical implications.