192 karma · joined August 14, 2007
https://manual.manticoresearch.com/Searching/KNN#Auto-Embeddings-(Recommended)
and hybrid search: https://manual.manticoresearch.com/Searching/Hybrid_search#Hybrid-search
With these two features you can set up RAG with hybrid search in about the time it takes to install the server and insert the documents.Manticore is really performant and uses far fewer resources than similar search engines like elastic.
I may be over optimistic but "car-dealer speak" sounds like something an AI could be trained on and access to inventory might be a couple of tool calls to the appropriate APIs.
I've done a few other experiments with my foreign speaking friends and it appeared to me that dogs understand the language their owners speak primarily.
https://github.com/cg123/mergekit
you can slice off layers and blend models with different strategies.I think the ones on el Camino actually came a couple years later. One tiny one in front of Safeway and another down south a bit. I think this is the one he ranks as number one.
I really recommend trying anyone of these places if you can. Really the simple ranchero style is delicious and unlike other Mexican food you’ve had. Sorry not a lot of vegetables but you can skip the soda and get an agua Fresca or carrot juice to make up for it.
I had a 1972 Triumph Bonneville which had a "tickler" button on each carb instead of a choke. That meant to start it up you would press each tickler button until a bit of gas shot out invariably on your hand but also the engine and the sometimes hot exhaust. Only after this ritual was performed could you jump on the kickstart (no electric start). So you end up smelling like gas.
Q: Why to the British drink warm beer? A: Because Lucas makes electrics.
A lot of the article is about "you'll never find 2nd" which is in large part to the weird shifting pattern. Reverse is where first usually is, first is where second usually is and the rest are in a kind of off by one pattern from there. This was actually considered a feature since supposedly it allows the driver to make the shift from first to second faster.
The vagueness of the stick is very true and something every 914 owner can relate to. Those first couple weeks you spend some time hunting for the right slot. I went from first to fifth many times before a muscle memory was developed and I didn't have to think about it.
StarSpace is a general-purpose neural model for efficient learning of entity embeddings for solving a wide variety of problems:
Learning word, sentence or document level embeddings.
Information retrieval: ranking of sets of entities/documents or objects, e.g. ranking web documents.
Text classification, or any other labeling task.
Metric/similarity learning, e.g. learning sentence or document similarity.
Content-based or Collaborative filtering-based Recommendation, e.g. recommending music or videos.
Embedding graphs, e.g. multi-relational graphs such as Freebase.
Image classification, ranking or retrieval (e.g. by using existing ResNet features).
In the general case, it learns to represent objects of different types into a common vectorial embedding space, hence the star ('*', wildcard) and space in the name, and in that space compares them against each other. It learns to rank a set of entities/documents or objects given a query entity/document or object, which is not necessarily the same type as the items in the set. https://mxnet.apache.org/versions/1.8.0/api/perlActual code is here
http://www.limerent.com/projects/2020_11_EigenGrandito/
but discussion is surprisingly interesting. https://github.com/scikit-learn-contribWe should start thinking about an ML life cycle were data is ingested, data labeled labeled, models trained, model tested, model deployed and monitored. Rinse, lather, repeat.
Of course a catalog is edited or curated which makes it biased but maybe more human editorial direction isn't such a bad thing in our current world of algorithmic optimized, SEO, echo chamber, click-bait choices.
Browse-ability allows non-focesed search and informs choices that are made at a later time from passive data gathering.
https://www.amazon.com/Cool-Tools-Possibilities-Kevin-Kelly/...
is a worthy successor to Whole Earth Catalog in both content and spirit.
- Consider injecting information with "oracles" An oracle is a kind of virtual user that likes one thing and only one thing. For example they only watch movies that have been tagged sci-fi. This sci-fi oracle adds information about sci-fi-ness to your data which is useful for several things. It helps with the cold start problem as new items can be automatically tagged by the appropriate oracles and get past the zero information horizon quickly. Also you can measure a users sci-fi affinity by measuring that users similarity to the sc-fi oracle.
- Another way to think about co-occurrences is as connected nodes in a digraph. You have users and items and connections between them (user watched video). Start with an item and traverse all the links to the other side (all the users who watched this video) then for each user traverse to the items side (you can roll up the occurrences for a score) and you have similar items. Works equally as well for finding similar users.
- Create an "average user" and use that as a seed for new users. If we know nothing else we should expect a new user to be close to average. This means they will probably get recommended the most popular items but
- Find items with divisive scores or groups and ask new users their opinion on those items to find out about them. After a new user gets created consider asking them their opinion on five of these divisive items. Their ratings should swiftly put them in an informed space the way taking five steps down a binary tree does a lot to reduce search space.
- I like the way you use simple plus one smoothing for your scores. I'm not sure why this doesn't get used more often.
Good luck with the project!
Move these to the root of the volume you are backing up and burn dvd or whatever.
You can keep these indexes together on your live volume somewhere and search them all together
Glimpse indexes text files. Other files require preprocessing. There are solutions for this too but it’s more complicated.