243 karma · joined February 25, 2022
It's an embeddings database of Wikipedia abstracts with page view data integrated to enable filtering pages based on popularity in addition to similarity.
This is a pretty straight forward problem and a good fit for a standard text classifier as well.
Here is an example of fine-tuning a model with txtai: https://colab.research.google.com/github/neuml/txtai/blob/ma...
For example: SELECT id, text, date FROM txtai WHERE similar('machine learning') AND date >= '2023-03-30'
GitHub: https://github.com/neuml/txtai
This article is a deep dive on how the index format works: https://neuml.hashnode.dev/anatomy-of-a-txtai-index
For those specifically interested in text embeddings, here is a good analysis: https://medium.com/@nils_reimers/openai-gpt-3-text-embedding...
https://towardsdatascience.com/milvus-pinecone-vespa-weaviat...
https://www.tensorflow.org/lite
https://huggingface.co/muhtasham/olm-bert-tiny-december-2022
https://neuml.hashnode.dev/train-a-language-model-from-scrat...
But it doesn't look all that easy to stand up.
Another thread on HN (https://news.ycombinator.com/item?id=34653075) discusses a model that is less than 1B parameters and outperforms GPT-3.5. https://arxiv.org/abs/2302.00923
These models will get smaller and more efficiently use the parameters available.
There's already great local/FOSS options such as FLAN-T5 (https://huggingface.co/google/flan-t5-base). Would be great to see a local model like that trained specifically for chat.
Referenced snippet from the abstract:
With Multimodal-CoT, our model under 1 billion parameters outperforms the previous state-of-the-art LLM (GPT-3.5) by 16% (75.17%->91.68%) on the ScienceQA benchmark and even surpasses human performance.
Answer the following question using only the context below. Say 'no answer' when the question can't be answered. Question: {question} Context: {context}
Not sure this prompt works for all scenarios and models. But it can easily be changed and is a starting point.
More NLP based, but here is an article on an effort to build Transformers micromodels to run on embedded devices. The model in this example is under 1MB. Goal would be to ultimately convert this from ONNX to TFLite.
https://neuml.hashnode.dev/train-a-language-model-from-scrat...
The goal is to use existing high quality TTS models without a heavy install footprint.
This one re-ranks the output from an Elasticsearch index - https://colab.research.google.com/github/neuml/txtai/blob/ma...
The next major release have more examples using the local BM25 scoring module.