HNHacker News
TopNewBestAskShowJobs

xatalytic

192 karma · joined April 9, 2015

submissionscomments
xatalytic··on Databricks Releases 15K Record Training Corpus for Instruction Tuning LLMs
One of the creators here - yeah, the thing we have our eyes on is the vector not the point.

It’s astounding how adaptable these open models are, even with just a quarter of the Alpaca data. We’re a team of machine learning engineers and hackers, not an AI science lab, but that’s kind of the point frankly - this whole exercise appears to be far easier that it might at first seem.

xatalytic··on Databricks Releases 15K Record Training Corpus for Instruction Tuning LLMs
15,000 instruction tuning records generated by Databricks employees in seven of the behavior categories outlined in the InstructGPT paper (predecessor to ChatGPT). Coincides with the release of Dolly 2.0, which is trained exclusively on this dataset and demonstrates high quality (but not state-of-the-art) instruction-following behavior.

The data and models are licensed for commercial use, setting them apart from recent releases trained on data from OpenAI.

xatalytic··on Show HN: Semantic search for video
I’ll shoot you a line. As it happens I just came across Haystack (6.5k stars on GitHub) which looks like an awesome stack for this class of work.

https://haystack.deepset.ai/

xatalytic··on Show HN: Semantic search for video
Curious if you can share more about the stack. From another comment it sounds like you're using Whisper to generate the text from audio.

My default out of the box way to approach this would be something straightforward like a BERT-alike encoder to embed each target sentence in a FAISS index (hell, podcasts aren't long -- it could be brute force lookup, I suppose) or similar, with the same encoder running on the queries.

Something I've been playing with is Flan-T5 (https://huggingface.co/docs/transformers/model_doc/flan-t5), which has really strong out of the box question answering capabilities. I could see chunking in larger blocks and using the blocks as a context passage and the query as a question-oriented prompt. I've run some fine-tuning experiments with this setup for text generation (e.g. write me a summary of Huberman's key takes on dopamine) and find that the Flan-T5 model forgets a lot of its other capabilities when subject to fine tuning.

In any event, understand if you're not inclined to share, but love talking shop on this stuff.