1,176 karma · joined August 13, 2012
Email me at hn@vikas.sh, or check out my work at https://www.vikas.sh.
Nougat is a great model, and converts a lot of PDFs very well. I just wanted something faster, and more generalizable.
I've been making a course that has a little less theory, and a little more application here - https://github.com/VikParuchuri/zero_to_gpt . Videos are all optional (cover the same content as the text).
It balances theory and code, and builds from the foundation up, so you're never typing something without understanding it. Teaching method is text, diagrams, and code. Most lessons have optional videos, too.
It focuses on text models over image models (rnn, transformer, etc).
It's not 100% finished, but has enough to get you very far.
- They compare the performance of this model to the worst 7B code llama model. The base code llama 7B python model scores 38.4% on humaneval, versus the non-python model, which only scores 33%.
- They compare their instruct tuned model to non-instruct-tuned models. Instruction tuning can add 20% or more to humaneval performance. For example, WizardLM 7B scores 55% on humaneval [1], and I've trained a 7B model that scores 62% [2].
- For another example of instruction tuning, Stablecode instruct tuned benchmarks at 26%, not the 20% they cite for the base model [3]
- Starcoder, when prompted properly, scores 40% on humaneval [4]
- They do not report their base model performance (as far as I can tell)
This is interesting work, and a good contribution, but it's important to compare similar models.[1] https://github.com/nlpxucan/WizardLM
[2] https://huggingface.co/vikp/llama_coder
[3] https://stability.ai/blog/stablecode-llm-generative-ai-codin...
[4] https://github.com/huggingface/blog/blob/main/starcoder.md
Edit: it could also be misleading to directly compare humaneval pass@1 against codellama without the same generation methodology. (possibly against GPT-4, also, but I don't know their methodology).
It teaches you everything you need to train an LLM, including the basics of deep learning and linear algebra. You learn the theory and the application. It includes written explanations, code, diagrams, and videos.
I believe that learning should be challenging enough to let the concepts sink in, so it's not a course you can just skim. It also isn't a "just type this, trust me" type of course - I think it's important to always know why you're doing something and how it works.
I've written 11 lessons, and I'm up to transformers - only a few more lessons to go. It's been fun to write, but balancing time spent training models with writing the course has been hard. Hopefully I will get to finish it soon.
If you also had the same initial thought as me, this is an excellent article - https://blog.eleuther.ai/transformer-math/ .
But, the popularity of LLMs indicates that people want to complete tasks efficiently. I think these two goals are usually in conflict.
I didn't fully reconcile them in my post because I'm not sure how to. But it's something I'm thinking about.
- I wouldn't use anything higher than a 7B model if you want decent speed.
- Quantize to 4-bit to save RAM and run inference faster.
Speed will be around 15 tokens per second on CPU (tolerable), and 5-10x faster with a GPU.To start, just using an LLM API (like Anthropic or OpenAI), and a light wrapper like microsoft guidance will be enough for the AI piece. If you want to get more complex, you can add in semantic search with an embedding model and a vector database. But don't do that off the bat.
For your use case, you won't need ML off the bat, either. When/if you do need ML models like classifiers, I'd use scikit-learn.
For the queries, I would skip pandas, and just use SQL. You can use an LLM to turn natural language into SQL queries, then just show the query results in an interface. The hardest part will actually be mapping the queries into the interface, and vice versa.
For my stack, I used FastAPI for the backend, and SvelteKit for the frontend. I highly recommend this stack for LLM applications - the async paradigm works well for streaming LLM outputs, and you get nice reactivity on the frontend.
It is hard to scale background tasks when you hit a high level of concurrency, as you've mentioned.
Your press releases makes it sound like the initial setup of background tasks is hard, and doesn't mention the harder stuff.
The press release is unnecessarily hyperbolic, which turned me off, though:
> Deploying new jobs to production also requires tedious configuration of cloud infrastructure which often requires a handoff to another team or individual. Often weeks of developer time is spent on basic workflows, before anything complex like idempotency is handled. Using Inngest, developers can write, test, and deploy complex workflows to production in hours, not weeks — all without touching infrastructure or queues.
Using something like Dramatiq [1] with Redis, writing a background job takes minutes, and can be deployed alongside an existing Python web app. There are probably JS equivalents.
I think Inngest could be a useful service (I might have used it if I'd seen it a few weeks ago), but the comparison felt off for me - it made me feel like this wasn't solving a real problem.
Some tips:
- Most vector search is basically kNN under the hood, with some kind of compression. If you have too many embeddings in your DB, this starts to pull up irrelevant text very quickly. The key is to segment the DB using other data before doing the embedding search. Postgres extensions are good for this.
- The quality of the data you put into your embedding DB matters a lot.
- How you chunk text matters. Chunking by paragraph is much better than naive chunking, for example.
- This is a good benchmark for embedding models [2]
[1] https://www.endless.academyBut the general dialogue around AI-related tools is surprising to me. The production parts of the langchain, embeddings, etc tools can usually be built in a few hours with better observability, performance, and maintainability.
1. Install sentence-transformers [1]
2. Initialize the MiniLM model - `model = SentenceTransformer('all-MiniLM-L6-v2')`
3. Embed your corpus [2]
4. Embed your queries, then search the corpus
This runs on CPU (~750 sentences per second), and GPU (18k sentences per second). You can use paragraphs instead of sentences if you need more text. The embeddings are accurate [3] and only 384 dimensions, so they're space-efficient [4].Here's how to handle persistence. I recommend starting with the simplest strategy, and only getting more complex if you need higher performance:
- Just save the embedding tensors to disk, and load them if you need them later.
- Use Faiss to store the embeddings (it will use an index to retrieve them faster) [5]
- Use pgvector, an extension for postgres that stores embeddings
- If you really need it, use something like qdrant/weaviate/pinecone, etc.
This setup is much simpler and cheaper than using a ton of cloud services to do embeddings. I don't know why people make semantic search so complex.I've used it for https://www.endless.academy, and https://www.dataquest.io and it's worked well in production.
[2] https://www.sbert.net/examples/applications/semantic-search/...
[3] https://huggingface.co/blog/mteb
[4] https://medium.com/@nils_reimers/openai-gpt-3-text-embedding...
Email me at hn at vikas.sh if you have a service. I'd need an SLA for sure, and multi-file support would be nice to have.
It's really oversimplified, as I mentioned. A more granular look is:
- Project the vectors with a linear regression. In decoder-only attention (what we usually use), we project the same vectors twice with different coefficients. We call the first projection queries, and the second keys. This transforms the vectors linearly.
- Find the dot product of each query vector against the key vectors (multiply them)
- (training only) Mask out future vectors, so a token can't look at tokens that come after it
- At this point, you will have a matrix indicating how important each query vector considers each other vector (how important each token considers the other tokens)
- Take the softmax, which both ensures all of the attention values for a vector sum to 1, and penalizes small attention values
- Use the softmax values to get a weighted sum of tokens according to the attention calc.
- This will turn one vector into the weighted sum of the other vectors it considers important.
The goal of this is to incorporate information from multiple tokens into a single representation.In LLMs, this means go from prompt to answer. I'll cover inference only, not training.
I can't quite ELI5, but process is roughly:
- Write a prompt
- Convert each token in the prompt (roughly a word) into numbers. So "the" might map to the number 45.
- Get a vector representation of each word - go from 45 to [.1, -1, -2, ...]. These vector representations are how a transformer understands words.
- Combine vectors into a matrix, so the transformer can "see" the whole prompt at once.
- Repeat the following several times (once for each layer):
- Multiply the vectors by the other vectors. This is attention - it's the magic of transformers, that enables combining information from multiple tokens together. This generates a new matrix.
- Feed the matrix into a linear regression. Basically multiply each number in each vector by another number, then add them all together. This will generate a new matrix, but with "projected" values.
- Apply a nonlinear transformation like relu. This helps model more complex functions (like text input -> output!)
Note that I really oversimplified the last few steps, and the ordering.At the end, you'll have a matrix. You then convert this back into numbers, then into text.