HNHacker News
TopNewBestAskShowJobs

vikp

1,176 karma · joined August 13, 2012

I used to teach people, now I teach machines.

Email me at hn@vikas.sh, or check out my work at https://www.vikas.sh.

submissionscomments
vikp··on Marker: Convert PDF to Markdown quickly with high accuracy
Yes, nougat is used as part of the pipeline to convert the equations (basically marker detects the equations then passes those regions to nougat). It's a great model for this.
vikp··on Marker: Convert PDF to Markdown quickly with high accuracy
(author) Please feel free to open an issue if you try again. Poetry can be painful, I might just switch to a requirements.txt file in the future. (you can skip poetry if you want by just pulling everything in pyproject.toml into a requirements.txt file also)
vikp··on Marker: Convert PDF to Markdown quickly with high accuracy
I chose markdown because I wanted to preserve equations (fenced by $/$$), tables, bold/italic information, and headers. I haven't looked into epub output, but this ruled out plain text.
vikp··on Marker: Convert PDF to Markdown quickly with high accuracy
Author here: for my use case (converting scientific PDFs in bulk), nougat was the best solution, so I compared to it as the default. I also compare to naive text extraction further down.

Nougat is a great model, and converts a lot of PDFs very well. I just wanted something faster, and more generalizable.

vikp··on Marker: Convert PDF to Markdown quickly with high accuracy
Author here - this is one of the reasons I made this. Also see https://github.com/VikParuchuri/libgen_to_txt , although I haven't integrated marker with it yet (it uses naive text extraction).
vikp··on Deep Learning Course
The deep learning book is a great choice, as many have mentioned.

I've been making a course that has a little less theory, and a little more application here - https://github.com/VikParuchuri/zero_to_gpt . Videos are all optional (cover the same content as the text).

vikp··on Ask HN: Resources to brush up from 'Intro to ML' to current LLMs/generative AI?
I'm building a course that teaches deep learning from the ground up - https://github.com/VikParuchuri/zero_to_gpt .

It balances theory and code, and builds from the foundation up, so you're never typing something without understanding it. Teaching method is text, diagrams, and code. Most lessons have optional videos, too.

It focuses on text models over image models (rnn, transformer, etc).

It's not 100% finished, but has enough to get you very far.

vikp··on Llama 2 on togetherAI is as bad of a privacy nightmare as OpenAI
This is not a random website - together is a prominent AI research company. Their chief scientist invented flashattention, which is used to train most LLMs. They release open source datasets/models like RedPajama. And they've published research around speeding up model training and newer model architectures.
vikp··on Refact Code LLM: 1.6B LLM for code that reaches 32% HumanEval
This post is misleading, in a way that is hard to do accidentally.

  - They compare the performance of this model to the worst 7B code llama model.  The base code llama 7B python model scores 38.4% on humaneval, versus the non-python model, which only scores 33%.
  - They compare their instruct tuned model to non-instruct-tuned models.  Instruction tuning can add 20% or more to humaneval performance.  For example, WizardLM 7B scores 55% on humaneval [1], and I've trained a 7B model that scores 62% [2].
  - For another example of instruction tuning, Stablecode instruct tuned benchmarks at 26%, not the 20% they cite for the base model [3]
  - Starcoder, when prompted properly, scores 40% on humaneval [4]
  - They do not report their base model performance (as far as I can tell)
This is interesting work, and a good contribution, but it's important to compare similar models.

[1] https://github.com/nlpxucan/WizardLM

[2] https://huggingface.co/vikp/llama_coder

[3] https://stability.ai/blog/stablecode-llm-generative-ai-codin...

[4] https://github.com/huggingface/blog/blob/main/starcoder.md

vikp··on Wikipedia search-by-vibes through millions of pages offline
There are plenty of great embedding models that are on the order of a few hundreds megs (even outperforming ada-002). See the leaderboard here - https://huggingface.co/spaces/mteb/leaderboard. Local/offline is only growing.
vikp··on Beating GPT-4 on HumanEval with a fine-tuned CodeLlama-34B
Got it, thanks - and thanks for the model! I'd be interested in the results if anyone benchmarks without sampling.

Edit: it could also be misleading to directly compare humaneval pass@1 against codellama without the same generation methodology. (possibly against GPT-4, also, but I don't know their methodology).

vikp··on Beating GPT-4 on HumanEval with a fine-tuned CodeLlama-34B
Did you use the same pass@1 generation method as in the code llama paper (greedy decoding)? I couldn't find this in the blog post.
vikp··on Ask HN: Tell us about your project that's not done yet but you want feedback on
I'm writing a deep learning course called Zero to GPT - https://github.com/VikParuchuri/zero_to_gpt .

It teaches you everything you need to train an LLM, including the basics of deep learning and linear algebra. You learn the theory and the application. It includes written explanations, code, diagrams, and videos.

I believe that learning should be challenging enough to let the concepts sink in, so it's not a course you can just skim. It also isn't a "just type this, trust me" type of course - I think it's important to always know why you're doing something and how it works.

I've written 11 lessons, and I'm up to transformers - only a few more lessons to go. It's been fun to write, but balancing time spent training models with writing the course has been hard. Hopefully I will get to finish it soon.

vikp··on Numbers every LLM Developer should know
I clicked because I thought they were defining LLM developer as "someone training LLMs", but instead they define it as "someone integrating LLMs into their application".

If you also had the same initial thought as me, this is an excellent article - https://blog.eleuther.ai/transformer-math/ .

vikp··on A vision for the AI web: the real web 3.0
That's right, although I used "tasks", which I think are higher level than individual search terms. I think there's also potential for personalization - learning the patterns that work specifically for you.
vikp··on A vision for the AI web: the real web 3.0
I'm torn on this. Sharing your own personal website and thoughts is a core part of the web, I agree.

But, the popularity of LLMs indicates that people want to complete tasks efficiently. I think these two goals are usually in conflict.

I didn't fully reconcile them in my post because I'm not sure how to. But it's something I'm thinking about.

vikp··on Llama2.c: Inference llama 2 in one file of pure C
I would use textsynth (https://bellard.org/ts_server/) or llama.cpp (https://github.com/ggerganov/llama.cpp) if you're running on CPU.

  - I wouldn't use anything higher than a 7B model if you want decent speed.
  - Quantize to 4-bit to save RAM and run inference faster.
Speed will be around 15 tokens per second on CPU (tolerable), and 5-10x faster with a GPU.
vikp··on Ask HN: Applying Open Source ML and LLM in side projects – where to start?
I was in a similar boat, and I built a project called Endless Academy - https://www.endless.academy/ . It helped me both brush up on some new tools, and scratch an itch to build something.

To start, just using an LLM API (like Anthropic or OpenAI), and a light wrapper like microsoft guidance will be enough for the AI piece. If you want to get more complex, you can add in semantic search with an embedding model and a vector database. But don't do that off the bat.

For your use case, you won't need ML off the bat, either. When/if you do need ML models like classifiers, I'd use scikit-learn.

For the queries, I would skip pandas, and just use SQL. You can use an LLM to turn natural language into SQL queries, then just show the query results in an interface. The hardest part will actually be mapping the queries into the interface, and vice versa.

For my stack, I used FastAPI for the backend, and SvelteKit for the frontend. I highly recommend this stack for LLM applications - the async paradigm works well for streaming LLM outputs, and you get nice reactivity on the frontend.

vikp··on Build a semantic search engine in Python
I've noticed that semantic search tutorials only use cloud vector databases and embedding APIs, so I figured I'd take a stab at writing a simpler guide. It's possible to get good accuracy and speed entirely locally.
vikp··on Inngest raises $3M seed to build the reliable workflow platform for every dev
I actually agree, which is why I wrote the comment. It is easy to create background tasks. Dramatiq can even handle some of the cases you mentioned - multi-step, fan-out, and retries.

It is hard to scale background tasks when you hit a high level of concurrency, as you've mentioned.

Your press releases makes it sound like the initial setup of background tasks is hard, and doesn't mention the harder stuff.

vikp··on Inngest raises $3M seed to build the reliable workflow platform for every dev
I think this is an interesting service, and could be a nice way to write a SvelteKit or Next.js app with background workers.

The press release is unnecessarily hyperbolic, which turned me off, though:

> Deploying new jobs to production also requires tedious configuration of cloud infrastructure which often requires a handoff to another team or individual. Often weeks of developer time is spent on basic workflows, before anything complex like idempotency is handled. Using Inngest, developers can write, test, and deploy complex workflows to production in hours, not weeks — all without touching infrastructure or queues.

Using something like Dramatiq [1] with Redis, writing a background job takes minutes, and can be deployed alongside an existing Python web app. There are probably JS equivalents.

I think Inngest could be a useful service (I might have used it if I'd seen it a few weeks ago), but the comparison felt off for me - it made me feel like this wasn't solving a real problem.

[1] https://dramatiq.io/

vikp··on 20x faster than pgvector: HNSW index in Postgres with pg_embedding
Yes, I use retrieval for Endless Academy [1] , and it works well.

Some tips:

  - Most vector search is basically kNN under the hood, with some kind of compression.  If you have too many embeddings in your DB, this starts to pull up irrelevant text very quickly.  The key is to segment the DB using other data before doing the embedding search.  Postgres extensions are good for this.
  - The quality of the data you put into your embedding DB matters a lot.
  - How you chunk text matters.  Chunking by paragraph is much better than naive chunking, for example.
  - This is a good benchmark for embedding models [2]
[1] https://www.endless.academy

[2] https://huggingface.co/blog/mteb

vikp··on Serverless Semantic Search, Free tier only
I was surprised, too, but then I realized they all work at Qdrant.

But the general dialogue around AI-related tools is surprising to me. The production parts of the langchain, embeddings, etc tools can usually be built in a few hours with better observability, performance, and maintainability.

vikp··on Serverless Semantic Search, Free tier only
I'd rather run ~10 lines of code locally than setup 3 cloud services and a lambda function, but to each their own...
vikp··on Serverless Semantic Search, Free tier only
Just run it on CPU, on your own machine. That's the cheapest way. You could also rent a free/cheap VPS, and even parallelize across multiple machines/cores if you need it.
vikp··on Serverless Semantic Search, Free tier only
This tutorial is very complex. Here's how to get free semantic search with much less complexity:

  1. Install sentence-transformers [1]
  2. Initialize the MiniLM model - `model = SentenceTransformer('all-MiniLM-L6-v2')`
  3. Embed your corpus [2]
  4. Embed your queries, then search the corpus
This runs on CPU (~750 sentences per second), and GPU (18k sentences per second). You can use paragraphs instead of sentences if you need more text. The embeddings are accurate [3] and only 384 dimensions, so they're space-efficient [4].

Here's how to handle persistence. I recommend starting with the simplest strategy, and only getting more complex if you need higher performance:

  - Just save the embedding tensors to disk, and load them if you need them later.
  - Use Faiss to store the embeddings (it will use an index to retrieve them faster) [5]
  - Use pgvector, an extension for postgres that stores embeddings
  - If you really need it, use something like qdrant/weaviate/pinecone, etc.
This setup is much simpler and cheaper than using a ton of cloud services to do embeddings. I don't know why people make semantic search so complex.

I've used it for https://www.endless.academy, and https://www.dataquest.io and it's worked well in production.

[1] https://www.sbert.net/

[2] https://www.sbert.net/examples/applications/semantic-search/...

[3] https://huggingface.co/blog/mteb

[4] https://medium.com/@nils_reimers/openai-gpt-3-text-embedding...

[5] https://github.com/facebookresearch/faiss

vikp··on Ask HN: Is the AI tool market a too crowded now?
I think it's more about potential TAM than the current product. The hypothesis is probably that there will eventually be a winner in each vertical, and that $$ can mint the winner.
vikp··on Show HN: I made an in-browser code editor with code replay and REPL
This would be interesting to me. There are a few options now, like Judge0, but the language versions are pretty out of date. Self-hosting is not a good time investment at the moment.

Email me at hn at vikas.sh if you have a service. I'd need an SLA for sure, and multi-file support would be nice to have.

vikp··on Ask HN: Can someone ELI5 transformers and the “Attention is all we need” paper?
You mean "Multiply the vectors by the other vectors. This is attention - it's the magic of transformers, that enables combining information from multiple tokens together. This generates a new matrix."?

It's really oversimplified, as I mentioned. A more granular look is:

  - Project the vectors with a linear regression.  In decoder-only attention (what we usually use), we project the same vectors twice with different coefficients.   We call the first projection queries, and the second keys.  This transforms the vectors linearly.
  - Find the dot product of each query vector against the key vectors (multiply them)
  - (training only) Mask out future vectors, so a token can't look at tokens that come after it
  - At this point, you will have a matrix indicating how important each query vector considers each other vector (how important each token considers the other tokens)
  - Take the softmax, which both ensures all of the attention values for a vector sum to 1, and penalizes small attention values
  - Use the softmax values to get a weighted sum of tokens according to the attention calc.
  - This will turn one vector into the weighted sum of the other vectors it considers important.
The goal of this is to incorporate information from multiple tokens into a single representation.
vikp··on Ask HN: Can someone ELI5 transformers and the “Attention is all we need” paper?
Transformers are about converting some input data (usually text) to numeric representations, then modifying those representations through several layers to generate a target representation.

In LLMs, this means go from prompt to answer. I'll cover inference only, not training.

I can't quite ELI5, but process is roughly:

  - Write a prompt
  - Convert each token in the prompt (roughly a word) into numbers.  So "the" might map to the number 45.
  - Get a vector representation of each word - go from 45 to [.1, -1, -2, ...]. These vector representations are how a transformer understands words.  
  - Combine vectors into a matrix, so the transformer can "see" the whole prompt at once.
  - Repeat the following several times (once for each layer):
  - Multiply the vectors by the other vectors.  This is attention - it's the magic of transformers, that enables combining information from multiple tokens together.  This generates a new matrix.
  - Feed the matrix into a linear regression.  Basically multiply each number in each vector by another number, then add them all together.  This will generate a new matrix, but with "projected" values.
  - Apply a nonlinear transformation like relu.  This helps model more complex functions (like text input -> output!)
Note that I really oversimplified the last few steps, and the ordering.

At the end, you'll have a matrix. You then convert this back into numbers, then into text.

← PreviousPage 2 of 6Next →