HNHacker News
TopNewBestAskShowJobs

aazo11

325 karma · joined March 20, 2019

submissionscomments
aazo11··on The Accuracy of On-Device LLMs
I tested on-device LMs (Gemma, DeepSeek) across prompt cleanup, PII redaction, math, and general knowledge on my M2 Max laptop using LM Studio + DSPy.

Some observations

- Gemma-3 is the best model for on-device inference - 1B models look fine at first but break under benchmarking - 4B can handle simple rewriting and PII redaction. It also did math reasoning surprisingly well. - General knowledge Q&A does not work with a local model. This might work with a RAG pipeline or additional tools

I plan on training and fine-tuning 1B models to see if I can build high accuracy task specific models under 1GB in the future.

aazo11··on AI's Version of Moore's Law
The trend is that the length of tasks AI can do is doubling every 7 months.

Accompanying YT video https://www.youtube.com/watch?v=evSFeqTZdqs

aazo11··on Lossless LLM compression for efficient GPU inference via dynamic-length float
This is a huge unlock for on-device inference. The download time of larger models makes local inference unusable for non-technical users.
aazo11··on Local LLM inference – impressive but too hard to work with
A better solution would train/finetune the smaller model from the responses of the larger model and only push to the inference to the edge if the smaller model is performant and the hardware specs can handle the workload?
aazo11··on Local LLM inference – impressive but too hard to work with
Thanks for calling that out. It was 32GB. I updated the post as well.
aazo11··on Local LLM inference – impressive but too hard to work with
Very interesting. I had not thought about gaming at all but that makes a lot of sense.

I also agree the goal should not be to replace ChatGPT. I think ChatGPT is way overkill for a lot of the workloads it is handling. A good solution should probably use the cloud LLM outputs to train a smaller model to deploy in the background.

aazo11··on Local LLM inference – impressive but too hard to work with
They look awesome. Will try it out.
aazo11··on Local LLM inference – impressive but too hard to work with
Exactly. Why does this not exist yet?
aazo11··on Local LLM inference – impressive but too hard to work with
By "too hard" I do not mean getting started with them to run inference on a prompt. Ollama especially makes that quite easy. But as an application developer, I feel these platforms are too hard to build around. The main issues being: getting the correct small enough task specific model and how long it takes to download these models for the end user.
aazo11··on Local LLM inference – impressive but too hard to work with
I spent a couple of weeks trying out local inference solutions for a project. Wrote up my thoughts with some performance benchmarks in a blog.

TLDR -- What these frameworks can do on off the shelf laptops is astounding. However, it is very difficult to find and deploy a task specific model and the models themselves (even with quantization) are so large the download would kill UX for most applications.

aazo11··on Show HN: GitHub-assistant – Natural language questions from your GitHub data
Great question! The purpose of github-assistant is to showcase the technologies that make it easy to build a tool/feature like this, not necessarily for it to be a stand-alone service. With dlt/Relta/LangGraph/assistant-ui we spin this up in about 10 days. For example:

- The GitHub graphql API limits to 100 items to be queried at a time and has pretty opaque secondary rate limits. Building this with cURL would take effort. dlt handles all this complexity to set up a robust pipeline by providing a connector to the GitHub API. - Creating semantic layers manually from a relational dataset and leveraging it in a text-to-sql pipeline to prevent hallucinations (similar to those we highlighted in our Medium post) would take lots of manual effort, which Relta streamlines. - Creating a chat front-end with charts was made easy by assistant-ui

Hope this makes sense.

aazo11··on Show HN: GitHub-assistant – Natural language questions from your GitHub data
Yes in the future. We share the source code in both commercial and non-commercial engagements already. Drop me a line at amir [at] relta.dev if interested.
aazo11··on Show HN: GitHub-assistant – Natural language questions from your GitHub data
There will new data from the graphql API added over time. Would love your feedback on which data you like to see added https://docs.github.com/en/graphql
aazo11··on Show HN: GitHub-assistant – Natural language questions from your GitHub data
No this currently only answers questions from the GitHub graphql API.
aazo11··on Show HN: GitHub-assistant – Natural language questions from your GitHub data
We pull data from the GitHub API which includes data that that is not available from GitHub.com pages. Currently only PR, Issues, Commit and Star data is being loaded. You can also read more here https://medium.com/relta/github-assistant-49ae388ad758
aazo11··on Show HN: GitHub-assistant – Natural language questions from your GitHub data
Was able to reproduce and pushed an update. Thanks for calling this out.
aazo11··on Show HN: GitHub-assistant – Natural language questions from your GitHub data
Hi -- strange that didn't work. Overall, the semantic layer is designed to provide very tight guardrails and not hallucinate. You can see the agent suggest changes to the semantic layer if you give the produced answer a thumbs down.

The idea is for the system to provide answers that have close to 100% accuracy, but make it a single click for developers to to improve the semantic layer.

aazo11··on Show HN: We open sourced our entire text-to-SQL product
Right now the supported Vector stores are Chroma (which you can self-host), Pinecone and Astra. Adding a new vector store is quite easy: you just need to extend the VectorStore class (https://github.com/Dataherald/dataherald/tree/main/services/...) and set it as the Vector store module to be used in the environment variable https://github.com/Dataherald/dataherald/blob/main/services/...
aazo11··on Show HN: We open sourced our entire text-to-SQL product
There are organizations using Dataherald in production right now.

The latency is ~20-30s and it takes some set up, so as long as those are not blockers it can be used in prod.

aazo11··on Show HN: We open sourced our entire text-to-SQL product
Yes when you connect Dataherald to a DB it scans it and you can do exploratory queries.
aazo11··on Show HN: We open sourced our entire text-to-SQL product
ORMs generally map around entities and dimensions. Users generally ask about metrics and measures, which can be expressed in aggregations and group bys.

How ould the NLP+ORM system do this?

aazo11··on Show HN: We open sourced our entire text-to-SQL product
While the engine response is not accurate all the time, the engine returns a confidence score. We have never encountered cases where a deployment with necessary training data indicates a .9 confidence score on an incorrectly generated SQL.
aazo11··on Show HN: We open sourced our entire text-to-SQL product
Tables, columns and views are scanned at configuration time (or based on an API trigger) and stored in the data store and a vector store, not on every run.

They are then retrieved and injected based on relevance to the query.

aazo11··on Show HN: We open sourced our entire text-to-SQL product
Added the License
aazo11··on Show HN: We open sourced our entire text-to-SQL product
As I wrote on the original thread, we recommend using the RDBMS row-level security features.

This blog discusses how to do that on Postgres

https://www.2ndquadrant.com/en/blog/application-users-vs-row...

aazo11··on Show HN: We open sourced our entire text-to-SQL product
We recommend users leverage row-level security features built into modern RDBMS so the query results only return data for a given user.

You can read more on how to do that on Postgres here https://www.2ndquadrant.com/en/blog/application-users-vs-row...

aazo11··on Show HN: We open sourced our entire text-to-SQL product
Currently does not but looking to add support. Would love to connect and learn more about your use case.
aazo11··on Show HN: We open sourced our entire text-to-SQL product
Sure will reach you out. Currently Dataherald blocks DML or DDL commands from being generated/executed.
aazo11··on Show HN: We open sourced our entire text-to-SQL product
This is not a text to semantic layer but it does far more than just inject schema into the prompt:

- the engine keeps an updated catalog of the data (low cardinality columns, their values etc) - taps into query history and finetunes the model to the schema - allows uploading context from unstructured sources like docs and data dictionaries - has an agent which collects all relevant info, generate the SQL, tries to retrieve a few rows to recover from errors and provides an confidence score to the generated SQL

aazo11··on Show HN: We open sourced our entire text-to-SQL product
The agent is LLM agnostic and you can use it with OpenAI or self-hosted LLMs. For self hosted LLM we have benchmarked performance with Mixtral for tool selection and CodeLlama for code generation.
Page 1 of 3Next →