HNHacker News
TopNewBestAskShowJobs

convexstrictly

1,147 karma · joined March 26, 2023

submissionscomments
convexstrictly··on Gigatoken: Fastest Tokenizer
Tokenization is done on the CPU. Models never see the raw characters. That's why you get trick questions like the number of r's in strawberry.

There are many research papers on models using characters directly. One challenge is that effective context length is smaller.

convexstrictly··on Gigatoken: Fastest Tokenizer
GitHub: https://github.com/marcelroed/gigatoken
convexstrictly··on Gemini Flash 2.0 Thinking Experimental
Video Demo:

https://x.com/OfficialLoganK/status/1869789822384255300

convexstrictly··on Gemini Flash 2.0 Thinking Experimental
"Just when you thought it was over... we’re introducing Gemini 2.0 Flash Thinking, a new experimental model that unlocks stronger reasoning capabilities and shows its thoughts.

The model plans (with thoughts visible), can solve complex problems with Flash speeds, and more ..."

- Logan Kilpatrick

https://x.com/OfficialLoganK/status/1869789822384255300

convexstrictly··on M4 MacBook Pro
GB per second
convexstrictly··on ThunderKittens: Simple, fast, and adorable AI kernels
Simran Arora: "Join us for a livestream this Thursday, Halloween/Diwali, and join our channel on the GPU Mode Discord server to hang out with us/get involved:"

https://discord.com/login?redirect_to=%2Fchannels%2F11894982...

convexstrictly··on ThunderKittens: Simple, fast, and adorable AI kernels
CUDA + ThunderKittens 4.5 hour tutorial

https://www.youtube.com/watch?v=xcpEl0cGCC4

convexstrictly··on Building GPT2o – Part 1: Audio
Results https://x.com/sbeastwindy/status/1801525876267372874
convexstrictly··on LSP-AI: open-source language server serving as back end for AI code assistance
Aider uses Treesitter to improve code generation. https://aider.chat/2023/10/22/repomap.html

Aider: https://github.com/paul-gauthier/aider

It is state of the art on SWE-Bench and SWE-Bench Lite. https://aider.chat/2024/06/02/main-swe-bench.html

convexstrictly··on Llm.c – LLM training in simple, pure C/CUDA
Candle is a minimalist ML framework for Rust with a focus on performance (including GPU support) and ease of use

https://github.com/huggingface/candle

convexstrictly··on LISA: Layerwise Importance Sampling for Memory-Efficient LLM Fine-Tuning
An author claims better performance than LoRA in 50% of the time.

https://twitter.com/Rui45898440/status/1772996453557997924

convexstrictly··on NTIA AI Open Model Weights RFC
The federal government requests comments on regulation of AI models with openly available weights. The deadline is March 27, 2024.

Earlier thread. https://news.ycombinator.com/item?id=39494760

convexstrictly··on Show HN: GPU Price list inspired by diskprices.com
It would be good to know on which listings Amazon is the seller. Filtering by that criterion may also be useful.

For professional cards, I've noticed dihuni.com has good prices. I have never purchased from them and have no idea what dealing with them is like.

convexstrictly··on Star Trek prompt optimal for grade school math on Llama-70B
The results are from the paper

The Unreasonable Effectiveness of Eccentric Automatic Prompts

https://arxiv.org/abs/2402.10949

convexstrictly··on Mistral Large
Pricing

input: $8/1M tokens

output: $24/1M tokens

https://docs.mistral.ai/platform/pricing/

convexstrictly··on NTIA Solicits Comments on Open-Weight AI Models
Jeremy makes compelling arguments. Here are some more mundane corollaries:

It is only a matter of time before you and your company are affected by the pending regulations. In the future, almost all software products will be using AI models, in the same way that most software products use the Internet today, whereas they did not in the 1990s.

Imagine you had to license Oracle software because MySQL or PostgresSQL could not offer certain capabilities, or are less capable because of regulation.

Now also imagine that your products have to agree with the political world view of either Sundar Pichai or Elon Musk. And if you need capabilities only present in one of the commercial alternatives, you don't even that that choice.

convexstrictly··on NTIA Solicits Comments on Open-Weight AI Models
Some information here.

https://www.ntia.gov/federal-register-notice/2024/dual-use-f...

convexstrictly··on NTIA Solicits Comments on Open-Weight AI Models
The comments will inform the drafting of regulations on open weight models under the Biden executive order on AI using his powers under the Defense Production Act.

Fact Sheet: https://www.whitehouse.gov/briefing-room/statements-releases...

Full Details: https://www.whitehouse.gov/briefing-room/presidential-action...

convexstrictly··on BitDelta: Your Fine-Tune May Only Be Worth One Bit
"By enabling the use of a single high-precision base model accompanied by multiple 1-bit deltas, BitDelta dramatically reduces GPU memory requirements by more than 10x, which can also be translated to enhanced generation latency in multi-tenant settings."
convexstrictly··on BitDelta: Your Fine-Tune May Only Be Worth One Bit
https://github.com/FasterDecoding/BitDelta
convexstrictly··on Time is encoded in the weights of finetuned language models
Twitter summary: https://twitter.com/ssgrn/status/1738256456250470853

Github: https://github.com/KaiNylund/lm-weights-encode-time

convexstrictly··on Groqchat
Everything you say makes sense. Training is definitely more compute intensive than inference.

Training is both memory throughput and compute constrained. Much research in speeding up training goes into optimizing HBM to SRAM communication. The equivalent for your chips would be communication from the SRAM of one chip to the SRAM of another, where it sounds like your architecture has a major memory throughput advantage over GPUs. So I assume you don't have a proportional compute advantage?

By the way, it's great to see a non von Neumann architecture showing a major performance advantage in a real world application. And your chips are conceptually equivalent to chiplets; you should have a major cost advantage on bleeding edge process nodes if you scale up manufacturing. Overall very impressive!

convexstrictly··on Groqchat
Could you explain the blockers to getting back-propagation working well on your chips?
convexstrictly··on Zoology 1: Measuring and Improving Recall in Efficient Language Models
Research suggesting that much of the power of the transformer architecture comes from associative recall over long sequences that does not require scaling model dimensions. They design state space models that narrow the gap.

Overview https://hazyresearch.stanford.edu/blog/2023-12-11-zoology0-i...

Zoology 2 https://hazyresearch.stanford.edu/blog/2023-12-11-zoology2-b...

Monarchs and Butterflies: Towards Sub-Quadratic Scaling in Model Dimension. Scaling in model dimension as opposed to sequence dimension scaling in the previous posts. https://hazyresearch.stanford.edu/blog/2023-12-11-truly-subq...

convexstrictly··on TinyGSM: Achieving >80% on GSM8k with small language models
"... we find that a duo of a 1.3B generation model and a 1.3B verifier model can achieve 81.5% accuracy, outperforming existing models that are orders of magnitude larger."
convexstrictly··on GDlog: A GPU-accelerated deductive engine
The paper claims it builds upon the concepts in HashGraph, an efficient CUDA hashtable implementation.

HashGraph (2019) https://arxiv.org/abs/1907.02900

Anyone know what the most performant CUDA hash table implementations are these days?

convexstrictly··on GDlog: A GPU-accelerated deductive engine
Github repo

https://github.com/harp-lab/gdlog

convexstrictly··on Androids built to meet the labor demands
Twitter thread with video introduction

https://twitter.com/1x_tech/status/1730610445541638378

convexstrictly··on Sam Altman likely to start company with researchers from OpenAI: Bloomberg
Emily Chang from Bloomberg reports: Satya Nadella was "blindsided" and is furious. Details in paywalled article:

https://www.bloomberg.com/news/articles/2023-11-18/openai-al...

Chang does not make it clear whether the source is close to Sam Altman.

convexstrictly··on Ask HN: Is paid ChatGPT Plus worth it?
If you were logged into Bing, those prompts may be in your history. They can be viewed using the Edge browser.

A few weeks ago, I had spotty service with Bing Chat where it would keep resetting the conversation which I assumed was due to load. In general all these LLM services are in constant flux because they are tuning both the models and UI. They feel like alpha quality products in terms of stability.

Page 1 of 3Next →