HNHacker News
TopNewBestAskShowJobs

jumpCastle

42 karma · joined December 30, 2016

I design channel codes for flash memories.
submissionscomments
jumpCastle··on OpenAI o3 and o4-mini
It was a good benchmark until it entered the training set.
jumpCastle··on OpenAI o3 and o4-mini
But open source like aider
jumpCastle··on The Llama 4 herd
Deepseek v3 was FP8
jumpCastle··on The Llama 4 herd
It's quite similar to muP

https://github.com/microsoft/mup

jumpCastle··on Humans have caused 1.5 °C of long-term global warming according to new estimates
Stock prices cannot go up without if the planet is destroyed
jumpCastle··on Nvidia reportedly delays its next AI chip due to a design flaw
Attention was invented because Bengio lab had to be disciplined about a black box (google had more compute)
jumpCastle··on Exo: Run your own AI cluster at home with everyday devices
Aren't services like runpod solve half of these concerns?
jumpCastle··on Firefox 128 enables "privacy-preserving" ad measurements by default
Are the updates really slow?
jumpCastle··on Ilya Sutskever to leave OpenAI
Spend more time with my side projects.
jumpCastle··on 2023 ACM Turing Prize awarded to Avi Wigderson
You can try his book. https://www.math.ias.edu/avi/book
jumpCastle··on GPT-4 Turbo with Vision is a step backwards for coding
You can use older models with the api.
jumpCastle··on How Chain-of-Thought Reasoning Helps Neural Networks Compute
Also the parameters are optimized also with loss of future tokens in the sequence.
jumpCastle··on Learning Theory from First Principles [pdf]
Work pretty well for all classes of problems?
jumpCastle··on Hi everyone yes, I left OpenAI yesterday
Rsu exist before ipo?
jumpCastle··on Memory and new controls for ChatGPT
Use api, embed history and retrieve.
jumpCastle··on OpenAI suspends ByteDance's account after it used GPT to train its own AI model
Use model output as training data. For better performance you can get some top log probs and minimize kl divergence.
jumpCastle··on Mistral: Our first AI endpoints are available in early access
Without open weights why would anyone care about them? By the time they could compete with gpt4 there's probably be gpt5 already.
jumpCastle··on GPT-4 generates simple app from Whiteboard photo
The future is the AI also writes the high level descriptions of stuff.
jumpCastle··on Exllamav2: Inference library for running LLMs locally on consumer-class GPUs
But you can fine tune gpt 3.5 turbo, so your comparison is not clear.
jumpCastle··on Discord.io breached, 760k user accounts for sale on darknet
I still use infinity for reddit though.
jumpCastle··on Gzip and KNN Outperforms Transformers on Text Classification
Gzip every query with all training data can get more expensive.
jumpCastle··on Scaling Transformers to 1B Tokens
Silly is a charitable interpretation.
jumpCastle··on Scaling Transformers to 1B Tokens
Title with a 10 digits number, meaningless first page figure and no experiments related to the main claim. Did a rogue author posted it without permission again?
jumpCastle··on Threads has passed 2M sign ups in the first 2 hours
They no longer require sign-in. Got scared of the reaction.
jumpCastle··on GPT-Migrate converts repos from one lang/framework to another
Is Azure as simple to use as OpenAI? Their documentation are much less simple.
jumpCastle··on The Secret Sauce behind 100K context window in LLMs: all tricks in one place
For an answer I would expect it to get the same reward for both question orderings. So naively I would expect it to not be affected by the ordering.
jumpCastle··on The Secret Sauce behind 100K context window in LLMs: all tricks in one place
Read codebase and implement feature. Read paper and prove conjecture.
jumpCastle··on The Secret Sauce behind 100K context window in LLMs: all tricks in one place
https://github.com/google-research/long-range-arena
jumpCastle··on The Secret Sauce behind 100K context window in LLMs: all tricks in one place
nat.dev
jumpCastle··on The Secret Sauce behind 100K context window in LLMs: all tricks in one place
For web access there's nat.dev
Page 1 of 3Next →