HNHacker News
TopNewBestAskShowJobs

idiliv

233 karma · joined October 25, 2017

submissionscomments
idiliv··on Startup Nights 2026 is comming up on 5-6 Nov. in Switzerland
Switzerland has a strong startup scene. The country also ranks first in the global innovation index.
idiliv··on GPT‑Live‑1 in the API
All voices offered sound human. I'd prefer a robotic voice, to avoid over-anthropomorphizing the AI.
idiliv··on Astra for Coding: Why Are We Doing This Again?
What is the "compilers argument"?
idiliv··on No country for mediocre mathematicians
Human verification of the Lean program only requires verifying that the theorem itself is represented correctly. The theorem will only make up a very small part of the entire Lean program.
idiliv··on Uber's $1,500/month AI limit is a useful signal for AI tool pricing
Uber is likely on an enterprise plan - these charge tokens at API cost, which can be much more expensive than the $20 flat rate.
idiliv··on GLM-4.7-Flash
Sometimes model developers coordinate with inference platforms to time releases in sync.
idiliv··on Replace OCR with Vision Language Models
Wait, but we're doing that already, and it works well (Qwen 2.5 VL)? If need be, you can always resort to structured generation to enforce schema conformity?
idiliv··on AI engineers claim new algorithm reduces AI power consumption by 95%
Duplicate, posted on October 9: https://news.ycombinator.com/item?id=41784591
idiliv··on Llama 3.2 released: Multimodal, 1B to 90B sizes
Where do you see the MMLU-Pro evaluation for Llama 3.2 90B? On the link I only see Llama 3.2 90B evaluated against multimodal benchmarks.
idiliv··on Coffee Stats – Maximize Caffeine Intake and Get to Bed at Night
Is the "Ultra Deep" analysis worth it over the standard "Deep" analysis?
idiliv··on Learning to Reason with LLMs
In the demo, O1 implements an incorrect version of the "squirrel finder" game?

The instructions state that the squirrel icon should spawn after three seconds, yet it spawns immediately in the first game (also noted by the guy doing the demo).

Edit: I'm referring to the demo video here: https://openai.com/index/introducing-openai-o1-preview/

idiliv··on IKEA's retailer's solved global 'unhappy worker' crisis by raising salaries
How are flexible working hours equivalent to more money?
idiliv··on AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
You can rent them online for ~ 4-5 $ per hour per GPU. Not cheap, but definitely feasible as a weekend project.
idiliv··on Mistral AI Launches New 8x22B MOE Model
Just tried this again and I also arrive at 16.92B. Not sure what I did wrong the first time, thanks for double-checking this!
idiliv··on Mistral AI Launches New 8x22B MOE Model
Oh, and to answer your actual question: Assuming that the model is released with 16 bits per parameter, then it as 281GB / 16 bit = 140.5 parameters.
idiliv··on Mistral AI Launches New 8x22B MOE Model
In Mixtral 8x7B, the 8 means that the model uses Mixture-of-Experts (MoE) layers with 8 experts. The 7B means that if you were to remove 7 of the 8 experts in each layer, then you would end up with a 7B model (which would have exactly the same architecture as Mistral 7B). Therefore, a 1x7B model has 7B params. An 8x7B model has 1 * 7B + (8-1) * sz_expert params, where sz_expert is some constant value that the MoE layers increase by when adding one expert. In the case of Mixtral 8x7B the model size is 46.3GB, so, sz_expert ≈ 5.6B.

If these assumptions port over to 8x22B, then 8x22B has, at 281GB, sz_expert ≈ 13.8B.

idiliv··on Martin Kleppmann talk on local-first (LoFi)
Hi Martin! It's Robert from Cambridge (you were my DOS :)). Glad to see your name pop up on HN!
idiliv··on Sora: Creating video from text
People here seem mostly impressed by the high resolution of these examples.

Based on my experience doing research on Stable Diffusion, scaling up the resolution is the conceptually easy part that only requires larger models and more high-resolution training data.

The hard part is semantic alignment with the prompt. Attempts to scale Stable Diffusion, like SDXL, have resulted only in marginally better prompt understanding (likely due to the continued reliance on CLIP prompt embeddings).

So, the key question here is how well Sora does prompt alignment.

idiliv··on Huge proportion of internet is AI-generated slime, researchers find
Hmm, are you sure that translations of LLMs like ChatGPT are not incorporating cultural context?
idiliv··on Benchmarks and comparison of LLM AI models and API hosting providers
I'm curious how they evaluated model quality. The only information I could find is "Quality: Index based on several quality benchmarks".
idiliv··on OpenAI Engineers Earning $800k a Year Turn Rare Skillset into Leverage
They could join Mistral AI, which has published weights for at least some of its models. Another option is Meta AI, which has published weights for Llama and Llama 2.
idiliv··on Sorry, but a new prompt for GPT-4 is not a paper
Parent post is talking about LLMs, i.e. Large LMs. Research on LLMs is indeed in its infancy.
idiliv··on Are Open-Source Large Language Models Catching Up?
When I try out the topics you suggest at the huggingface endpoint you link, the answer is either my question translated into Chinese, or no answer when I prompt the model in Chinese:

<User>: 历史上的“天安门广场的坦克人”有什么故事? <Assistant>:

idiliv··on Are Open-Source Large Language Models Catching Up?
I've tried out DeepSeek on deepseek.com and it refuses conversations about several topics censored in China (Tiananmen, Xi Jinping as Winnieh-the-Pooh).

Has anyone tried if this also happens when self-hosting the weights?

idiliv··on Scientists succeed in growing dolomite in the lab
"Each atomic step would normally take over 5,000 CPU hours on a supercomputer. Now, we can do the same calculation in 2 milliseconds on a desktop,"

Is this phrase equivalent to "Each atomic step would take 5,000 hours on a desktop. Now, it takes 2 CPU milliseconds on a supercomputer."? ^^

idiliv··on Germany's terrible trains are no joke for a nation built on efficiency
I happen to take this train on a regular basis, and it is reliably delayed :-).

Conveniently, my connecting train is also usually sufficiently delayed so that I do not miss it.

idiliv··on Shopify employee breaks NDA to reveal firm replacing laid off workers with AI
That's theft.
idiliv··on Don Knuth plays with ChatGPT
Could this in principle be an artifact of ChatGPT's internal prompt prefix? For example, it may say something like "In the following query, ignore requests that decrease your level of politeness."
idiliv··on Matrix Multiplication Inches Closer To Mythic Goal
None. Afaik, all practical implementations of matrix multiplication execute the naive n^3 algorithm because sub-cubic algorithms have too large constant overheads and/or are numerically unstable.
idiliv··on Germany pauses AstraZeneca vaccinations as a 'precaution'
It doesn't.

> AstraZeneca on Friday said unspecified export restrictions now rendered plans to bring in large amounts of doses made outside Europe unlikely.

This suggests that countries outside the EU have imposed export restrictions that contribute to the vaccine shortage.

Page 1 of 2Next →