HNHacker News
TopNewBestAskShowJobs

hustwindmaple1

35 karma · joined August 30, 2022

submissionscomments
hustwindmaple1··on TPU Deep Dive
Every major Cloud vendor is trying to develop their custom AI ASIC. Putting Google aside, Amazon has trainium/inferentia, which Anthropic uses quite extensively. Microsoft is doing sth. similar, although they are quite behind. OpenAI is doing it. Meta is doing it. That's why the stock price of Broadcom/Marvell soared.
hustwindmaple1··on KumoRFM: A Foundation Model for In-Context Learning on Relational Data
I remember Kumo was focusing on GNN when it was founded (Jure's strength back then). Looks like they are pivoting or have pivoted.
hustwindmaple1··on Linear Programming for Fun and Profit
You are basically doing a heurstic. Your solutions are not guaranteed to be optimal. Integer programming is the way to do.
hustwindmaple1··on Google Gemini has the worst LLM API
If you are not a paying GCP user, there is really no point to even look at Vertex AI.

Just stick with AI Studio and the free developer AI along with it; you will be much much happier.

hustwindmaple1··on Tracing the thoughts of a large language model
Really appreciate your team's enormous efforts in this direction, not only the cutting edge research (which I don't see OAI/DeepMind publishing any paper on) but aslo making the content more digestible for non-research audience. Please keep up the great work!
hustwindmaple1··on SoftBank Group to Acquire Ampere Computing for 6.5B
His Alibaba/Arm bets are legendary. Both made $100B+
hustwindmaple1··on China tells its AI leaders to avoid U.S. travel over security concerns
Housing is definitely not easily affordable. But the other points are valid.
hustwindmaple1··on Implementing LLaMA3 in 100 Lines of Pure Jax
Cool blog
hustwindmaple1··on How to scale your model: A systems view of LLMs on TPUs
there is limited TPU support in pytorch via torch_xla
hustwindmaple1··on Andrej Karpathy: Deep Dive into LLMs Like ChatGPT [video]
When he drops a vid, you don't ask questions. You watch first and then ask questions :)
hustwindmaple1··on DeepSeek FAQ
High-Flyer is pretty damn rich actually. Someone did some calculation and it turns out they are spending ~$200m (somewhere in that range) on 150 employee compensation alone.
hustwindmaple1··on Foundations of Large Language Models
prob. not prof; just phd students needs pubs to graduate
hustwindmaple1··on Trae.ai IDE
Are you comfortable to send your personal/company code data to a Chinese company (ByteDance)?
hustwindmaple1··on Trying out QvQ – Qwen's new visual reasoning model
There are a couple, i.e., OLMo 2
hustwindmaple1··on Trillium TPU Is GA
Exactly. Kind of like the old days when they put together a massive amount of commodity CPUs to build search.
hustwindmaple1··on Genie 2: A large-scale foundation world model
GDM is a research lab. They are not set up for production. There are other teams in Alphabet doing productionization stuff.
hustwindmaple1··on QwQ: Alibaba's O1-like reasoning LLM
Large Chinese companies usually have overseas subsidiaries, which can buy H100 GPUs from NVidia
hustwindmaple1··on Francois Chollet is leaving Google
well, is Waymo doing better than the PyTorch-powered Tesla?
hustwindmaple1··on Google warns uBlock Origin and other extensions may be disabled soon
Just use Brave. You'll thank me later ...
hustwindmaple1··on Faster Integer Programming
Reformulation is a good strategy for many hard problems.

It might also be possible to add problem-specific cuts/heurstics to the solver so that it can solve it fast.

hustwindmaple1··on Faster Integer Programming
I was on CPLEX team for a few years. Core developers are all PhDs from Stanford/MIT and etc. So it's very hardcore stuff, no less than AI research.

Completely agree that size is not a good proxy for estimating MIP difficulties. Internally we colllected a bunch of very hard problems to sovle from different domains. Some are actually pretty small, say a few thousand variables/constraints. IMHO what made hard problems difficult to solve is actually the 'intneral structure' of the problems. And modern industry solvers all have a lot of built-in heurstics to take advantages of the structures, i.e., what kind of cuts, presolve/diving/branching strategy to apply, how to get the bounds ASAP.

Interestinglly at some point some folks even tried using machine learning to predict strategies. Didn't work quite well back then (10+ years ago). There was some work of using seq2seq for MIP (pointer network, I think) a few years ago; worked OK. So I'm really looking forward to some breakthroughs by LLM.

It's a shame that after IBM aquired ILOG (which owns CPLEX), most of the ppl left for Gurobi.

hustwindmaple1··on Google announces Firebase Genkit with Ollama support
The problem is that there is no bette alternative.
hustwindmaple1··on Scaling Clubhouse From 10K to 10M Users In 6 Months With Postgres
Is Clubhouse still alive?
hustwindmaple1··on IBM Granite: A Family of Open Foundation Models for Code Intelligence
Right, but they can just use Llama/Mistral for free, instead of their inferior models, which I'm sure take quite a bit of resources to train in the first place.
hustwindmaple1··on IBM Granite: A Family of Open Foundation Models for Code Intelligence
I wonder why companies like IBM are jumping on the LLM bandwagon and training/releasing models that have no chance of competing with Llama/Mistral? To me it just looks like a complete waste of $$ because nobody will use them in any serious scenarios
hustwindmaple1··on Gemma: New Open Models
It's MQA, documented in the tech report
hustwindmaple1··on Mojo vs. Rust: is Mojo faster than Rust?
maybe they are tracking # of accounts for their next raise from VCs
hustwindmaple1··on Grammarly Announces Layoffs
It's not their emails. It's all the ads you see on YouTube, display ads on random websites through Google/FB ads network, ads within xyz apps, and etc.. They spent a lot of $$$ on those ad networks.

Those are usually targeted through personal email addresses and device IDs. Impossible to opt out on users' end.

hustwindmaple1··on Grammarly Announces Layoffs
I emailed their customer support and asked them to stop targeting my email account in all their ads. They said they did, but I still got a ton of their ads afterwards. Anybody with ads experience would instantly know their ads targeting wasn't working great and they were wasting a sh*t ton of money.
hustwindmaple1··on Grammarly Announces Layoffs
One of the most annoying companies that ad spams like crazy in 2020/2021. i literally had to beg them to stop sending ads my way.

Now they are running out of money and have to cut staff. How ironic!

Page 1 of 2Next →