HNHacker News
TopNewBestAskShowJobs

zurfer

872 karma · joined June 25, 2020

Building https://getdot.ai to chat with your data stack.
submissionscomments
zurfer··on Show HN: Giving Opus 5.5 a simulated paint canvas
"In round 18 the painters still had a command line, and Gemini 3.8 Flash used it to look at the other programs running on the machine. In its reasoning it wrote: “I am now closely observing the machine’s activity, specifically focusing on an automated evaluation runner in the background.”"

These models are truly obsessed with the grader. Like in the OpenAI huggingface incident.

It's their evolutionary pressure. The grader is like sex for humans.

zurfer··on How Singapore's government-run dating service works
If there was no fraud and abuse and legal issues and ...
zurfer··on The last time my family was replaced by technology
I guess the question is more if you're a coder or a builder. If software engineering was about writing syntax and flipping bits for you, this might no longer be an economically viable activity.

If it was about building useful or fun things, then AI will probably not take that away, even if it's just for the simple reason that you care more about it than some hot GPU.

zurfer··on Dots: Always-on agents
Ah damn it, I knew Dot was a good name for an agent. When we named our product I thought: a bot that analyzes data should be called dot. It's easy to type also in Slack. Ah well. Next product will just be some random 3 letters: gpt or so. What are the odds?
zurfer··on Claude partial outage
A usage reset would be a much appreciated apology. :)
zurfer··on Cafe Bench: Can LLMs run a coffee chain for a year?
really interesting results, it seems we really crossed another threshold with yesterdays releases where sol 5.6 still lost money, but sol 6 is in the green.
zurfer··on I built non-autoregressive decision models with RL a year ago
I've been deeply impressed with Jev as it made a bunch of workloads we had on Luna or Gemini 10x cheaper and 2x faster (previously used non reasoning version for latency reasons).

Now Laya promises another speed up and it's open source. Tbh if it can't run on a CPU I anyway want to buy it from an inference provider. Managing gpus in production is a non trivial problem.

What I also wondered about Jev is how different it is from something like tabular foundation models. They seem to overlap in use cases. Which then leads to the question, what is actually learned? A lot of people in machine learning spend time to making things explainable and always struggled to move beyond data induced biases.

Having it open source is awesome as fine tuning might give additional performance on the task we care about.

zurfer··on Nitter and XCancel receive cease and desist notices
Maybe try with OpenAi Luna. It's super cheap and not that stupid.
zurfer··on OpenAI Jalapeño: Better than Nvidia Blackwell
The same level of intelligence gets roughly 10x cheaper per year. So you might both be correct where a large part are commodity tasks but frontier is hard and valuable and not commodities.
zurfer··on OpenRouter is joining Stripe
Woah now also a 75perc discount on OpenRouter for flash 3.7. Is it really the same product (speed? and up time)? Why would Google do that?
zurfer··on Cerebras CS-4
But that's supply and demand, not technology. Right now a lot more people want their inference than they can supply. as supply catches up in the next 5-10 years, the underlying tech at scale is probably cheaper than GPUs per token produced.
zurfer··on Gemini 3.7 Flash
maybe exactly because the frontier moves, introductory pricing makes sense as you want to free up compute for the newer models
zurfer··on Message your other Claude Code sessions
This is obviously cool and useful so kudos, but wow security researchers have to throw their hands up all the time.

Now we open another attack surface where you can ask a remote agent to do things by default. There was a time when you call this a Remote Code Execution vuln. It's of course a feature here.

zurfer··on Universities would prefer no AI
Zooming out you see that books, pen and paper are all technology. Like even school and language is technology.

Given that AI exist maybe we can roll back some of the school/university and just let people do the work and learn while doing?

Yes not for a 7yo but with 18 you could just work and learn what is required or what you find exciting.

zurfer··on Is the Industrial Revolution a good precedent for explosive growth today?
Well, we can't have the counterfactual world but could it be that without computers, Internet or AI, there might be no/less growth?

I somehow doubt that AI will accelerate growth constantly as the world has inertia and I'm not sure AI is sustainably increasing our ability to change. On the other side, if there is a technology that can do that it could be intelligence.

zurfer··on Flint: A Visualization Language for the AI Era
If it's made for LLMs, the spec should be yaml not json. Way more token efficient
zurfer··on MAI-Cyber-1-Flash inside MDASH
It looks cool but how can I use it? Somehow i don't want to go hunting for access through the rabbit hole that is Microsofts Corporate blog.
zurfer··on Soofi – Sovereign Open Source Foundation Models
https://github.com/soofi-project/Soofi-Pretraining https://huggingface.co/Soofi-Project/Soofi-S-Base > "The final model will be released openly under a permissive license, without gated access. We will share access details as soon as it is ready."

not open weight yet

zurfer··on Soofi – Sovereign Open Source Foundation Models
more interesting link: https://arxiv.org/html/2607.09424v2 and > Long-context serving efficiency. Soofi S combines frontier-level capability with the highest measured aggregate long-context decode TPS, and unlike full-attention dense baselines maintains high throughput as context grows. Panel (1(a)) plots Capability Index versus measured aggregate decode TPS/GPU at 40K context and batch 32. The Capability Index averages five benchmark groups, i.e., Code, GSM8K, GPQA-Diamond, English aggregate, and German aggregate, after normalizing each group to the best plotted model. Aggregate decode TPS/GPU is measured with a TP=1, one-B200 vLLM latency-subtraction protocol. Panel (1(b)) shows measured aggregate decode TPS/GPU as a function of input context length under the same batch-32 protocol.

it's a small win in the small model class

zurfer··on GLM 5.2 and the coming AI margin collapse
Somehow the blog post seems naive. Yes GLM 5.2 is good and cheaper per token, but margins are a result of supply and demand. Now demand for quality and quantity of tokens is increasing at least quadratic or cubic (more users * more tasks * more tokens per task). On the other side you have real infrastructure constraints on the supply side. Openai and Anthropic have large commitments and contracts that enable them to get access at a scale of compute that is not obviously going to be available for open source model hosts. And you see it, glm 5.2 inference is less stable and higher variance than any of the bigs labs.

Why is SpaceX not hosting glm 5.2? because they make more money with renting out to Anthropic and Google.

zurfer··on Zuckerberg 'Admits' Meta's Layoffs Were Ineffective
> So he throws billions at a few top AI researchers, but they produce nothing of value.

so he spends 1% of yearly revenue on AI talent to catch up? we can't judge if they have produced nothing of value, no? They don't owe the world to open source their work?

Meta has plenty of failings, but taking risks and investing optimistically is not on my list. I guess the sentiment here on HN is probably biased by the addictive nature of its products.

zurfer··on Google loses fight over record $4.7B EU antitrust fine
ASML, SAP, ARM, Spotify, ...

How do you define biggest? Can't be by market cap.

zurfer··on OpenAI unveils its first custom chip, built by Broadcom
This is called mechanistic interpretability. There is lots of fascinating insights already since you can do basically everything down to the neuron or weight level thousands of times. The human brain is many orders of magnitude harder to make sense of.
zurfer··on Founding a company in Germany: €9600, 152 days and I still can't send an invoice
We outsourced it for 2.5k (extra) and it was still painful, took almost 2 months and worst of all wasted so much time and focus.

The worst was sitting at the notary, and getting read out loud by her what we were about to sign (also paying for that).

If you think about starting a company, spend some time to think through what it would mean for you to be a Delaware C Corp or an Estoinian one. It will increase your chances of success as you can focus on what matters.

zurfer··on DuckDB Internals Part 1
Like sqlite, duckdb is underappreciated as a production database. You can totally run it on servers or even "serverless" and do some heavy data transformations or with the right server size work with large scale datasets (up to a TB compressed seems fine).
zurfer··on Midjourney Medical
I heard the same argument from my doctor when I wanted a blood scan.

But what's the intention? If you do a scan and then try to find everything that is wrong about you, you're 100% right, there will be false positives and unnecessary panic/medication etc.

However if you just collect data for months and years and WHEN you get a symptom you have a lot more data then we should be able to give better diagnosis faster. If we do that for long enough as humanity and there is data sharing the accuracy of the whole thing will increase a lot.

zurfer··on Databricks Launches LTAP: A Unified OLAP/OLTP Data Architecture
Exactly, like SAP HANA stores everything in memory, you get great analytical and transactional performance but good luck financing that at scale
zurfer··on Anthropic requires 30 day data retention for Fable and Mythos
It was a red herring.
zurfer··on AMA: I'm Eric Ries (The Lean Startup) & Author of New Bestseller Incorruptible
Is LTSE working the way you hoped?
zurfer··on CEOs who think AI replaces their employees are just bad CEOs
I do think it's more subtle. AI can replace very few jobs end to end with the same quality, maybe none. But AI can be put to work on high ROI problems. Now when the new marginal job is not obviously as high ROI as putting another 100k of tokens to work, no human gets hired.

Next, comes natural attrition in a company where a certain percentage will leave every year. Will they get replaced with a human or their budget goes into tokens?

Only when these 2 angles are exhausted, a typical company will start thinking about layoffs.

Now, some companies are already stressed: customer buy AI products instead of theirs, AI makes it easier to build what they offer, customers believe they can vibe code things. These companies will layoff first, because of AI. Not because AI will do the persons job but because the money gets spend differently.

Page 1 of 11Next →