HNHacker News
TopNewBestAskShowJobs

alansaber

824 karma · joined June 13, 2025

submissionscomments
alansaber··on Getting the most out of Opus 5.5 in Claude and Claude Code
"Don’t ask it to show its reasoning in the reply" “Explain why you chose this approach in three sentences” says it all really
alansaber··on Language models for text classification: From bag-of-words to Jev
Agreed as well, this is my preferred framing, it only makes sense to get into semantic vectors once you realise the very real limits of discrete vectors and bag of words.
alansaber··on Dots: Always-on agents
Just sounds like a more user friendly interface for projects (rather than a folder, how about the adorable green dot for my cybersecurity questions?)
alansaber··on A Privacy Analysis of Web and Mobile Conversational AI Agents [pdf]
AI companies rushing an implementation? Surely not :).
alansaber··on Show HN: HN.watch – Videos of all Hacker News posts
I like video as a format because it's pattern-breaking. Consuming thousands of very similarly produced videos sounds even more depressing than consuming LLM blogs and articles. EDIT: That aside, this is of course extremely impressive.
alansaber··on Sonnet 5.5
When they inevitably drop allocation after post-launch hype dies down.
alansaber··on Sonnet 5.5
Always key to include the one bench where the smaller model inexplicably outperforms the larger model
alansaber··on Maybe don't let Muse run your Facebook Marketplace account
They're building for future capabilities. It's a logical approach to leveraging the data they have.
alansaber··on Maybe don't let Muse run your Facebook Marketplace account
I don't much care for it on desktop either
alansaber··on Thinking fast and slow in AI: The role of metacognition (2021)
Given how LLMs access compressed knowledge from their model weights, the similarities make sense
alansaber··on Ember-1
Exactly, competition between both frontier model companies and the chinese labs is the primary factor suppressing consumer prices.
alansaber··on There is more to code review than (automatable) detection
Outdated/bad CICD is surely extremely common in every org
alansaber··on There is more to code review than (automatable) detection
So I generally agree with this with some exceptions: these proficiencies will still exist, but move away from standard engineering roles, to highly specialised individual roles, just as there are still specialists in assembly and other low level languages. They'll just have roles maintaining the nuclear arsenal etc etc where artisanal skill is still an absolute requirement.
alansaber··on Prompting Claude Opus 5.5
"a lot of the models work was essentially coordinating everything." - I don't see anything wrong with that personally. It's still extremely challenging to build a model harness, and having a model-mediated everything is clearly wishful thinking. It's exhausting to keep up with, but also somewhat exciting, all depends on your perspective of course.
alansaber··on The Normalization of Inexplicable Failures
TBF OP makes a good point: skipping out on the whole fine-tune sing and dance and diving directly into a good general classifier is a barrier to developing actual, deep intuition for the problem space (a pretty prevalant anti-AI argument).
alansaber··on Unsealed Briefs in Authors’ Case v. Microsoft/OpenAI
Not even openai want to get flamed on show hn
alansaber··on Opus 5.5 is good at explainer videos
Recent models have been a big step up on graphics generation. Not exactly sure why, it would be neat if they would release any quantity of technical blogs.
alansaber··on Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)
I'm amazed they didn't test xhigh thinking mode explicitly to ensure it didn't exceed the 128k thinking budget allocation. I guess pace of development gets away from everyone, even OpenAI.
alansaber··on GPT-6 Sol and Luna
I felt that was about 5.5. IMO 5.6 Sol was overindexed: more verbose, prone to overengineering.
alansaber··on GPT-6 Sol and Luna
It's extremely variable because the products are roughly equivelant, and a lot of the quality of service depends on their inference capacity at any given hour/day.
alansaber··on GPT-6 Sol and Luna
It would be extremely funny if the explosion in SVG generation capability in particular was a result of this benchmark
alansaber··on AI coding has made CI a bottleneck, so we reworked ours to keep up
Scope explosion. AI is really bad at polish, requiring human intervention. So everyone reallocates effort to breadth not depth, effectively mirroring AI capability.
alansaber··on AI Has No Wisdom and Neither Will You
"New technology: cool but the supply chain is vulnerable". A tale as old as time.
alansaber··on AI Has No Wisdom and Neither Will You
"My project isn't open sourced yet" underlines the problem. There has been a massive explosion of software that works, but badly. The average user experience suffers as a result.
alansaber··on Transformers Explained Visually
Not dumb questions. These design decisions are based on years of applied testing, more than any theoretical result.
alansaber··on Grok 4.7
As anthropic/openai subscription allocations get squeezed you'll see more people using "second rate" closed models like grok. The token allowance with a Cursor subscription is crazy.
alansaber··on Attention is all you have
This is why I like books. No attempt at attention hijacking. Wheras every other modern app tries to claw your eyes out.
alansaber··on Don't Use AI to Write
Yes but on the sliding scale of how deterministic is this task- writing, not very. coding, can execute and write some tests (even if gamed/bad).
alansaber··on Don't Use AI to Write
This is a neat workflow. Incentivises you to proactively think about your writing, even as you cut it.
alansaber··on Don't Use AI to Write
Well said, the issue being, the form factor of many AI products disincentivises this behaviour. To give an example, using a CLI tool with a frontier model to write a blog (actually fairly common), you have to re-open it in another editor, potentially reformat it or convert MD>DOCX, etc etc. There's a lot of noise and mess, and people naturally tend towards "not bother" and 1-shot some "amazing" prompt (because the value of their elbow grease is disproportionately low, unless they actually have something to say- many do not).
Page 1 of 24Next →