HNHacker News
TopNewBestAskShowJobs

mmasu

197 karma · joined March 1, 2021

submissionscomments
mmasu··on Stateless MCP has recaptured my interest
In enterprise MCP allows users to access resources that could either be unsafe or impractical to consume via API or CLI. It is a powerful pattern, supported by virtually all clients (Cursor, Claude Code, Codex, whatever) and easily implemented in custom harnesses. If you don’t need it you don’t, but it has many useful applications. Stateless will make it a lot more practical to expand applications.
mmasu··on Any text-to-SQL benchmark should address difficulties of real-world data stores
also, the same metric can be used in different ways and for different purposes across a business, so you might need different validation rules and definitions.

Maybe one separate semantics validation layer could help, but costs 2x or possibly Nx if you need to recover and turn a wrong query into a correct one

mmasu··on Throwing AI-generated walls of text into conversations
I am so tired of these people, but it’s so sad they don’t understand themselves how ridiculous they are
mmasu··on Access to frontier AI will soon be limited by economic and security constraints
I think “bought” here is to be read as “financed from”, not bought in the literal sense.
mmasu··on The paper computer
this is such a great idea! well done
mmasu··on Tell HN: I'm 60 years old. Claude Code has re-ignited a passion
how would you suggest someone who just started their career moves ahead to build that “taste” for lean and elegant solutions? I am onboarding fresh grads onto my team and I see a tendency towards blindly implementing LLM generated code. I always tell people they are responsible for the code they push, so they should always research every line of code, their imported frameworks and generated solutions. They should be able to explain their choices (or the LLM’s). But I still fail to see how I can help people become this “new” brand of developer. Would be very happy to hear your thoughts or how other people are planning to tackle this. Thanks!
mmasu··on Iranian Students Protest as Anger Grows
maybe the fact that Persians != Arabs will improve their odds. Recent uprisings had more luck (i.e. Bangladesh), even if it’s too early to fully assess their success
mmasu··on Claws are now a new layer on top of LLM agents
I am reading a book called Accelerando (highly recommended), and there is a play on a lobsters collective uploaded to the cloud. Claws reminded me of that - not sure it was an intentional reference tho!
mmasu··on Anthropic officially bans using subscription auth for third party use
I agree with you - we are saying the same thing, by restricting their API or making less developer friendly, they want you to be captive in their UI. This might not be true for Anthropic or OpenAI - another child commenter made a comment about ads in CLI, I would not be surprised if in a while we will have product placements in LLM responses exactly as we have it in movies - not a plain ad but just a slightly less subliminal suggestion.
mmasu··on Anthropic officially bans using subscription auth for third party use
I think that these companies are understanding that as the barrier to entry to build a frontend gets lower and lower, APIs will become the real moat. If you move away from their UI they will lose ad revenue, viewer stats, in short the ability to optimize how to harness your full attention. It would be great to have some stats on hand and see if and how much active API user has increased decreased in the last two years, as I would not be surprised if it had increased at a much faster pace than in the past.
mmasu··on "Token anxiety", a slot machine by any other name
I will give you an example I heard from an acquaintance yesterday - this person is very smart but not strictly “technical”.

He is building a trading automation for personal use. In his design he gets a message on whatsapp/signal/telegram and approves/rejects the trade suggestion.

To define specifications for this, he defined multiple agents (a quant, a data scientist, a principal engineer, and trading experts - “warren buffett”, “ray dalio”) and let the agents run until they reached a consensus on what the design should be. He said this ran for a couple of hours (so not strictly overnight) after he went to sleep; in the morning he read and amended the output (10s of pages equivalent) and let it build.

This is not a strictly-defined coding task, but there are now many examples of emerging patterns where you have multiple agents supporting each other, running tasks in parallel, correcting/criticising/challenging each other, until some definition of “done” has been satisfied.

That said, personally my usage is much like yours - I run agents one at a time and closely monitor output before proceeding, to avoid finding a clusterfuck of bad choices built on top of each other. So you are not alone my friend :-)

mmasu··on France Aiming to Replace Zoom, Google Meet, Microsoft Teams, etc.
yesterday in an article here on HN i read a wonderful dutch proverb:

“trust arrives on foot and leaves on horseback”

seems it’s applicable to this case too. Sad to see decades of work being tore apart in a few months.

mmasu··on RIP Low-Code 2014-2025
a friend (and colleague, disclaimer) pushed this recently to github. It passes data through a duck fb layer exactly to avoid context bloat:

https://github.com/agoda-com/api-agent

worth taking a look to see multiple approaches to the problem

mmasu··on Data Leak Exposes 149M Logins, Including Gmail, Facebook
isn’t this easy for a potential attacker to mitigate, i.e. dropping from the address everything after the plus? it’s a known trick for gmail so i would not be surprised if an attacker knew how to get to the “real” address by cleaning it up.
mmasu··on Iran Protest Map
I disagree with this - there have been overthrowings that did not require weapons in the field (i.e. Egypt, Tunisia), while widespread weapons were likely to cause civil wars (Lybia, Syria). In these cases however the role of the army was key in forcing the rulers out (and in Egypt to replace them), which might be unlikely in the case of Iran.
mmasu··on Claude Code creator says Claude wrote all his code for the last month
I use the house analogy a lot these days. A colleague vibe-coded an app and it does what it is supposed to, but the code really is an unmaintainable hodgepodge of files. I compare this to a house that looks functional on the surface, but has the toilet in the middle of the living room, an unsafe electrical system, water leaks, etc. I am afraid only the facade of the house will need to be beautiful, only to realize that they traded off glittery paint for shaky foundations.
mmasu··on Measuring AI Ability to Complete Long Tasks
this is a very good point, however the risk of writing bad or non extensive tests is still there if you don’t know what good looks like! The grind will still need to be there, but it will be a different way of gaining experience
mmasu··on Measuring AI Ability to Complete Long Tasks
I remember a very nice quote from an Amazon exec - “there is no compression algorithm for experience”. The LLM might as well do wrong things, and you still won’t know what you don’t know. But then, iterating with LLMs is a different kind of experience; and in the future people will likely do that more than just grinding through the failure of just missing semicolons Simon is describing below. It’s a different paradigm really
mmasu··on The Grand Egyptian Museum's Astonishing Arrival
Morsi did not have much time to do anything in all frankness - Sisi has now been there for 10+ years
mmasu··on Python 3.14 is here. How fast is it?
I too started with your tutorial - thanks a million
mmasu··on Vibe engineering
While I don’t agree with you, I keep a healthily skeptical outlook and am trying to understand this too - what is the hard data? I saw a study a while ago about drops in productivity when devs of OSS repos were AI assisted, but sample size was far too low and repos were quite large. Are you referring to other studies or data supporting this? Thanks!
mmasu··on Helsinki records zero traffic deaths for full year
is it also possible that one of the side effects of this are that people driving recreationally become sometimes exceptionally good at it? see how many great f1/rally pilots Finland has generated. Clearly not good when this happens while drunk tho
mmasu··on Build an AI telephony agent for inbound and outbound calls
what I mean is: building these systems is nontrivial, but if done well it can help. Imagine non being in an endless queue on a phone call when having to do a simple task through a customer center call, or having a phone reminder with more information and less noise than from a written notification. The fact that I failed at it (for lack of experience and resources) does not mean it should just be shrugged off as useless or impractical. Some companies offer this service and it works just fine for narrow use cases.
mmasu··on Build an AI telephony agent for inbound and outbound calls
we tried to build something similar lately for outbound calls (for simple reminders to partners) and faced massive issues using gpt-4o-realtime-audio. Noise detection, turn detection, random telephony issues (we were using Twilio too), prompt not holding together, and more.

We dropped the project because it would have resulted in a terrible experience for the person on the other side of the phone. Building these things is non trivial.

The plan would have been to A/B test and see what the response would have been (watching NPS and business metrics uplift). Human handoff was always the plan in case things got too tricky for the LLM to handle.

I see some hostility here towards this project and while I share many concerns, it is very naive to think that these services won’t be massively leveraged going forward. An AI agent can handle things as well as humans (not in our case but there are good services out there, i.e. Parloa) and the key elements are the same as all the other agentic based workflows:

- narrow use cases

- human in the loop ready to pick up/steer/correct

we will see a lot more of this and as LLM capabilities improve, it will only get better - it is inevitable at this point and might (_might_) result in a better experience for customers in some cases.

Nevertheless I also see the possibility that we will go full circle and we will always reach for a human, maybe showing up in person in a physical office to make sure cases or requests are handled well… or not :-)

mmasu··on Study mode
yesterday I read a paper about using gpt 4 as a tutor in italian schools, with encouraging results - students are more engaged, get through homework by receiving immediate and precise feedback, resulting in non-negligible performance improvements:

https://arxiv.org/abs/2409.15981

it is definitely a great use case for LLMs, and challenges the assumption that LLMs can only “increase brain rot” so to say.

mmasu··on Peep Show is the most realistic portrayal of evil I have seen (2020)
please Jez, don’t talk about crack!!
mmasu··on What the Fuck Python
the Python ecosystem for AI is far ahead than R, sadly (or not :-) )
mmasu··on OpenAI slams court order to save all ChatGPT logs, including deleted chats
yet - likely subscription grade will stay ahead of the curve, but we will soon have very decent models running locally for very cheap - like when you play great videogames that are 2/3 years old on now “cheap”machines
mmasu··on Cursor 1.0
why would you open 3 IDEs all at once :-)
mmasu··on My AI skeptic friends are all nuts
While I agree with your suggestion, the comparison does not hold: calculators do not tell you which numbers to input and compute. With an LLM you can just ask vaguely, and get an often passable result
Page 1 of 3Next →