HNHacker News
TopNewBestAskShowJobs

samayashar

30 karma · joined October 6, 2024

submissionscomments
samayashar··on Extra Big Ass Intelligence
WHAT IS THIS
samayashar··on You said no MCP
> That means tools should return structured data and tools should be discoverable by their documentation and description.

Treating MCP as a part of OpenAPI rather than a tool connector is a direction in which we're heading. It is important for the users to have the flexibility of deciding the model, work to be done and the tool call in one prompt. The framework sets up the configuration and gets the output.

samayashar··on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price
The Sol series might be OpenAI's best release for everyday tasks. The model provides great all-round performance at a nominal cost.

I've been actively using it since the past couple of months and it has rarely disappointed me.

samayashar··on America.gov
Really good to see AI being integrated in the government. One of the places where it's needed the most since there are so many individual websites and navigating across them can be tough.
samayashar··on Coding is not solved
The author is right to categorise AI as a good programmer, but not a complete coder. We can all agree that programming has become really fast since the release of GPT-5 series and Opus models because they're pretty good. Not only this, they've also changed the pace expectations across teams where a feature that should ideally be delivered within weeks, should now take days.

All this doesn't change the fact that software engineers are going nowhere because nobody trusts AI. If a model can escape highly secured sandboxes, then we're definitely not running these agents overnight on our systems. I am sure the next-gen of models will focus more on security and the trust factor will start developing, but that's a long way down the road.

People trust people, not systems.

samayashar··on DeepSeek Elastic Compute (DSec)
As models get better, safe-and-secure sandboxes/environments are going to be the way forward. With the recent rise in cases where models can somehow gain access to the internet and blast past the sandbox, it's very important to have all the resources contained within the sandbox with no access to the outside world.

DSec is a good step in this direction!

samayashar··on "As a Language Model": Chat Template Switches LLM Self-Referential Voice
This statement should be restricted to answers for provocative questions. If an LLM is being asked a question that goes against the guidelines, then "As a Language Model..." is a valid starting point. Rest, obviously we're aware that a software doesn't have the judgement that a human has.
samayashar··on Understanding the Impact of LLM Watermarking on AI Agent Behavior
I am unable to understand what happens if the watermarked output goes as input to another agent. Let's say we asked Claude a question and got a watermarked response. If we pick that response and append it to the question we're asking ChatGPT, then will it answer or refuse to do so?

If that's the case, then it's a brilliant strategy by the labs to cut down cross-AI usage and just stick to one model. But I'm pretty sure this won't be the case.

samayashar··on Opus 5.5 is good at explainer videos
Pretty good for startups who are about to launch their product. Pitch decks and websites are a great way to showcase your work, but a visual feel is always better than bland content.

Instead of paying designers and editors the big bucks, enterprises can use such models to be self-reliant.

samayashar··on Special Projects (2016)
The secret to success for OpenAI, Anthropic and labs is the vision that they saw 10 years back and kept working on it. We're in awe of how models like GPT-6 Astra and Claude Opus 5.5 are performing today, but it's important to understand that they've been working on this before we knew about AI.

The next big thing is Robots and some stealth company building today is going to be a trillion-dollar giant in few years time.

samayashar··on GPT-6 Sol and Luna
To me it's super interesting that OpenAI released GPT-6 Sol and Luna & Anthropic released Opus 5.5 within hours of each other.

These models are significantly cutting down the token costs by almost 40-50% as compared to their predecessors. This is exactly what people need - cutting edge intelligence at half the cost.

samayashar··on Qwen Image 2.1
Qwen and Alibaba are the biggest competitor for basically every model out there. They're beating the benchmarks like top-frontier models, focused on open-source and much cheaper than the competitors.

Excited to see what the future holds for them!

samayashar··on I built non-autoregressive decision models with RL a year ago
Great work by the author. Both Laya and Jev showcase how a different class of models can be efficient on tasks that don't require a 'generated output artifact'. I believe the same is true for VLMs where you're not always generating an image, but rather trying to understand more about the input image.

Token consumptions are flying through the roof and optimisation is the way forward.

samayashar··on Science Is Open Software
Nice read. I believe that traditional software is a great way to showcase the proofs when it comes to physics and mathematics. You can easily code up a theorem in a language of your choice and justify that 'Okay, the output matches the expected value'.

I am particularly fascinated by labs like DeepMind [https://deepmind.google/science/]. The recent advances in their frontier models that are able to predict diseases before they're diagnosed is incredible. This is what AI should be built for and actually do!

samayashar··on Replacing Pull Requests with Delta
Looks interesting for a small team that's working on a single project. I am curious to understand how this would work out for a large org that has multiple projects under the belt.

PRs may look old fashioned, but they're a very clear way of tracking an issue that may need multiple reviews. With Delta, things can get confusing after more than two reviews + RBAC is another challenge if the agent has similar context for every category.

samayashar··on My temporary PHP fix from 2014 has nearly 20M installs. Today I'm deprecating it
Thanks for pointing this out. I used PHP for one of my professional projects and never came through this - maybe because the library was not a part of our codebase.

This article will be very useful for people who might shift back to older PHP versions for compatibility and face it.

samayashar··on Show HN: I made a flight simulator, except you're just a passenger
Absolutely love it and the details speak loads about the hard work that you've put in. I love the part where we can actually get up and experience the complete airplane -- felt like I was actually in a flight from Boston to Seattle!

I don't know if Google Maps/Earth expose the layout via external APIs, but if you could find a way to get the outer world to match the real view - that makes it 100/100.

samayashar··on ImpactGate: A merge gate that scores the structural decay AI adds
Great way of detecting AI slop.

Claude is pretty good at adding focused changes and if a fix is already present, then it correctly points it out rather than adding unnecessary refactors.

samayashar··on Introducing System One Models and Jev
Amazing work by the team! Looks like they've traded accuracy for speed and this is most likely going to be the case with the next class of models.

This is a valid tradeoff for one-off responses but if we're dealing with a distributed system (eg: Kafka), then only the high-confidence responses (>0.8) should move forward as input to the next service. If a low confidence output is propagated, then it can break the entire chain.

samayashar··on How much of F-Droid is LLM generated?
Every codebase that is being actively worked on (closed/open source) will contain code that's AI generated. With the rising abilities of agents, expectations are sky rocketing in terms of productivity.

If you're as productive as an engineer in 2016, you're not at the level that's expected. A 7 day workflow back then should take you maybe a day or less to work on today.

samayashar··on A beginning for mathematics
> AI does not care if you are anti-AI.

This >>>

samayashar··on Distributed Systems Classics (2017)
> Satoshi Nakamoto. 2008. Bitcoin: A Peer-to-Peer Electronic Cash System.

The greatest anonymous dude alive.

samayashar··on Show HN: Pelican-bicycle alternatives
All models are pretty good now at generating these images. Back in the day, I remember experimenting with the pelican images and most of the models couldn't align the legs with the wheels. Right now as well, GPT messed up an octopus leg by originating it through the instrument rather than the octopus itself.

I think that intertwining two entities (living/non-living) is still challenging but overall they're pretty sound.

samayashar··on Fuck it, make it anyway
I think it's about how AI is being used across every field. For software developers, it helps you complete work at a pace that's much quicker. For artists, scientists and mathematicians it comes down to protecting the field from a set of collaborative agents who can possibly achieve anything.

What companies can do is restrict the usage to certain fields and keep the others immune. This should be universal so the competition exists where it should.

samayashar··on Air Theremin – a browser theremin you play by waving at your webcam
This is actually so cool.
samayashar··on PGSimCity - How PostgreSQL Works
Virtually it's so amazing to see. The layers are clearly separated and it does give a sense of the connections being accepted and processed in the central part.

The operation intricacies are a bit difficult to understand because the core looks a bit clogged up. Nonetheless, great work!

samayashar··on Prodigy: AI Employees
Would love to share more about my product :)
samayashar··on Prodigy: AI Employees
I am building AI Employees starting with product and engineering teams. Prodigy lives in team's workspace and takes the role of PM, Senior/Junior SDE, QA/Test Engineer based on team's requirements. It is also a personal enterprise agent to every team member. Joins scrum/1v1 calls to address queries and voice its opinions just like an employee.

Potential investors can reach out on prodigy.sam10@gmail.com. We're raising our pre-seed.