HNHacker News
TopNewBestAskShowJobs

aliljet

1,538 karma · joined February 22, 2016

contact me here: pav.gup@gmail.com
submissionscomments
aliljet··on Claude Memory
I really want to understand what the context consumption looks like for this. Is it 10k tokens? Is it 100k tokens?
aliljet··on With deadline looming 4 of 9 universities reject Trumps pact to remake higher ed
Having casually attended one of these schools, I'm so confused about why they are even in this group. What is making this group of schools best suited for this sort of blackmail?
aliljet··on Claude Skills are awesome, maybe a bigger deal than MCP
If this is true, what is the Playwright Skill that we can all enjoy with low token usage and the same value?
aliljet··on IRS open sources its fact graph
I wonder how this can be used with an LLM to provide interesting tax advice? I'd love to regularly ask questions of the tax code...
aliljet··on Claude Haiku 4.5
What is the use case for these tiny models? Is it speed? Is it to move on device somewhere? Or is it to provide some relief in pricing somewhere in the API? It seems like most use is through the Claude subscription and therefore the use case here is basically non-existent.
aliljet··on Claude Sonnet 4.5
These benchmarks in real world work remain remarkably weak. If you're using this for day-to-day work, the eval that really matters is how the model handles a ten step action. Context and focus are absolutely king in real world work. To be fair, Sonnet has tended to be very good at that...

I wonder if the 1m token context length is coming for this ride too?

aliljet··on Context is the bottleneck for coding agents now
For me? It's simple. Completely empty the context and rebuild focused on the new task at hand. It's painful, but very effective.
aliljet··on Context is the bottleneck for coding agents now
There's a misunderstanding here broadly. Context could be infinite, but the real bottleneck is understanding intent late in a multi-step operation. A human can effectively discard or disregard prior information as the narrow window of focus moves to a new task, LLMs seem incredibly bad at this.

Having more context, but leaving open an inability to effectively focus on the latest task is the real problem.

aliljet··on The Harvard-Emory ECG Database
I work deep in the weeds on startups using datasets like these and I'm always curious about how people drive value from this raw data. Who is designing things that use this data for anything interesting in the United States? The whole system is designed NOT TO PAY for these advancements. What value if any exists in this data beyond academic research?
aliljet··on Apple Watch Ultra 3
Is it? Across what metric?
aliljet··on Python: The Documentary [video]
For better or worse, I find Python beautiful. And here's an easter egg for everyone (it comes up in the Documentary too):

  import this
Even on HN, indentation reigns king.
aliljet··on Claude for Chrome
Having played a LOT with browser use, playwright, and puppeteer (all via MCP integrations and pythonic test cases), it's incredibly clear how quickly Claude (in particular) loses the thread as it starts to interact with the browser. There's a TON of visual and contextual information that just vanishes as you begin to do anything particularly complex. In my experience, repeatedly forcing new context windows between screenshots has dramatically improved the ability for claude to perform complex intearctions in the browser, but it's all been pretty weak.

When Claude can operate in the browser and effectively understand 5 radio buttons in a row, I think we'll have made real progress. So far, I've not seen that eval.

aliljet··on Building A16Z's Personal AI Workstation
I don't want to be insulting here, but have you sat down with a partner at a VC before? You may be surprised to discover their skill is rarely deeply technical...
aliljet··on Building A16Z's Personal AI Workstation
This article is a great way to showcase A16Z standinging head and shoulders above other VCs with REAL technical expertise in the partnership. Love reading this kind of stuff, but the article really needs a price to put this in perspective. It would be the VC that would ignore price, lol, but roughing this out it looks like it costs 45k to build this thing. Seems, at first glance, that this is a cost efficient way to dodge buying a Kia Carnival and get a tier-1 GPU workstation..
aliljet··on Claude Sonnet 4 now supports 1M tokens of context
How do you know that?
aliljet··on Claude Sonnet 4 now supports 1M tokens of context
This is definitely one of my CORE problem as I use these tools for "professional software engineering." I really desperately need LLMs to maintain extremely effective context and it's not actually that interesting to see a new model that's marginally better than the next one (for my day-to-day).

However. Price is king. Allowing me to flood the context window with my code base is great, but given that the price has substantially increased, it makes sense to better manage the context window into the current situation. The value I'm getting here flooding their context window is great for them, but short of evals that look into how effective Sonnet stays on track, it's not clear if the value actually exists here.

aliljet··on Perplexity Makes Longshot $34.5B Offer for Chrome
For the simple minded among us, can someone explain why this would be worth 34.5 billion dollars? Wouldn't the fork (https://www.perplexity.ai/comet) be sufficient?
aliljet··on GPT-5: Key characteristics, pricing and system card
I'm curious what platform people are using to test GPT-5? I'm so deep into the claude code world that I'm actually unsure what the best option is outside of claude code...
aliljet··on GPT-5 for Developers
Between Opus aand GPT-5, it's not clear there's a substantial difference in software development expertise. The metric that I can't seem to get past in my attempts to use the systems is context awareness over long-running tasks. Producing a very complex, context-exceeding objective is a daily (maybe hourly) ocurrence for me. All I care about is how these systems manage context and stay on track over extended periods of time.

What eval is tracking that? It seems like it's potentially the most imporatnt metric for real-world software engineering and not one-shot vibe prayers.

aliljet··on GPT-5
The eval bar I want to see here is simple: over a complex objective (e.g., deploy to prod using a git workflow), how many tasks can GPT-5 stay on track with before it falls off the train. Context is king and it's the most obvious and glaring problem with current models.
aliljet··on GPT-5
Honestly, you're probably right. It's quickly become a pretty weak eval, but the guy that's running that eval is excellent. I'd much rather the evals people were using to test these things looked more like classic/boring engineering problems: deploy to dev/test/stage/prod with digital ocean, cloudflare, github, and a common git flow. Boring problem, I know, but that problem is wildly complex when you start to add a few extra dimensions (frontend vs backend, ports shifting between deployments, local deployments, etc.).
aliljet··on GPT-5
It's very unclear if OpenAI has been casually leaking things to create buzz, but a few days ago there was a pretty stunning pelican on a bike attempt: https://old.reddit.com/r/OpenAI/comments/1mettre/gpt5_is_alr...

In practice, it's very clear to me that the most important value in writing software with an LLM isn't it's ability to one-shot hard problems, but rather it's ability to effectively manage complex context. There are no good evals for this kind of problem, but that's what I'm keenly interested in understanding. Show me GPT-5 can move through 10 steps in a list of tasks without completely losing the objective by the end.

aliljet··on Claude Code weekly rate limits
Maybe this is an unpopular opinion, but it seems like Anthropic has quietly 4x'd the real cost of the Pro plan. There are 168 hours in a week, and if I'm able to (safely) bet on 40 hours of use, realistically, I just lost 75% of the value of the plan.

What are the reasonable local alternatives? 128 GB of ram, reasonably-newish-proc, 12 GB of vram? I'm okay waitign for my machine to burn away on LLM experiments I'm running, but I don't want to simply stop my work and wake up at 3 AM to start working again..

aliljet··on Qwen3-235B-A22B-Thinking-2507
I see the term 'local inference' everywhere. It's an absurd misnomer without hardware and cost defined. I can also run a coal fired power plant in my backyard, but in practice, there's no reasonable way to make that economical beyond being a toy.

(And I should add, you are a hero for doing this work, only love in my comment, but still a demand for detail$!)

aliljet··on I'm Peter Roberts, immigration attorney who does work for YC and startups. AMA
What is the likelihood today that Trump and his allies in Congress may improve the ability for H1B holders to obtain green cards? This was repeatedly raised during his campaign and his first weeks in office and it sounded like this was going to be a thing....
aliljet··on AWS launches Kiro, its Cursor clone
I'm a little confused about how pricing works here. What is an 'agentic interaction' and how does that translate to dollars? And how does this work with models that are differently priced???
aliljet··on Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model
If the SWE Bench results are to be believed... this looks best in class right now for a local LLM. To be fair, show me the guy who is running this locally...
aliljet··on Grok 4
[edit to focus on pricing, leaving praise of Simon's post out despite being deserved]

Simon claims, 'Grok 4 is competitively priced. It's $3/million for input tokens and $15/million for output tokens - the same price as Claude Sonnet 4.' This ignores the real price which skyrockets with thinking tokens.

This is a classic weird tesla-style pricing tactic at work. The price is not what it seems. The tokens it's burning to think are causing the cost of this model to be extremely high. Check this out: https://artificialanalysis.ai/models/grok-4/providers

Perhaps Grok 4 is the second most expensive and the most powerful model in the market right now...

aliljet··on Buffett to step down following six-decade run atop Berkshire
How much of Berkshire actually relied on Buffet toward this announcement? I'm increasingly suspect of 90+ or 80+ adults operating the machinery of massive entities. Lots of examples of this.
aliljet··on Hackers Crack Subaru's Connected Services to Access Data, Door Locks and More
I am so curious about taking over running the services that perform this for my car? Shouldn't I be able to issue commands to my car myself?
← PreviousPage 4 of 9Next →