HNHacker News
TopNewBestAskShowJobs

vikramkr

6,185 karma · joined February 9, 2017

submissionscomments
vikramkr··on Tao: Open math problems being non-renewably mined by AI
> Replicating a paper is just as valuable scientifically as publishing it, but how many careers advance through replication?

A lot. In fields where knowledge is incrementally building on previous work the reason the whole field hasn't collapsed from the replication crisis is that usually the results that are really high impact are replicated in as an initial step in new research building on it. It's almost never the focus of the paper but you'll often find a quick mention in methods/supplemental of some previous work that was verified to be valid by a replication of a key technique etc. you'll have crisis where old tools are found to be problematic and findings end up revisited etc. Plus fields like clinical research where there's an awful lot of focus on replicating findings using staged clinical trials with increasing statistical power to determine if new interventions work - that's driven by regulatory requirements grounded in good science and a lot of people make careers in just that.

vikramkr··on Tao: Open math problems being non-renewably mined by AI
Not an expert by any means but the assumption here as I understand it is that the arxiv worthy PDF would not be acceptable or meaningful for impossible to understand proofs. And the lean proof would be meaningless unless the specific expression being proven is human understandable as the direct translation of the question the human is asking in formal form. So proving the negation is not a thing but if you make a subtle mistake in translating the statement you want to prove then obviously the QI is going to be proving the wrong thing. And otherwise you're relying on the correctness of lean as a system and on identifying/preventing if the proof is adversarially exploiting bugs in lean to falsely prove things.
vikramkr··on Please don't rearrange our shoes when we turn up, paramedics in Japan urge
Can't speak for op for but me no - it was just ingrained as something disrespectful to do to books. Sometimes you've got a taboo like eating with your left hand that has a pretty clear reason for coming into existence (you use your left hand to clean up in the bathroom and before hand soap in the era of everyone dying from cholera might as well keep those functions as physically separated as possible) but this is more just something you would be disrespectful to do
vikramkr··on Carbon-aware electricity pricing, measured daily on 38 grids
I mean yeah. Nuclear waste is much safer than fossil fuel waste - deaths caused, radioactivity, etc
vikramkr··on Path to Astra: critical capabilities and frontier safeguards
I'm so confused by the point you're trying to make. There's a lot of rhetoric about Georgia and an investigation about how the country list is 1:1 with some mysterious list from 1996 - but like - it's an export control list? Yeah, Washington approved the sale of missiles to Georgia. That's how that list works. Washington has to approve it. Openai is not Washington. They're perfectly reasonably erring on the side of caution and potentially over-complying with export controls. And if they then have to get approval from the feds to export to Georgia - well - our current administration has provided many reasons to over comply with trade related controls and not exactly been a champion of encouraging cross country trade right now. Idk what you expect from OpenAI or why you think it would matter at all that Georgia is a democracy or an ally of the US or is closely aligned with us. Ask Canada and NATO how much that's counted for with this administration.
vikramkr··on A third of Perplexity's citations don't contain the number they're cited for
> the intent is good

What is that supposed to mean? They're trying to be an llm search engine that's not some radical new concept

vikramkr··on Qwen3.8-Max
no the answer is `npx ccusage` or any of the other of the trillion ways to see how many tokens you're getting and what the current subsidization rates are.
vikramkr··on Qwen3.8-Max
no the $200/mo subs are definitely infinitely cheaper. If you're stuck paying enterprise API prices though that's not the case. So for personal or business premium plan use there's no competition but api rate/enterprise there is. Still a ton of spend happening on e.g. bedrock and via api.
vikramkr··on Apple caught off guard by AI demand for Mac Mini and Mac Studio
if you're one of those enterprises it's perfectly reasonable to assume you're more likely to go have your procurement and legal teams negotiate with google for one of those weird boxes to run gemini on prem (https://cloud.google.com/blog/products/ai-machine-learning/r...) or other enterprise-y nonsense vs buying consumer hardware to run a chinese large language model in your network. I do not envy the person at a company having to get approval model by model because of weird open source licensing terms dealing with the "how do we know chinese models are safe" question (hopefully they're at least getting asked in the context of hooking it into an agent harness so there's some sort of plausible risk that necessitates the conversation - i can very much imagine it getting shut down to even run in a sandbox because chinese model + people being scared after the huggingface stuff). There's a lot of reasons to assume enterprise wouldn't be interested, and it's very very cool that they are IMO. Anthropic cut claude code rate limits - they dont seem to be see open source llms as a market threat (rhtetoric to the white house aside) yet but enterprises being willing to run consumer hardware to run local models could make things a bit more tangible
vikramkr··on Apple caught off guard by AI demand for Mac Mini and Mac Studio
>"the majority of the market for Apple devices does trust that they will not produce and sell a computer incapable of support their use"

the problem with that argument is that the vision pro exists, where they clearly overestimated demand, and where even among people with interest in VR and disposable income, the compromises on battery life and weight were actually too much to bear. Forecasting is just hard and you always need to be especially skeptical when you yourself are doing the forecasting for something you want to succeed. All those arguments for the neo line up in retrospect hindsight is always 2020

vikramkr··on Apple caught off guard by AI demand for Mac Mini and Mac Studio
I don't think you need the scare quotes. product lifecycles can take years and the actual neo demand was actually pretty insane. They were using it to soak up demand for binned a18 chips and it would have been irresponsible to forecast that they'd have the demand they did when there's another perfectly reasonable universe where 8 gigs was too much of a compromise and it flopped.

this article's also about enterprise demand specifically. That's a bit surprising to me as well frankly. I'd have thought the primary market for mac studios would be hobbyists/enthusiasts with a bunch of disposable income who are willing to pay 18k for a 512 gb machine to run glm 3.5 flash or 9k to run deepseek v4 flash locally. It's competing with a $200/mo subscription or renting server gpu time for open source models during a memory shortage - and idk if it's going to be powerful enough to train or fine tune so it's really just inference. seems reasonable to be surprised

vikramkr··on Don't enthusiastically agree to rewrite the system at work
I mean that's mostly true in theory - and we're assuming good faith from the author of the blog that they were honest and clear in their communications of course (definitely strongly agree with you on that point). But that's still a very optimistic view of the power dynamic. It's generally not meaningful to say that someone technically has the power to do xyz (rejecting a timeline) because nothing literally physically stops them from doing it when there are severe consequences (getting fired) that come from it.

Also, > tell them it's their responsibility if it fails, so they can't blame it on you when it does fail.

That's definitely not how any of this works. They can absolutely blame whoever they want and often have the power dynamic in place to get away with it. Especially in this case where they'll just say the author was given a senior title and a raise so it's on them to deliver and their attempts to shift the blame to management won't work to cover up the engineers failure to deliver etc. You just have no leverage as an engineer joining a company that doesn't care or doesn't need devs to the extent they go extended periods of time without a single one. And if you don't build that leverage within the company you pretty much have no option but to go somewhere else if put in a bad situation.

vikramkr··on Continuous Diffusion Language Models (CDLM's)
They were worried it would make it easy to generate a flood of misinformation. They were correct.
vikramkr··on How to build a diffusion language model
Google was at least trying. Wouldn't be surprised if the others were experimenting with it too. The bar is going to be a lot higher now for for any diffusion model to go from experiment to product since it needs to compete with stuff like glm 5.3 flash and Luna on cost/efficiency for a given quality of output, which is not going to be easy. If it was easy Gemini diffusion would have landed - if it requires a bunch of money and effort it has a much higher bar to make it to market, if it requires some clever breakthrough you have no way of knowing where that's going to come from or what it'll look like/if it'll even seem important when it happens
vikramkr··on SK Hynix CEO sees memory chip shortage lasting until 2030
That's like 3 and a half years away? I can see that - fabs don't get built overnight and bubbles aren't under any obligation to pop on a timeline. Wild that 2030 is not that far away. Sometimes it feels like we never left 2020.
vikramkr··on Why your boss doesn't seem to care about your tech career anymore
It's a new culture. Combination of AI making junior engineers redundant, new grad skill levels cratering if they used AI to skate their way through college, and filtering through job applications being functionally impossible with huge floods of perfect ai generated resumes etc leading people to fall back to hiring through their network instead of applicants from open listings (and senior engineers are going to have better developed networks). And at least personally (+ for people around me) - AIs increased productivity comes with increased cognitive load and workload.

Plus you don't want to take a risk on training a new grad anymore because usually they'll leave after a couple years, which was a good thing since you'd backfill with a new grad some other company had trained/ideas circulated in the industry/early career engineers built their networks and got exposure to different tech stacks and company styles etc. that's really tough now too since nobody else is hiring and training juniors so you would be wasting time and effort training some for some other company to hire away - nash equilibriulm vibes.

So yes something shifted- the new grad/junior market is in the gutter rn. And if AI plateaus in this range of capability idk what happens as the current generation of engineers age out/retire. The juniors that aren't being trained today are going to lead to a vacuum of senior engineers tomorrow.

vikramkr··on Don't enthusiastically agree to rewrite the system at work
That is a _very_ optimistic view of both the power dynamic at play and that level of accountability executives in that context would face. This is a company that previously had no developer - so being the person who best understands the code is not meaningful. If it was a company where that mattered, they wouldn't have survived without an engineer on staff. This is not a situation where the dev is going to have the actual power to "accept" timelines.

And in terms of accountability/feeling the impact of their poor decisions - the blog is already describing them throwing the engineer under the bus, telling them they need to pull their weight, blaming them for mistakes, etc. Management isn't going to feel the impact of anything. Theyve got a convenient fall guy right there and they're already setting him up to take the blame if stuff goes south.

Sometimes stuff just sorta sucks ¯\_(ツ)_/¯

vikramkr··on Google Bows to Trump, Renames 'Lake Ontario' to 'Lake America' in Google Maps
I really doubt they would have changed their policy for either of those lol. It's easiest to go along with whoever is in charge - they aren't in the business of overtly challenging the political order - they have lobbyists for that
vikramkr··on Google Bows to Trump, Renames 'Lake Ontario' to 'Lake America' in Google Maps
This is standard practice no? Anything with disputed names they show whatever name is the official one in the user's region. This renaming is absolute garbage but he is the president and Congress isn't doing anything so it is what it is for now. If you don't like it - I mean if you can change a name once you can change it twice - 2028 is just around the corner
vikramkr··on The growing divide between AI hype and software engineering reality
Weird article. Feels very 2025 for a 2026 article (both in ways it reads too anti AI to me and ways it seems not critical enough of ai)

>Also stop saying “please” to an LLM. It does not have any feelings.

It emulates having feelings and it's a next token predictor. The next token in a dataset where you respond to an engineer rudely and call them stupid is rarely said engineer locking in and delivering incredible code. You have to play along to get the output you want.

> Again, I recommend people try running small LLMs locally where temperature and other settings are fully exposed and configurable to see this themselves.

This is like saying you should experiment with a paper airplane to see why fighter jets are overrated. Also the general understanding of temperature is not super solid here - it's not just that its "too boring" without it - random sampling is required for the models to work.

>Why are benchmarks showing they're still improving?

Ironically this section is far too generous to LLMs and benchmarks. lLMs cheat and companies benchmax. Don't trust benchmarks. They lie

Also in the what I do section - are these using local LLMs too? "Sometimes looping llms on itself can make it fix its own errors" feels very 2025 - the modern state of things is more "we've given up on one shots and getting it to produce the right answer immediately, set it up with a test harness so it can fix its own mistakes and let it loop otherwise it won't work." Also "asked fellow developers to make sure their code is well structured, easy to follow and documented. LLMs unfortunately make it easier for people to cheat in this regard" - really? Easy to follow? Maybe gpt models with good steering but trying to get an anthropic model to speak coherently and clearly and write documentation that isn't incomprehensible slop is a Herculean effort

vikramkr··on Don't enthusiastically agree to rewrite the system at work
> "I joined this company after they'd gone a stretch without any developer, and the bug queue showed it."

I think maybe if you're at a company like that, you should be very cautious in suggestions/timeline/etc about literally anything and everything technical. Certainly much more than you would be at a company that has a basic level of software engineering competency. It's kind of a different world...

vikramkr··on I Signed Up for Claude Pro, Why I'm Canceling (and What I'm Using Instead)
"I stopped using software x for y"

> Blog post clearly generated by software x

vikramkr··on Why your local LLM feels dumber than it is
I find that I remember models being a lot better than they were, even when I remember them being not very good - because of a novelty factor ("whoa it can do that now?") mostly. And then I go back and look at them and its like, what how did I find this impressive.

A funny example - I remember thinking "yeah sonnet 3.5 is a really good coding model"

https://stack.convex.dev/using-cursor-claude-and-convex-to-b...

>Prompting Cursor to Scaffold my App: FAIL This was my first hurdle.

>It became immediately apparent that I would not be able to prompt my way through the entire process.

>While the tooling we have is undeniably powerful, it's not yet capable of completing most nontrivial tasks

It couldn't run pnpm install lmao. Opus 4.5 was a crazy jump

vikramkr··on Anthropic's best AI model struggles to attract users as cheaper tools thrive
Sonnet 5 is a trash model and stupidly expensive if you accidentally set reasoning tokens high - more expensive than fable - it absolutely should not be the default lmao. If the common coding tasks you use ai for is doable with sonnet or local qwen - you're either not using Claude code (if you are, you'll very quickly see that sonnet 5 in Claude code is not a model for "occasional subagent use" - spinning up subagents is the only thing it's good at and it does it way too much. It can spin up subagents and waste huge amounts of tokens but it can't write good code lol.) or you've got Claude code workflow that is very human in the loop where you are significantly steering and controlling the models. And in that case your default should be to use gpt. Claude models are stupid slow.
vikramkr··on Why your local LLM feels dumber than it is
Was there a more recent refresh or is this the model from a year ago? The frontier models were barely functional and almost useless a year ago (gpt oss was pre opus 4.5!) - I would be very surprised if the original drop is anything more than totally obsolete/irrelevant at this point
vikramkr··on Stop Making TUIs
Sure but in this case we're the devs. We're in charge of that and would include them if we're making a gui instead of a tui and want keyboard accelerator mappings.
vikramkr··on Feature Request: Support AGENTS.md
Oh ffs is that why it suddenly started running into a ton of permissions errors trying to read and write files outside it's sandbox (I think the auto mode classifier blocks bash commands that would be allowed as read commands) and runs into all this nonsense where it uses bash to read a file then later tries to use the write tool and gets blocked on "must read file before writing it" and stuff? I thought I was going crazy yesterday- like had something changed ov5or had I just somebody not noticed it was failing tool calls that badly for months until yesterday but it makes sense if it was just because of that system prompt update. That's so god damn annoying idk how many tokens are getting wasted in the past couple days on these failed tool calls but it's not a trivial number
vikramkr··on DeepSeek API Pricing Update
Always gonna exist during a period of rapid exploration and experimentation. The js/web dev world slowed down a lot and entered a steady state eventually - been years since react took over and nothing's displaced it since
vikramkr··on Lost my phone at the office. Claude suggested tracking Bluetooth signal strength
Find my exists for a reason. Idk when the last time I had my phone not on mute is. Years? How often do you hear ringtones in public anymore?
vikramkr··on Auto mode is now the default in Claude Code
I don't think the marginal new user is anxious about approving messages - I think they're quickly annoyed by permissions prompt they don't understand and quickly get in the habit of approving everything or figuring out how to set bypass permissions on
← PreviousPage 2 of 34Next →