The difference is that the worker you hire would be a human being and not a large matrix multiplication that had parameters optimized by a a gradient descent process and embeds concepts in a higher dimensional vector space that results in all sorts of weird things like subliminal learning (https://alignment.anthropic.com/2025/subliminal-learning/).
It's not a human intelligence - it's a totally different thing, so why would the same test that you use to evaluate human abilities apply here?
Also more directly the "all sorts of other things" we want llms to be good at often involve writing code/spatial reasoning/world understanding which creating an svg of a pelican riding a bicycle very very directly evaluates so it's not even that surprising?
And why should the workers who work on the bridge be denied happiness and satisfaction from their work? Building and creating physical stuff is incredibly rewarding in concept for so many people - especially in a culture that values/glorifies physical and manual labor like parts of the US. I mean bob the builder is a popular kids show and "all boys are fascinated by big trucks and construction projects" is both an incredibly common stereotype and to a significant extent just a true statement.
If you want to learn it the best way is probably to come up with a personal project idea that requires it specifically? Idk how much you'd get out of a book but you could always do a side project with the specific goal of doing it just to learn a particular stack or whatever
And what about the predictions of energy use that did pan out, like air conditioning and stuff? Also in 1999 how many personal computer companies were restarting nuclear power plants to fuel their projected energy consumption? Feels like a weird argument to make when the investments into AI I fra are literally measured in gigawatts. Feels like a weird argument in general - ai consuming lots of energy isn't some weird degrowth conspiracy theory
In addition to stuff like that they also handle it with rate limits, that message that Claude would throw almost all the time when they were like "demand is high so you have automatically switched to concise mode", making batch inference cheaper for API customers to convince them to use that instead of real time replies. The site erroring out during a period of high demand also works, prioritizing business customers during a rollout, the service degrading. It's not like any provider has a track record for effortlessly keeping responsiveness super high. Usually it's more the opposite.
People like it when an IPO pops. It's a good news story and it makes all the banks who participated happy. If it was priced perfectly it'd get reported as the stock was flat, if it's a bit underpriced then you get headlines as the hot new stock that's taking off
The bit about smaller facilities being exempted in the current proposals does seem like a genuine issue and a great way to end up with hundreds of unregulated inefficient small data centers exploiting loopholes and causing all sorts of issues.
There's also a chance that the primary photosynthesiizers on each happened to be purple for a while (purple earth) and the ancestors of plants absorbed red/blue and ignored green because they were getting leftovers. Also, even now, iirc the limiting step in oxygenic photosynthesis is by far rubisco's incorporation of CO2, so there's no immediately obvious fitness function that would be optimized by just increasing the efficiency of light harvesting.
That's calculating value against not having LLMs and current competitors. If they stopped improving but their competitors didn't, then the question would be the incremental cost of Claude (financial, adjusted for switching costs, etc) against the incremental advantage against the next best competitor that did continue improving. Lock in is going to be hard to accomplish around a product that has success defined by its generalizability and adaptability.
Basically, they can stop investing in research either when 1) the tech matures and everyone is out of ideas or 2) they have monopoly power from either market power or oracle style enterprise lock in or something. Otherwise they'll fall behind and you won't have any reason to pay for it anymore. Fun thing about "perfect" competition is that everyone competes their profits to zero
I mean obviously it depends, a lot of that 3mb is probably doing all sorts of ad and tracking stuff which, well, welcome to the internet. There's a point where trying to avoid JavaScript and build tools is just plain harder (especially with large teams) than like, a basic react app with typescript. You can use astro or just render out HTML and serve that. No reason to not use the tools you're already familiar with. If it's a webapp, the demands of complexity on the app mean you almost certainly need a framework (especially working with a team) to manage that. If it's a blog, if you make the active choice to make it slow and painful to load as a developer and put in the work to add all that bloat, it doesn't matter that much anyway?
These articles always suggest that minimal HTML and jQuery is enough. Like, I guess for a blog, but like, not for web apps? That's the other weird thing - they're always framed like the entirety of the internet is blogs or whatever. How many software engineers are working on blogs? Everything is a webapp now, the vast majority of the time most users spend interacting with the internet is through complex applications like Spotify, YouTube, Gmail, Netflix, Google docs, notion, linear, etc etc etc. I guess there's an audience for complaining about react for some reason
Did he predict they'd always be necessary? He mostly seemed to predict the opposite, that we're at the early stage of a trajectory that has yet to have it's Linux moment
If you close off the market to US tech giants, maybe they'll have some amount of market dominance at home, but I would doubt that would mean they've "caught up" tech wise. There would be no incentive to compete. American EV manufacturing is pretty far behind Chinese EV manufacturing, protectionism didn't help make a competitive car, it just protected the home market while slowly ceding international market after international market
Always hard to tell what and where specifically the impacts of any research will be. Looking at how they discuss the creating of this lnp formulation in the paper:
> We therefore modified the lipid composition of the LNP to enhance potency. First, the ionisable lipid DLin-MC3-DMA (MC3) was replaced with SM-102, an ionisable lipid previously shown to lead to greater cytosolic mRNA delivery through enhanced endosomal escape [30]. Second, the SM-102-LNPs were further modified using ß-sitosterol, a naturally-occurring cholesterol analogue associated with enhanced mRNA delivery [31], to create a formulation referred to as LNP X (Fig. 1b).
Neither of those papers they cited were focused on HIV/T cells specifically. Maybe this will have no applications, or applications in HIV only, or be generally useful for hard to transfect cells, or be bad for HIV for some unforseen reason but useful elsewhere, or be a dead end, you never know. But yeah maybe if folks run into similar challenges for approaches to dealing with those viruses, maybe there's something from this work that could help them, who knows?
I mean, they're cheaper models and they aren't as much if a pain about rate limiting as Claude was/they have a pretty solid depenresesrch without restrictive usage limits. IDK how it is for long running agentic stuff, would be surprised if it was anywhere near the other models, but for a general chatgpt competitor it doesn't matter if it's not as good as opus 4 if it's way cheaper and won't use up your usage limit
The excessive comments might help the model when it's called again to re edit the code in the future - wouldn't be surprised if it was optimizing for vibe coding and the redundant comments reinforce the function/intent of the line when it's being modified down the line
That's not remotely correct- how would you be able to ignore tokens? That's literally what defines the context size of an llm, the larger the more memory. Generating each token is literally the compute you're doing. Your hardware limits how many tokens per second you can produce with a particular model. It's literally what you're consuming electricity to produce. That's like saying you can drive your own car for free without worrying about how many miles you've driven.
I don't think it's obvious that any of these model providers are even profitable right now. I'm also not sure what there is to "figure out" - it's an expensive technology where the cost scales per token, so they charge per token? would you rather they burned even more money giving it away for free until everyone was dependent on it and then hyper enshittified to try and not go broke like so much of the rest of tech?
looks like it - it's such a minor and brief mention in the paper for the article to focus on it so much lol. They probably should have cited it, looks like they decided it was minor enough (or forgot) that they didn't put it in their software used/citation. Super commonly used though I wouldn't be stunned if most of its uses never got cited- just a quick check if it thinks deleting some section or doing some sort of fusion is gonna cause a problem, or if you've got something without a PDB structure finding a site to mess with that looks like it's not gonna cause any problems. You can't count on it blindly obviously but it's super helpful. Like if it's pretty confident about some section of a protein it hasn't seen before, the weird stuff you're studying might not be folded properly by the model, but if you want to stick a handle onto the protein to grab it with whatever it can let you know where's the least likely to be a waste of time and money to try.
The real issue is politics - grids are absolutely going to be required for all the folks who can't generate enough solar on their own roofs, industry, cities, restauraunts, etc. Plus how else are you going to make use of wind, grid scale utility solar installations, etc. I have a feeling many countries in the world (especially china) will not have much trouble forcing the grid to do what's needed and subsidizing shared infrastructure with taxes as a shared societal good. If we insist on not doing that though, the grid system as is is not going to be able to financially and logistically figure out this transition, which is probably a competitive disadvantage for us long term if our own energy grid is stopping us from competing on energy because of the way it's structured.