Perhaps the US are trying not even to acknowledge it. A reputable newspaper would normally include "the Department of War was contacted for comment" but maybe that's not appropriate for Reuters.
11,329 karma · joined January 2, 2014
Perhaps the US are trying not even to acknowledge it. A reputable newspaper would normally include "the Department of War was contacted for comment" but maybe that's not appropriate for Reuters.
I can't find a better source for just how fast the Prague escalators were, but as an occasional user they seemed more than twice as fast as London or Paris which TFA gives as 65-75 cm/s.
They felt dangerous to step onto, but I never saw anyone actually have an accident. The metro in Prague is also deeper than most.
[0] https://www.expats.cz/czech-news/article/prague-s-speedy-len...
Run the same protocol again, but have the agents think they had limited resources or that HuggingFace was rate limiting them, and they'd find something you'd consider smarter.
Computers don't have a sense of elegance by default. Elegance emerges from constraints.
Storing it in underground salt caverns is more secure, not necessarily cheaper.
If he pairs the edited human-readable part with a real barcode copied from a real license in someone else's name, then anyone inspecting the license will see the documentation matches his claim, and if they also use this site to check for fake barcodes it will confirm the barcode was really issued by the California DMV.
> the ZNB field is not empty and not garbage: it contains a well-formed 71-byte DER ECDSA signature, correctly Ascii85-encoded, with the right prefix and a plausible length. But it fails the cryptographic check instantly, because it was signed with somebody else's key.
Seems doubtful! I expect the forgers used a real signature from another card instead, so it has the right key but the wrong data. Reverse engineering the process as the author did and making up their own key wouldn't be of any value to the forgers.
> I built a little demo to check the signatures across California, New York, and Virginia: take a picture of the barcode and check it here.
This is not wrong, but should come with a little warning. A real verifier needs to additionally check the encoded data matches the human-readable data on the front of the card.
If it's 98%, the average household loses 0.2% of their income gambling.
Probably I already ask them to spend 2+ days attending HR or compliance training or listening to senior management tell them about sales targets.
But the point is, that 1% extra productivity requires the sometimes staggering cost of making the software 10x or 100x more reliable.
Three nines reliability is great for most purposes. 8 hours downtime a year.
If your system produces money at a constant rate, it captures 99.9% of the available money. Even two nines or one nine might be pretty good on that basis, when the alternative is spending 2x or 10x as much - let's build another unreliable system with that money that captures some other independent market opportunity.
Poor reliability is a problem where you need to chain many systems together, or where the cost of a single failure is very large compared to a success. Or - as happens commonly because of load - if your periods of unreliability are correlated with periods of maximum opportunity, like an e-commerce site failing on Black Friday or a trading system failing when the market is most busy. But if you don't have one of those cases, evaluate whether investing in reliability is actually worth it to you.
GitHub is an example where two nines of reliability ought to be OK. The argument against it is that it's bad marketing to have an unreliable service, especially one aimed at software engineers. And if GitHub is largely a marketing play by Microsoft anyway (do they really make back its cost in enterprise subscriptions?) then marketing considerations need to drive its reliability.
Everything is locked behind multiple layers of permissions. It takes weeks to get the correct permissions set up even for my actual job (which I see any time changing roles here, or when onboarding new hires). Getting the permissions to jump in and improve some other part of the system that I don't officially own - even though I might get granted them if I ask nicely, because nobody knows who is actually meant to have what permissions - is so much higher friction than asking the "right" person.
I don't believe this.
You refer to "subagents", so this is not just an LLM but an LLM with some kind of agentic harness. Any reasonable harness and prompt, given internet access and appropriately prompted to succeed on this task, is more than capable of firing up Lichess or chess.com and relaying moves back to you. The free levels will be enough to beat you.
A frontier model can also likely one shot a chess engine that plays at your level, again if given an environment in which it can do that.
I completely believe the LLM on its own can't play a full game of chess at your level. Though I'd bet that with enough reinforcement learning it is possible to train a pure transformer architecture to do that. We just don't do it because there are other approaches that play chess much better.
The region that was once East Prussia has now mostly been ruled by Russia since longer than almost anyone there has been alive.
> About a third of all non-big-box/amazon online shopping runs through Woocommerce installations
And as for this:
> 2.84M vs 4.34M is quite a wide margin, wouldn't you agree?
No, I wouldn't. I think that's within the margin of error for the typical methodology of this kind of research. They're the same number.
Source? This source [0] puts Shopify 10x ahead of WooCommerce on GMV. They're close on number of installations. That roughly matches my experience working in this space in 2022, though it's definitely possible our customers weren't representative of the market as a whole.
[0] https://www.sanjaydey.com/shopify-vs-woocommerce-why-brands-...
One shot generating a problem of exactly the difficulty the user has in mind is a very difficult problem for anyone. "IT security students" span a wide range of capabilities, but I would expect most of them are worse than qwen3.8-2.7b at this kind of work.
Changes that are needed to get the project to run right in the dev environment, but that should never be committed, are a smell: I move them to config that doesn't get committed.
Increasingly with agents I'm likely working on multiple features in parallel (I haven't adopted git worktrees but I probably should). Even then I try to lay out code so concurrent changes touch disjoint parts of the codebase, so I can do "git add Widgets/FooWidget/" and the ten files modified under there will be the right things to commit.
Maybe 30% of the time I find I have done work that belongs in multiple commits and I need to break it up. Even then I often end up doing "git add -p ." to interactively pick the parts I want to add.
I don't think there was anything particularly ballsy, or perhaps even reckless, about agreeing to test the taxiing behaviour. It never crossed his mind that this could lead to a situation that got out of his control. Starting the airplane was a complete accident.
Getting the airplane airborne once it was running out of runway - and then landing it again - did take balls, but is not really the part of the incident that the review board was there to scrutinise.
That's exactly the service the commodities markets provide: they will absorb a small trade without anyone incurring the costs of shipping precious metals, and they know who to call to get it shipped once there's enough supply/demand imbalance to make shipping it worthwhile.
That reasonably covers pretty much all cases, especially in a state where lots of people have guns.
That said, this isn't a security vulnerability, it's just a bug. To meet a reasonable threshold for being a security issue, you need to show a real system that has an issue caused by this, and then the vulnerability is in that system, rather than in Python.
I'll grudgingly allow that a buffer overflow or an SQL injection possibility - in a library advertised as safe against that kind of bug - is a security issue, because there's so much history of turning those into real exploits. But a choice of library or language that makes those bugs easier to write - the idna library, or C or PHP say, is not itself a security issue.