1,656 karma · joined September 8, 2018
EDIT: To add, if you get it slightly wrong... you're fine, it'll be adjusted for in next year's taxes.
[0] ...because it has to, in order to investigate evasion.
It's the same for Magnus Carlsen. Even with Queen odds, Stockfish is literally unbeatable for the best players in the world. It's just too strong at evaluating all kinds of random tangent moves (and ensuing positional advantage) which no human player can possibly pay attention due to the time required.
Stockfish vs any human is like Carlsen vs other players by about 3-5 orders of magnitude[0]. It's that stark.
[0] A wild pun appears.
EDIT: To avoid having to respond to each responder, fair comments about Queen odds. Maybe I was thinking Rook odds? Also, I kinda lumped Stockfish in with all the other engines, but I realize there are other engines with different properties ofc.
There was a renaissance during Covid and due to 'The Queen's Gambit' where it gained much more mainstream popularity, but... Chess AI was already far far (like 1000+ Elo) ahead of human players at that point.
The thing is... chess is humans playing (communicating) with humans and that's what keeps it interesting. Check out the view counts of chess AI tourneys vs. human tourneys.
We're headed for even rougher times if LLMs are going to be taken "at their word".
[0] Now, watching her testimony, it seems she was willfully complicit and it's a disgrace that she isn't in prison, but that's tangential. The "system" was responsible, but without personal accountability, nothing will change. So it goes.
Gotta tie some actual prestige/shame to not following up.
(As someone pointed out in the comments, the title of this specific post could have been less heart-attack-inducing, but still... :) )
I seem to remember there was a "you failed 5 PIN entries in a row, please wait 500000 seconds before you retry" on Apple phones. So, you probably also want a sensible max... which makes exponential a bit pointless. Just do a basic fixed delay + (large, e.g. 0.5 x the delay) jitter and you'll be fine for most things. You can add a bit of cumulative delay if it's really costly to do retries.
I remember talking to a business type who was so incredibly proud that GPT could turn his incredibly vapid thoughts into bland pointlessness on LinkedIn... not realizing that everyone in his feed was doing the same thing. We never did business with him for ... reasons, obviously :)
It really applies to anything with a huge number of options, I remember similar things from back when I used Gentoo and the wild suggestions I'd get for what exact CC/C++ compiler options to use...
So, ideally you'd want each separate test case to be compiled separately, but even then you wouldn't be safe! ... because any UB in that test (or the code it's testing!) could lead to a random pass.
UB is good in some ways, but other ways it's really really bad.
EDIT: I will say: If you have a UBSAN turned on for testing, etc. you're reasonably safe... but not fully. There's a lot of stuff they don't catch because it's essentially impossible.
Every property-based testing system (invented ca. 1980) will explore boundary values. The semantics (or lack thereof) of C and C++ can make this difficult to actually test for because the compiler is allowed to say "test passed" to any input leading to UB.
EDIT: Don't get me wrong, my European nation did a lot of bad shit, but don't put the idea of an "expat" on us.
If doesn't matter what 'p' is in their example. The point is: if 'f' is undefined behavior (rather than just impl-defined), then the optimizer concludes that the "if p { f() }" can never happen... which means that we're allowed to assume that 'if p { ... } else { ... }' (in the first part of the example) will always take the else branch. The compiler will optimize accordingly and just always call g() unconditionally.
[0] I still use Arch, btw. :)
[0] Arch Linux, btw, because that must be mentioned.
I can generate a lot of tests amounting to assert(true). Yeah, LLM generated tests aren't quite that simplistic, but are you checking that all the tests actually make sense and test anything useful? If no, those tests are useless. If yes, I don't actually believe you.
It's the typical 10 line diff getting scrutinized to death, 1000 line diff: Instant LGTM.
Pay attention to YOUR OWN incentives.
That's not my experience... mostly it's about first interrogating the actual problem with the customer and conditions under which it occurs. Maybe we even have appropriate logging in our production application? We usually do, because you know, we usually need to debug things that have already happened.
(If it's new/unreleased code, sure fine, let's find a debugger.)