162 karma · joined July 25, 2016
Two of the three strip titles are hallucinated and two of the three strips are bad examples. Haley is mute in strip 403 and does nothing. Strip 578 is the start of the arc that shows the behavior Gemini is talking about, but has things going wrong so it's not a good example either.
Claude picks a good strip but also hallucinates the strip title: https://claude.ai/share/56be379d-c3da-443e-b60f-2d33c374eba8
Though I still prefer Claude for this since it's better at citing sources.
So if a model can solve every question but takes 10x as many steps as the second best human it will get a score of 1%.
However it's possible that consumers without a sufficiently tiered plan aren't getting optimal performance, or that the benchmark is overfit and the results won't generalize well to the real tasks you're trying to do.
There are plenty of other signs this story is likely fake. The author claiming to be posting from a library on New Years' Day (most government buildings are closed) and was responding over 10 hours on the account. He's using a throwaway and a "burner laptop" at the library, but he also says he put his two weeks' notice in yesterday (also odd that this is on New Years' Eve) which would make identifying him trivial.
Fake stories get to the front page of Reddit every day, I wish journalists were pointing out the actual signs not to trust something to act as a better example.
It is perfectly valid to say "the best case is O(n)" which means the best case scales no worse than linearly. The "no worse" here is not describing the best case, but rather the strictness of your bound. You could say a sort algorithm is O(n!) since, yes, it does scale no worse than n! but it's not particularly helpful information.
Big O notation is used imprecisely frequently (see also people saying a dataset is O(millions) to describe scale).
AI is not a precise term. If you can make a product that feels intelligent to the user, why does the implementation matter?
I think both games have their interesting parts and it's fun to play both.
Early in the game the expected value of a trip around the board is positive because of low rents of undeveloped properties and the net gain from Chance and Community Chest, plus passing Go. Cash flow into the game is positive as a whole.
Later in the game, when property is developed you're going to be paying bigger rents. This is going to cause players to mortgage or sell houses to pay the big rents (also their own liquid cash supply is likely lower from developing their own properties), each of which is a money sink since you don't get full value for these. This causes money to exit the game, eventually leading toward bankruptcy.
However, this requires that players actually get to develop their properties. If the game hits a lock due to no one trading and no one getting a natural monopoly, the game can go pretty much indefinitely.
A good example of training/serving skew.
This study was a 100 person survey with no experimental evidence to back up people's beliefs. I've been working with educators for a while, and they've said that "learning styles" are just a myth with no solid evidence. All this survey says is that people think they learn best via one method, but people are notoriously bad at knowing what is an effective way of learning. In surveys people overvalue the "intensity" of a learning experience [1], which makes them think things like a 3-week bootcamp are more effective, when really spaced repetition is a better use of time.
At least in the US, there are many easier ways we could improve our education methods before resorting to ML-driven customization.
Alsup was suggesting Waymo to drop the patent portion of their suit and focus solely on the trade secret portion of their case. The trade secrets were pared down as mentioned in that snippet, but that's in addition to dropping 3 of the 4 patent infringement claims.
Minus-Celsius is a former tournament Monopoly player and his posts on this thread summarize it better than I could: https://www.reddit.com/r/boardgames/comments/42wze4/how_to_w...
Three years into my job, I'm mostly over it. I still have to deal with coworkers laughing that I use Windows on my personal computer, but at least I can eat the cake I made this weekend.