481 karma · joined June 22, 2010
They are expensive to run, in terms of GPU cycles, but they are noticeably better than the previous models.
It's also hard to constrain them well. If you want 95% accuracy, it takes some tuning work. If you also want to avoid 1% total batshit nonsense (repeat "chicken" 50 times), then you have to check for that. Earlier models were sometimes wrong, but they were not quite so aggressively wrong as the 1% case of LLMs.
That's just my anecdotal experience, but it leaves me both optimistic about applications in the right spaces and worried that people are just shipping something that's OK 75% of the time and calling it a product.
Of course, that worked because 1) I was really only doing one project, not juggling multiple ones, 2) there weren't all that many dependencies (Numeric, plotting, etc.), and 3) I was already up to my eyeballs in the build system with SWIG and linking to the actual compute code, so I knew my way around the system.
But every now and then I just shake my fist at the clouds then mutter darkly about just installing the dang thing and maybe not taking on so many dependencies. :-)
That being said, those are very much edge cases.
More damning, from my POV, is that you can't get ref counting of things like C objects or file handles, since they're just string handles. But there are a lot of uses that don't need that
For example, I couldn't find docs on its display stack, so when I wanted to display a graph but with images instead of text for the nodes, I got stuck. Sure, I could dive into the code and eventually get something that worked, but I wouldn't be sure I was coding to the API or to the implementation. It crossed over to "playing around with this tool" to "have to do some real work to understand", so I ended up dropping it.
I mean, today I spent time:
- Digging into the analysts dashboards for some time series that seemed off. Generated a plot from my own metrics stash, filed a bug w/ the analysts about the differences. Replied to some doc comments saying that I'd done so.
- Code review of a few changes.
- Write a quick jupyter notebook (well, colab) analyzing a different problem. Basically reading the data for an example into the notebook, writing a bit of code to visualize it, fiddle until I had some sense of what the problem was.
- Decide I should regenerate a dataset with some different processing. Spent maybe an hour doing the code changes and getting them reviewed. Then build the binary, run it to generate a dataset, then launch a few-hour processing job to see if it works better.
- Write some notes on how I should decide if the new version is better than the old version.
- Fire off email to a few people. Asked one person about existing viz tools, updated another about my earlier colab experiments and asking if they know of anyone who's done a similar analysis.
Out of all this, I think pair programming would have maybe been tolerable for a bit in the middle when I was actually making code changes, but everything else was either based on huge amounts of internal context or was completely ephemeral analysis.
I think I must be doing a different kind of thing. :-)
Permanent DST makes no sense to me. Maybe it's my astronomy background, but "noon" means something, something that involves the position of the sun and the earth. We quantize that to timezones for coordination, but it doesn't mean it's meaningless.
If we stop switching, fine, but don't mess with noon. Just change your schedule to 8-4 or whatever. Permanent DST seems like wanting everyone to be above average. Or deciding that everyone would be happier if they're taller, so we're shrinking the foot by 10%.
And then some summer jobs writing dBase III and FoxPro, some actual education, and here I am.
Those books absolutely set the tone and sparked that first interest.
As someone who's worked in science and finance (modeling, not accounting), floats work just fine, thank you very much. The modeling/accounting split in finance is a legit point of confusion, though.
It's a very hit-or-miss industry, so there will be periods when you're not doing anything, and then you're left with this pool of extremely high-talent people doing small productions, waiting tables, and the like.
I've know that the first thing dice rollers seem to grow is 3D simulated dice, so clearly the market is there.
But this would be (for me) replacing an assortment of wood blocks, lego minifigs, scribbled-on index cards, and assorted tokens with a whole lot of additional prep time.
But in my experience the PM is technical enough to understand what's going on. (Can write some SQL to answer questions, possibly ex-engineer themselves, etc.) They're in the same meetings, same email threads, looking at the same set of OKRs, etc. It's part of the engineers' jobs (B/F/D whatever) to communicate their constraints and their ideas (both product and pure-tech) to the PM, and it's part of the PM's job to take those into account when advocating for what should be done.
Similarly, the more the engineers know each others' specialties, the better they can coordinate. It's probably more important for everyone to have "a little bit of product" in them, but that doesn't mean we don't need a product-specialist.
When it's time for quarterly planning, the PM's voice is definitely loud, but they're still just one voice in the room. They're the one accountable for the product, which gives them some leverage, but the other voices are there (TLs, managers, etc.)
Now I can see this going terribly wrong if the scale is off (only one PM for too many engineers), or if communication breaks down (PM scribbles a "design" on a napkin and faxes it over), or if only the PM is consulted for planning. But the problem there is that communication broke down, not that it's bad to have a PM.
This comes out in #13, #19, and #20 (well depending on exactly what they mean by "managing".) I find it really helps to have someone dedicated to the product side, so they can help with prioritization, generating ideas for what to try next, keeping track of the research, and tracking the launch process and metrics.
We could, I suppose, distribute all of that to the engineers, but there's value in that specialization.
(I generalize to "agile-folks" here because I feel like I've seen this before, but I can't think of exactly where.)
Their site and app are glacially slow, though. Not sure what's going on there. And for a while they've had some text-entry glitch that reverses text every now and then. But that's just griping.
Also, with kids it's nice having the weekly 3-4 gallons of milk just show up rather than schlepping them home from the bowels of Whole Foods.
First of all, the quote from Jack Menzel about figuring out home and work locations sounds 1) about typical for him and 2) exactly correct. If Google is giving you directions and popping up traffic along your route between home and work every day, it's pretty clear what's going on. There's nothing weird or unexpected going on there.
Second, the settings thing. I see this as two issues: one, multiple settings, and two, making the settings hard to find. For the first, both search and maps had their own settings (web whatever vs location/location history), and they didn't talk to each other. I'm sure the relevant VPs talked about relevant VP-things, which probably did not include the config page. Both were sure they had an option to turn location off (for their project), so that box was checked, and done. Yes, someone should have made sure there was a single button, not three, but the org chart was shipped instead.
The hard-to-find thing is harder. I, as a regular schlub engineer, think that sounds pretty sleazy, but I have no idea how true it is. If someone's A/B test said, oh user engagement was down on this arm, I can see that happening. It'd be a failure, but I can see that happening.
I guess at my level, it seems like all the people I'm working with take user privacy really seriously. If someone wants their data deleted, we go through a lot of hoops to make sure it's really gone. Any feature using user-data gets a privacy review and usually ends up requiring pretty strict differential privacy bounds.
This is a little unfair and possibly just ignorant, but my impression is that Google is far better at protecting user location info than the telecoms, who have more complete data from cell-tower triangulation and who are generally willing to sell that data to whoever, and yet they get a lot less attention for it.
My daily driver is my 16" MBP, and while I'm thinking of getting a plain-vanilla 10.2" iPad, I can't think of any cases where I'd want a bigger one.
The code was check-all-pairs, e.g.
for (int i = 0; i < container.size() - 1; ++i) {
for (int j = i + 1; j < container.size(); ++j) {
stuff(container[i], container[j]);
}
}
Which worked just fine for int size, but failed spectacularly for size_t size when size==0.I totally should have caught that one, but I just couldn't see it until someone else pointed it out. And then it was obvious, like many bugs.
All the treatments I've seen jump in to manipulation without really going in to the axioms used. That paper seems to do a lot better at fitting it together, so I'll certainly read it, thanks.
My assertion, that R's NaN is not "non-standard", seems upheld by the article. It's a quiet NaN with a payload, which is well-defined by the IEEE 754 standard.
As other posters pointed out, it's relying on a "should" behavior from the spec, which is risky but common. It sounds like disabling the "RunFast" mode cleared up their issues, which seems quite far from it being an "obscenely bad" design decision.
It's not terribly unusual to require IEEE 754 compliance in numerical code, like the usual options for avoiding --ffast-math -style stuff.
(IIUC, that is. It may be something like a "should" not a "must".)
Many menus are also available in PDF form, so we're trying to figure out if it's worth bothering with the PDF itself, or if we should just render to image and thus reduce the problem to the menu-photo one.
Does anyone know of any attempt at this?
Blah blah blah transformer something BERT handwave handwave. I should ask the research folks. :-)
It seems a bit overly-aggressive to accuse them of lying because some pop-sci journalist heard tensorflow and went straight to "artificial intelligence." And even there, it's an understandable and common hype.
It doesn't take away from the point that they're doing some really interesting and novel computational physics to make this thing work.