A single character code change
blog.pitest.org
blog.pitest.org
Here's what he changed:
> Set<Thing> aSet = new HashSet<Thing>();
to:
> Set<Thing> aSet = new HashSet<Thing>(0);
His explanation:
"Most of these sets were empty for the entirety of their life, but each one was consuming enough memory to hold the default number of entries. Huge numbers of these empty sets were created and together they consumed over half a gigabyte of memory. Over a third of our available heap."
I'm not sure he ever got around to explaining how it saved "half a million dollars", though.
I love the word bloviating for times like this. It's not really an onomatopoeia but something like it. It's also somehow clearly an insult.
Edit: jrockway has actually changed my perspective on recipe thing...if people want to put a six page backstory and photos on top of a recipe that's completely their right and I would never suggest that more quality (?) content is a bad thing. My problem, it seems, is with Google putting that as the top result when I'm searching for a recipe. Esp when I'm on my phone in a store getting spammed by some captive portal while trying to figure out if the damn cookies I'm making need baking soda or baking powder.
There might be other reasons for this on the net. And I wouldn't be surprised if new recipe bloggers write narrative because "that's how it's always been done."
[0]https://info.legalzoom.com/recipe-copyright-laws-20049.html
I'd argue it's the authors trying to differentiate (and add value to ) their copy of the recipe vs the 100s of others. And I suspect some do believe that adding dialogue or slightly changing the recipe covers them from copyright claims (not realizing recipes aren't covered) or provide them with some copyright protection of their own.
Others probably do it because that's the way it's always been done, but don't realize what they're trying to emulate...
It's the Martha Stewart / Alton Brown recipe for success, where the details are more interesting than the underlying recipe. But people mistake the extra detail for sentimental fluff (Martha Stuart) or technical babble that they don't understand (Alton Brown), and so they just perceive the formula as recipe plus fluff, without seeing the value that fluff should add.
Or they misjudge the value that they think they are adding with their own fluff.
Also probably relevant: recipes (along with product pages) are a couple of categories that Google treats drastically different than general web pages. And as a result has had unintended consequences for the worse, as far as overall quality of the web goes. I'm looking at you, affiliate links.
This. There are some blogs that really succeeds with the story-telling. I don’t mind the stories on Food Lab for example. But if it’s just an incoherent and fluffy word salad, then I’d rather just read the recipe straight away.
The experience narratives can certainly be constructive as well, like maybe describing different versions of dishes tried at various restaurants when trying to recreate a recipe. But even most of these aren't helpful, depending on the author.
And of course the worst is the people who just try to describe an appropriate weather condition or family feeling that goes along with a meal, ala the Giada. Of course Giada is proof that even some people can make that model work well. It's just that most people on the internet don't.
Some experts on Mashable agree [0]:
> SEO and marketing experts agree that Nelson's approach is a smart one, especially in such a saturated landscape. "Because a recipe usually consists of a listing of ingredients and steps, it’s often very difficult for a search engine to discern what this article is trying to convey,"
[0]: https://mashable.com/article/why-are-there-long-stories-on-f...
That makes sense to me. More keywords related to search topics means more hits.
>Nelson thinks there's an element of sexism to the critiques she sees about recipe writing.
Give me a break.
> "The feeling seems to be that they don't think these writers have something of value to offer," Nelson said.
I mean yes, most people do not care about your home life.
> By sharing stories on blogs, people get to know the types of foods [and] flavors that specific recipe creators enjoy. You figure out who is a good match for your own palate."
Do this after the recipe? Oh wait, that would mean I don't have to scroll past more advertisements... This is what everyone is complaining about. Getting to the actual recipe is quite annoying, especially on phones without ad blockers.
I realize sometimes the articles are pure fluff, but if you are passionate about what you are making, you want to read the how and why, the trial and error, and maybe even the history and backstory of the dish.
1) like http://bressanini-lescienze.blogautore.espresso.repubblica.i...
Followed by a story which contradicts the headline, revealing the whole thing to just be click bait.
> This Chrome browser extension helps cut through to the chase when browsing food blogs. It is born out of my frustration in having to scroll through a prolix life story before getting to the recipe card that I really want to check out.
WTF happened? Seriously? It wasn’t long ago things like Google and google maps were actually helpful. And establishments had websites with menus and opening hours.
But this is HN, so nobody knows what that is here. :)
Wait. No. That can't be it. It's about two critiques in drive thru greetings.
Wait. No. That seems irrelevant as well. It's a Yuletide thing? Answering the phone?
Wait. It's not easily googled. It's an obscure term of art apparently.
How would you know something is easily googled if you've never heard the term before?
BTW I did enjoy reading the article I found on "two-part greeting." [1]
[1] https://www.qsrmagazine.com/ordering/100-ways-improve-your-d...
I didn’t agree, and it didn’t fit my own writing style. But they insisted I do it, or they’d edit it in themselves.
Same with this post. The ‘context’ wasn’t particularly enlightening. I thought I’d be reading a programming war story but there wasn’t any pay off.
But I've subscribed to some food blogs because the recipes are decent and the stories make it interesting - otherwise, a series of recipes in my rss would just be incredibly boring. In this case it's closer to a tv personality you'd watch on tv than a recipe book.
I love Bon Appetit's youtube channel for example, and (ironically?) usually tune out at the end of their video when they get to the recipe at the end. Great point, thank you.
People love to ramble, but they also love being able to print a recipe to use in the kitchen.
It seems plausible to me that halting the re-architecting project would produce the savings claimed.
Maybe the better message is "try to understand the causes of your problem before you try to solve it."
I think this is more
Obvious and simple optimizations are beneficial.
I feel in the case of this optimization it also improves the code by making it more clear what is wanted.
I think that is likely often the case with simple optimizations.
To add to the summary:
"This particular trick also no longer works. As I found out while writing this article, modern JDKs do not allocate any space in a HashSet until something is added."
So.. The article might have been better organized (and edited to 1/4 length) and focused on the section, "Four one liners every Java programmer should know". Unfortunately, it'd still be a click-baity title.
Half a million seems like a good guess for how much the company would have lost paying developers and architects to re-design and re-implement the system, buy new hardware, new licenses for their third party software and the revenue they would miss out on by doing all that unnecessary work instead of building a profitable new feature.
The most useful writing, for both beginner and expert audiences, puts the most important content first, with supporting details later. Dispersing the lede late in the article serves the author/advertisers, not the audience, and is independent of whether the content contains context for beginners or a narrative style.
If anything I’d bet he underestimated the labor and opportunity costs.
The cynic in me wants to blame crippling licensing fees. Have you ever scaled up Oracle? Just typing that question gave me a nervous twitch.
But if you actually read the whole thing, it's clearly a meditation on the classic tradeoff of readability vs performance optimization, when micro optimizations are and aren't worth it, and how we should be deciding when to focus on one aspect or another. The single character change thing was just one particularly salient example that he uses to drive a point home.
Reading it without a preconceived idea and letting the article itself explain its purpose to me, I found it pretty well written. There's been plenty written about this tradeoff over the years and this was far from the worst I've seen.
/sigh
If somebody wants to know how to resize an image in Photoshop, you really just want to know which option in which menu and which default settings to use, followed by short (skippable) discussions of sampling algorithms and how resolution interacts with image size. YouTube discourages this kind of efficient information delivery, as Google's incentives drive it to encourage sustained attention. It's hard to build that kind of attention on "press this button and then that one." This is analogous to a laboratory session where you really want to learn by doing, and just need the recipe.
However, learning to be a good photographer is not an algorithmic task—though it may involve them. Being good at something complex involves a comprehensive view and a value system to understand what's important and how many factors interplay. This is calls for a soft version of the academic lecture—something YouTube can do very well.
YouTube can be fantastic at the latter: just avoid it for the former.
Short videos with specific titles, high like/dislike ratios that don't start with ultra fake enthusiasm. The back button is your friend for this ;)
Fucking hell, so much SEO-attracting bullshit instead of just 2 bytes (N and O). It even blahs blahs that "you are here because you want to know if the movie has post-credits scenes". And the next paragraph repeats the fucking question!
https://docs.oracle.com/en/java/javase/13/docs/api/java.base...
Probably the classic software vanity formula:
profit = efficiencyIt involved some expensive re-architecturing.
It required some very expensive new servers.
It required a lot of even more expensive new licences for the third party components.
It required delaying the features the business wanted until all this work was done.
Re-architecting the whole system, buying new licenses for third-party components, buying new servers (all explicitly named in the fine article) can easily cost half a million.
I learned not to read/watch things that don’t get to the point
Often times they never do!
The big takeaway for me here is to learn how your target language allocates memory for new objects.
Me either. TL;DR. I did skim enough to get to the one-character change, and I instantly knew what the issue was: they were using Java. Ok, ok, just kidding, yes, pre-allocating what you're not going to need is premature optimization -- don't do it, except when it pays off.
In this story, instead of doing meetings, why did no one fire up the profiler and actually try to figure out why the heap size was so big? If the senior architect's first thought is to come up with an elaborate proposal even before they understand the memory and processing budgets of their program, what are they a senior architect for? It's like trying to shave off micro seconds of processing time while your network call takes milliseconds.
Good style guides strike the balance where common caveats and easy to miss gotchas are avoided by making rules around them (and then it's no longer premature optimization because someone already did the math and decided on the optimization, so each new programmer using it is not optimizing but merely following a rule in the style guide), while leaving the complex cases for the programmer to optimize as and when needed.
Many people who scoff at opinionated style guides forget that in a big company you don't want 1000 engineers to spend time independently figuring out whether X needs to be optimized [1]. If someone can do the profiling, figure out what should be done in 99% of cases, make a style rule out of it explaining the rationale, it saves everyone the hassle. I believe Google does exactly this.
[1] Apart from the waste of time, not every engineer might be capable of actually figuring it out. Many of the intricacies of C++ for example aren't immediately obvious even if you have spent years with the language. It's more optimal to hire a few C++ super-experts, have them do the profiling, and come up with style guidance.
I think we've missed a similar lampoon of software developers about measuring things with computers. Lack of basic analysis skills has been a running lament for about as long as I can remember.
Some of it may be learned. We seem to fare better when we have accidents that require the additional expenditure of resources than when we point them out a priori. Letting things be surprises is a bad habit but one that often enough pays dividends.
I had sort of the opposite with a Java 1.2 code base. Some clever person has noticed that a bunch of vectors often had only 4-7 entries so they overrode the default (10) to 7. This is unfortunately greater than n/2 which will become important momentarily.
In Java collections tend to double in size on space exhaustion. Using a multiple instead of addition amortizes the insertion time to O(1). So if you start at 7, you go to 14 next.
As I’m sure you can guess, the number of entries crept up to 8 over time. Switching back to the defaults reduced the wasted space from 6 to 2 slots, and dropped our memory usage by more than 10%.
V8 allocates a whopping 64 bytes per empty object.
{}
But if you initialize it with a single property, the size shrinks to 40 bytes. [1] {v8Hack: null}
I too was creating lots of object that would stay empty, but V8 (incorrectly) assumed I was going to change them later so it reserved some extra space. [2]Unfortunately, there is no size hint like for Java's HashMap; the closest hint you can do is create your objects with dummy properties.
[1] https://stackoverflow.com/questions/59044239/why-do-empty-ob...
[2] https://www.mattzeunert.com/2017/03/29/v8-object-size.html
In particular, you should always aim to profile your app when running under a production load so that you do not have to make assumptions about its behaviour. Something like async-profiler[0] is good, since it avoids the safepoint bias issue and can also track heap allocations.
So it really is possible.
As a DBA, I routinely save that (and more) by doing query performance tuning. I aim for 500x-1,000x faster on physical hardware, and 10,000x in the cloud, thanks to EBS/pSSD latency.
And not using Hadoop.
But even worse, the "fix" was what I would consider a spaghetti hack. They create a HashSet that is mostly never used? The real fix is to create the HashSet when you know you need it and don't create it if you don't need it. It sounds like the underlying code has expectations that the HashSet exists and is valid, which in and of itself is bad code.
Check to see if it's null, and if it is null and you need it to not be null, then allocate it then, with size 0 or default capacity, whichever makes more sense.
How is that bad code? No NPE and other null-related issues, cleaner API, easier to test.
If this is a proof of anything, it’s that it can be optimized; not that it’s bad code.
One of the better cases for using appropriate comments...
There are some good points, but, like so many of these topics, it enters the world of "orthodoxy," where we all have to do things the same way, everywhere, and in all places.
Optimization (effective optimization, anyway) is a fairly intense process. For one thing, common sense doesn't really apply. What seems to be an optimization can sometimes do exactly the opposite, like breaking cache, or introducing thread contention.
It can also introduce really weird bugs, that are hard to track down, and strange code structures that relatively inexperienced maintenance programmers may have difficulty grokking.
It's all about the metrics. Profilers are your friend.
Also, in many cases, optimization isn't necessary at all. If the code controls response to a tab selection, then the code that redraws the tab is likely to be executed in a separate thread, anyway, and done at the pleasure of the OS. Why bother optimizing the lookup index to save a few microseconds?
Another matter, entirely, when we are iterating an Array that is many thousands of elements long. In some cases, using HOF can actually decrease the performance (but YMMV). I have sometimes had to replace a map() with a for.
We measure and find hot spots, and then concentrate on them.
If we are working with mobile or embedded, then we also need to worry about power consumption. That's a fairly fraught area, right there, and optimization can actually cause power drain.
In my experience, profilers are not used nearly enough. They also tend to be used when systems are at or near crisis already.
It's way more valuable to have profilers run at on every build (or at least every deployment) to see if there's regressions. Having performance metrics in production to continually track key metrics is critical. All software is built with assumptions that impact performance and all of a sudden those can be invalidated. There's a lot of O(N^2) code that's just fine right up until it explodes.
Some examples I've run across:
- A query API that could contain metadata saying "throw away this result and re-query with these values instead" that worked great for years because it only impacted a single-digit percentage of queries. One day the users decided to really start using this feature, and found that the impact of executing a query that took 100ms and running another 100ms query was way worse than doubling the average.
- A system where you could configure plugins to make "pretty" URLs by running a list of regular expressions. Again, this worked great for years until both the number of URLs and the number of plugins increased and the 25% of the page generation time was spent running regular expressions.
That might be some record both for inefficient coding and then for not caring much about inefficient coding and all the extra expense millions of WordPress users face to host such inefficient code? Either that or maybe few WordPress users make many edits to large pages?
I don't know if there are any canonical writeups on the various string implementations and their performance implications wrt modern hardware.
As much as developer time is more expensive than machine time it seems people are forgetting best practices and/or minor improvements that can have a big impact.
The form:
Boolean b = Boolean.valueOf(true);
Is slightly more efficient than the subjectively uglier:
Boolean b = new Boolean(true);
So couple of notes on that. The "new" operator is cheap, but still involves an allocation. Oddly enough, one of the strangest bugs I found back in the day was: Boolean a = new Boolean(true);
Boolean b = new Boolean(true);
System.out.println(a == b);
false
But with objects and wrapper types, the .equals() method is the correct way to compare two objects.This is another reason to avoid boxed primitive whenever possible.
Boxed primitives are useful when working with systems and languages where "null" is a concept. Take for example, interacting with a database where you have a boolean column that has true/false/null. That translates well to boxed primitive types.
Leads me to believe that there's lots more meaningful optimizations remaining that someone who better understands the system could implement.
Ah, but the object still exists. If large space of objects is needed that is sparsely used, the thing to do is to make that space itself lazy, not the individual objects.
https://hg.openjdk.java.net/jdk/jdk14/file/f77e9e27b68d/src/...
There's no hard data to show that a smaller value would give better performance for typical Java workloads. Resizing operations are quite expensive.
See this comment, and also the implementation of putVal()/resize(): https://hg.openjdk.java.net/jdk/jdk14/file/f77e9e27b68d/src/...
Perhaps things were different back in Java 1.3.
EDIT: Apologies, I completely missed this in the article, where it's explicitly called out.
This particular trick also no longer works. As I found out while writing this article, modern JDKs do not allocate any space in a HashSet until something is added. If I had kept applying this trick pre-emptively I would have been wasting my time.
I mean, you have conditionals in the insert instead of unconditional allocation at the start, so maybe a little worse for things like branch predictors, but ...