Drop millions of allocations by using a linked list
github.com
github.com
Even the dev with the fix wasn't rude about the original problem, he seemed pretty humble about it actually.
If you think the Ruby guys are such shitty programmers you should be able to dive into their codebases and find the plethora of problems to show them what's up... so either give them a pull request or STFU ;)
The Ruby community is very good interpersonally from my experience. It's a culture that I think comes from this:
We were talking about the overall culture of the rails community. Nor did I intend to imply the DHH was evil, just abrasive.
"nice" != "good"
Most people have neither the time nor inclination to fix problems with the language they use. If I buy a power saw and it turns out to be a POS, I just won't buy that brand again.
"Oh, I made a boneheaded error and now your code is 100x slower than it should be? Tough sh*t, fix it yourself. I owe you nothing!"
That's not a reasonable mentality to have if you want adoption. Obviously the Ruby devs don't feel this way, but you seem to. That sort of attitude is nonsense and short sighted.
No, you really don't. The fact is: they owe you something, but it's OK, since you give it away for free.
>If you don't want that responsibility then don't release your code or get out when things become serious.
Or else?
>That's not a reasonable mentality to have if you want adoption.
What if you DON'T want adoption? Or you want just the smart kids that can contribute back to adopt your code?
>Or else?
Or else people just don't use your software and don't buy into your ideology. If you don't care, fine, but obviously many OSS proponents do.
Actually EFF has little to do with OSS.
Perhaps you meant FSF. In any case, most OSS software that isn't GNU has very little if anything to do with FSF. Some OSS authors, especially those MIT-licence inclined, even hate its guts.
I never said that, but most projects welcome pull requests.
> I owe you nothing!
True, it's the short version of the MIT license.
> That's not a reasonable mentality to have if you want adoption.
Yes, maintainers should always try to take care of their projects and to accommodate the needs of their communities.
> Obviously the Ruby devs don't feel this way, but you seem to.
You are over-interpreting. But I hate seeing more and more OSS contributors getting burnt out and steamrolled by a self-entitled rout.
If you release some script or small program which was useful to you thinking that it may help someone else down the road, I agree, you don't owe your users anything. However, when you position your solution as something people should adopt (e.g. RoR, etc) then you do owe a level of quality and support to your users (not saying RoR doesn't, just a random example.)
I love OSS, but every time it bites a user in the behind it loses mindshare.
Ruby's standard libraries don't have much support for immutability, or deep cloning, or copy on write, or linking concatenation. There's no standard library way to ask for a snapshot of an object.
So in the early days for Ruby, an idiom was: if you're writing a method that takes a list, and you need to be sure your list doesn't change out from under you, then duplicate it, get it working, and if it becomes a bottleneck then optimize it.
Take off your Hats of Superior Coding. Any one of us, regardless of honorific titles, could have made this mistake, and you know it. Being steeped in CS Fundamentals does not immunize you against bugs.
Congratulations to tenderlove for finding the bug. Remember the details - it'll be a great war story in a few years.
I like to remember the past of computer programming as though it was once friendly and receptive to people of all skill levels. I owe quite a lot to the geeks who came before me and answered my stupid questions, gave me powerful tools to learn with, and accepted my contributions; flawed as they were. Without making a few mistakes along the way I wouldn't be where I am today.
I don't know whether I would have given up if the people I met were more smug and mean-spirited but my progress might have been slowed by it. Life is too short to waste bothering with people who are miserly with their good fortunes. It doesn't cost you anything to be nice and share your knowledge and wisdom. It may pay off when the person you're sharing with rises to stand upon your shoulders one day and pay homage to you.
Also, hubris being a virtue is crap. It may be decent for your self esteem but it makes you miserable to be around. Case in point: Larry Wall.
Merriam-Webster - hubris - "a great or foolish amount of pride or confidence"
It's just a synonym used for arrogance and smugness. Humbleness is a virtue.
I make a point to keep track of all technical debt in my projects so that I have an easy way to quickly identify opportunities for improvements when there is spare capacity, and also so that technical debt isn't left unaddressed.
//TODO:
//FIXME:
;)Things that we know need to be addressed, but the priority level is "when we have time" are better served living in the code. They tend to get lost in project management and ticketing systems, whereas they live until the code dies if they are in the code.
//kludge
I think the reason this rubs some people the wrong way is that implementations like this can be the result of the "no premature optimizations!" philosophy and its advocacy. I've encountered this firsthand at a number of organizations, and it apparently was somewhat endemic in the Ruby space.
Make something that works, and benchmark later to find that one magical hog that you can quickly change and then everything is optimal. Only it almost always ends up being a performance death by a thousand (million) cuts, performance and resource malaise so endemic that fixing it almost seems impossible.
In my experience, projects that adopt a slightly broader definition of "bug" have a better track record of improving on those fronts. Nobody like to have bugs around, after all.
I'm growing more impatient with the years. I have measured out my life with slow software. We talk about saving developer time with dynamic languages, but, as Flight of the Conchords sang, the sneakers don't seem to get much cheaper; what are your overheads?
There are very, very, very few people who actually do any work on core infrastructure projects. I don't blame them. The codebase of rubygems is... not exactly welcoming. I myself did some work a while back, got frustrated, and quit. But if you want the real reason, there it is.
Oh, and when you _do_ overcome all these barriers and then actually make an improvement, people will say "lol rubyists learning data structures" instead of saying "thank you for saving a bit of the most precious resource I have every single day, time."
Note that I wasn't only complaining. I just figured that there is some real reason for the slowness and that this reason might be known by the community (and I didn't find much googling for "why is ruby gems slow"). Call it preliminary research...
You'd think that, wouldn't you? The reality is... not that. :/
And yeah, this isn't really about you specifically, sorry if it came off that way. This is just a general problem in the Ruby OSS world. I can't speak to too many other OSS worlds, as the Ruby one is the one I've been most involved in. If you're looking for more research, Andre Arko is one of the other people who's actually doing work in this area, and he's given a number of talks about why Bundler is slow: https://vimeo.com/67807956 and related. He's been making a lot of strides, but it's not easy.
What a very polite way to put it.
IMHO that whole mess (rubygems + bundler) would ideally be replaced from scratch, removing the need for bundler in the process.
If any generous sponsor wants to improve Ruby as a whole, that's where their money should go. Imagine the productivity gains if everyones test-cycle was suddenly >10% faster, and nobody would have to waste energy on bundler/rbenv/rvm issues anymore.
Perhaps we could even fix the deployment nightmare in the process, with e.g. jar-style packaging, but now I'm really dreaming...
While I love a good 'burn the world down re-write,' it's a _lot_ of work and isn't guaranteed to succeed. It's been tried before, and in other languages too: check out wheels in Python.
That said, Rubygems did replace what came before it, and Gemcutter replaced what came before it... so it could be done. It's just non-trivial.
Same here, still have my napkin notes. It's actually one of my bucket list projects, if there wasn't the dreadful food-on-table constraint...
check out wheels in Python.
Valid point. Python is indeed a good example for an even worse situation. In fairness, most languages are still worse off than Ruby even now. I wouldn't trade maven, CPAN, etc. for bundler with all its warts.
However, there's also languages pulling ahead. npm seems to be slowly getting there (after a rough start) and the Go experience (godep), while still in flux and not directly comparable, is also something to draw lessons from.
Rubygems is a very old project and it was basically stalled for a long time, it has already made quite big strides recently.
But I don't think it's rubygems' role to fix what rvm/rbenv fix (multiple rubies). Jars don't do that either :)
Well, it all sticks together. I would say ruby-build should be retained as an external tool to conveniently fetch and install ruby versions.
However, we should very much replace the god awful environment magic that rbenv/rvm perform with native ruby/rubygems support for version/project-scope gemsets.
Rbenv is a well designed crutch - but still a crutch.
In an ideal world you'd checkout a ruby project, point any recent ruby-binary at its Gemfile, and it would download/install not only the required gems, but if necessary also the required Ruby version, as specified by the Gemfile.
It would store everything in './.ruby', which could optionally be backed by a common shared directory (~/.ruby) for space efficiency.
gem: --no-document --verbose
in your ~/.gemrc may speed up some things and show slowest steps (where '--no-document' means the same as deprecated '--no-rdoc --no-ri').Check the difference between "gem install sass" and "gem install sass --no-rdoc --no-ri" and be amazed.
I don't even know how to view gem documentation, and I've never wanted or needed to. --no-document should be the default
They could make it download the docs on first view
Source gems that include C extensions may well have compilation; there is also, IIRC, automatic doc generation in the usual, default gem installation.
> Why can I install Ruby packages using apt almost immediately, when gem/bundle install takes half a coffee break?
Because the apt packages are prebuilt for your OS and architecture, they don't need to resolve platform constraints, possibly build after download, and do doc generation after download.
It seems like this should have been trivially detectable, since the difference is so dramatic.
This reminds me of the gc_disable() PR that reduced composer install times by half a few months back.
--Donald Knuth
http://ruby-doc.org/stdlib-1.9.3/libdoc/rubygems/rdoc/Gem/Sp...
The 'trail' parameter, an array, was implicitly duplicated by applying the '+' operator on each recursion through a dependency.
trail = trail + [self]
That + operator looks so innocent, so seductively simple..It's very easy to pontificate in hindsight when somebody else has done the hard work of actually finding something that can be improved.
This was not a good comment to post to HN. When you toss a Molotov cocktail into an HN thread and "duck", here is what we are going to get:
web programming people aren't familiar with even basic CS
I stand before you as a counterpoint to your foolish generalisation
Oh cute, a web dev. You've just reinvented 1975 [...]
Would you like a pat on the head?
Stop whining
your comment is bullshit
you deserve a slap
For any large group like "Ruby people" or "web programming people", HN has many users who either identify with being part of that group or identify with not being part of it. Given the numbers, there will always be a few who are having a bad day or feeling defensive or what have you, enough to respond angrily when someone posts a slur. Then their counterparts feel attacked, and down the whole thing goes.Social padding prevents disputes like this from turning ugly, but we don't have that on HN. Even when you know another user from their comment history, that's not much information. Imagination fills in the gaps and then we imagine the other in the worst light. That's why the community here is fragile.
To be a good community member, please don't post things that could easily set the thread on fire. If after editing them out, your comment has nothing substantive left, please don't post it.
In Java and Ruby, the standard library uses an object `+` operator to mean concatenation (not numeric addition), and the implementation did immediate data copies, rather than doing reference linking and copy-on-write (or immutability).
Also one might argue that the problem is actually much worse with strings, because string concatenation is so common and the syntactic sugar of the '+' operator for strings encourages the "wrong" way.
As to using the wrong data type, that's really the programmer's fault. If you don't allocate enough capacity or use another data type (e.g. LinkedList) if you don't know the required capacity, you are doing a bad job.
I'm not really trying to knock any particular language or runtime here, the point I was trying to make is that nearly every language I've used has quirks that encourage convenience over optimization, and that just because you're coding in language foo, it doesn't mean you're off the hook when it comes to being intentional about the choice between them.
For example, any linear recursive algorithm that deals with the head of a list and the tail/rest/butHead of a list is optimized for a linked list implementation. So... to understand a linked list, you really need to be familiar with algorithms that bisect list sin this manner, and the reverse: To understand algorithms that bisect lists in this manner, you have to be familiar with a linked list.
So... Yes it’s a data structure, but it’s joined at the hip to the algorithms that operate best on it.
https://github.com/rubygems/rubygems/blob/800f2e63bc6174b5b4...
But really, I should move to web programming. I'll be an expert computer scientist there, probably.
/snark (apologies for being offensive, but good lord, what a silly statement - unless I missed the joke)
I stand before you as a counterpoint to your foolish generalisation, and guess what, I know plenty of other people that don't fit your stereotype either.
I'm not saying there's not an element of truth to your statement. See, here's the thing. People throughout the software (or any) industry have different collections of knowledge. There are an endless number of things to learn and each one of us is different and brings a different set of skills to the table.
To be great at web development you spend years learning the subtleties of developing for a vast domain of different platforms. There are bugs in the platforms decades old that I know the intimate details of and have workarounds for especially constructed to fit in with the other bugs in the other platforms we deal with. And that's a tiny facet of what you need to know.
You know what would really happen when you came over to web development? You'd find that a lot of the skills as a computer scientist aren't altogether useful.
You wouldn't be an expert computer scientist. You'd be a junior developer, probably.
The problem is, I think, this is NOT happening or at least to seldom because it "does not pay off", because "the aws boxes are sooo cheap!"
I'd say: the aws boxes are way too cheap and effectively misused. What should have been a way to scale up apps with reasonable effort turned into an energy and resources burning landfill. Granted, at least it is a shared landfill which really helps with from the resources/energy point of view.
(Juniors) Coders should be able to power the machines by driving a bicycle ergometer ... as a way to make the homo-sapience grasp the effects of their "mental work" ;)
But fundamental CS is a different thing.
"Annnndddd.... what reaction do you expect?"
Admitting your shortcomings and try to fix them is not an option at all?
When some central web frameworks do lamest mistakes like making an O(n^2) queue, it is frightening. And instead of deflecting any critique, one may try to fix this situation somehow.
goto:fail - https://news.ycombinator.com/item?id=7281378
shellshock - https://news.ycombinator.com/item?id=8365110
I'm sure whoever wrote this would feel a little chagrin when coming back to their code and seeing how inefficient it was (though it got the job done when the lists were small), and they probably would admit their shortcomings, why wouldn't they?
Condescending snark is really easy, and it's easy to say in retrospect and with time to reflect that most useful code has flaws and point them out - software is never finished, and there are a lot of different levels of experience and requirements. If rubygems had never become popular, this wouldn't even be an issue.
PS Rubygems isn't a web framework, it's a package management tool, so the straw man you're hacking away at is the wrong one.
I was talking about the last one here.
"PS Rubygems isn't a web framework, it's a package management tool, so the straw man you're hacking away at is the wrong one."
My statement about web frameworks was not about RubyGems, it was about web frameworks, and it was an example of the state of affairs in web development.
The story is about rubygems, not web dev.
If you actually cared, you'd be out there using your vaunted skills to improve the software other people use, instead of slagging off other developers on Hacker News.
/snark
Bluntly, web devs often screw up in the fundamentals of algorithmic operations and data. Web dev comes out of the horrific slap-it-up-i-tude of the HTML/Perl days of the mid-90s, and its tooling is still incredibly shoddy compared to desktop development. And the really fun part is? Web devs don't even get it. They often think they are the top of the food chain, with the best tools ever built. I still can't even find a tool to match VB 5's capabilities.
One example, from Chrome: https://groups.google.com/a/chromium.org/forum/#!msg/chromiu...
but, yeah, I've seen some awful non-web code. :)
Edit Could you clarify? Which part of what I said is "bullshit"?
Why do we keep conflating the two?
[1] Note that I said "most programming".
I like to think I take a certain amount of rigor in the choices of tools and processes and design philosophy that reduces the amount and impact of bugs... but if we get a customer complaint about our product we don't generally issue a recall and lose millions of dollars.
It's not that I spend any less time learning theory and application and it's certainly no less challenging in some cases than even mechanical engineering but... it's a liability thing.
Also, I don't write software for aerospace control systems.
I've seen companies advertise "software engineer," positions whose primary responsibilities included running a fleet of Wordpress blogs.
Just a matter of perspective I guess.
Certain segments of our industry are currently engineer heavy, such as embedded, and some appear to be technician heavy. I don't see that as a problem per se but it clearly causes friction occasionally as we tend to conflate them.
Ideally, they overlap.
Our field is maturing. These difference are only going to become more important.
+1
Back in time when I had to choose if I want to do CS degree I was already paid as a web developer and doing what I was going to study for. Then I asked some friends that were actually studying CS and they told me that they have a very tight schedule with work in the following disciplines : FORTRAN, ASM, C, C++ . I told them that's too low-level for me and probably I will not use it for my work so I decided to go into completely different sphere ( I graduated law ).
Now I'm not the best programmer lived on the planet, but certainly I still do what I love and I'm paid for web development as high as a senior person, because of my experience - not CS degree.
So... If you think you can make benefit to web development, why don't you start with a simple open source contribution to any of the existing projects and reduce Wirth's law [1] a bit with some low-level skills. But for something more complicated, please consider some experience first in that specific area.
Web development is interesting, because it tends to mix in people from a lot of different backgrounds — in particular, some of them come through the design side, and move down the stack. That's good, because it demonstrates the accessibility and flexibility of the stack; it's bad because it can result in suboptimal solutions to common problems.
You probably have the luxury of specializing in one particular language, and maybe a handful of processor architectures. You probably rely on one or a few libraries, and know them backwards and forwards. You probably have your own personal repository of code and tactics that you go back to often. In other words, your knowledge of programming is probably narrow and deep.
Web developers usually don't get that luxury. Top web developers today have to be fluent in a minimum of two programming languages (javascript and a backend language like Python, Ruby, PHP...), use several different fast prototyping approaches (css compilers and wacky templating systems), have a good working general knowledge of everything from the web browser to the web server, and cope with an environment in which at least one of those parts is changing on almost a daily basis. Web development is shallow and very, very broad.
I prefer systems programming but I've worked as a web developer off and on for several years now. I have a lot of respect for any web developer that's really good at it. They aren't lesser programmers at all, and I suspect a lot of them, if they decided to do it, could kick the pants off of most systems programmers.
Of course not. Stop generalising.
http://stackoverflow.com/questions/3811678/add-two-variables...
Edit: found it https://plus.google.com/u/0/+DougTyrrell/posts/br3kqg6Vet6
GoF is a (somewhat C++-centric) book about design patterns that's completely unrelated to the discussion at hand, may not be the best resource to learn about design patterns and has nothing to do with CS.
TAoCP is more like an encyclopedia; actually reading through even one chapter takes a significant amount of effort (if you want to get anything out of it). Try it. (Yes, I've worked with it.)
K&R is a 27 years old, thoroughly outdated book about C. There are better options[1].
[1] Try the 16 years old book "Expert C Programming" by Peter van Linden, which is excellent, even though outdated too.
I don't see any evidence that the Ruby community suffers more than any other development community from this sort of thing — that is, the ones where high performance is not the biggest concern, of course.
The allocations here (ruby) are reduced because the implementation of appending is horribly slow in the first place, using defensive cloning (I'm taking jph's word here).
http://ruby-doc.org/stdlib-1.9.3/libdoc/matrix/rdoc/Vector.h...
Yeah, down the line you will eventually have to do optimization, but you will prioritize.
Of course, if you have 10k users and it runs on 3 machines, you got a problem which no amount of boxes can solve.
Once a more holistic view is taken, wide spread total costs and benefits are taken into consideration, once costs are not only defined as money flowing out of my own pocket, once not only "Gesinnungsethik" but also and more importantly "Verantwortungsethik" gets applied, well,
in such a world we would probably wish, that Amazon would change its pricing policy to:
- get the first 2-5 AWS instances almost for free
- and pay for the next few 100 exorbitantly much more money.
We all would benefit from the cultural, technological and social changes this would help to spark, I think.
But who am I to question current culturally entrenched "economic" thinking...
It changes when the context are not "marginal EC2 instances" but instead energy and resources burning machines, used (often) by ignorant software developers and their organizations allowed and actually encouraged, partly even actively driven into such purely self beneficial behavior models.
For a definition of social: http://en.wikipedia.org/wiki/Social ... obviously driving software development into a scarcity of computing power would haven "social" consequences: in the development teams f.e. interactions and priorities would need to change dramatically.
But maybe software development turned almost into a "commodity" because of the commoditization of computing power available to even the most ineffective mental artifact aka program.
And that in turn was possible in large extent by off-loading the true costs of assembly / dis-assembly / disposal and the resources needed to build those machines onto people in underdeveloped regions of the world...
Now what could the social change be, the more expensive computing devices could allow for in those regions?
If people cared about "effectivity" not only via a detour to "uh, i need to recharge my phone, again?!" f.e. but essentially because because they would have to pay the true cost for their ineffective setup of hardware and software?
What could the social change be... ;)
It changes when the context are not "marginal EC2 instances" but instead energy and resources burning machines, used (often) by ignorant software developers and their organizations allowed and actually encouraged, partly even actively driven into such purely self beneficial behavior models.
Fair enough, and I agree with you that more reflection is needed on the ethics of our industry. No argument there.
That said, I think you're miscalculating the result of such a switch. The fact is that servers are pretty efficient.
Say one of the developers commutes to work, doing ~12 miles each way on her/his Prius. If he works for two days optimizing the code, the energy cost of his commute will be ~130kWh.
With that same energy, you can run a PowerEdge R420 on full power (CPU benchmark) for almost 40 days! And remember that each of those would power a bunch of EC2 instances.
The reason EC2 instances are cheap is because they're actually cheap, both in terms of energy and resources.
Sure, the servers are more efficient than they used to be, and as stated elsewhere yes, because they are shared they are more efficiently used than machines which are not shared.
"they're actually cheap, both in terms of energy and resources."
... well I disagree with this: they are cheap to us because we don't pay adequately for them: not for the energy, the labor and not for the rare earth elements f.e. All of which quite conveniently is actually payed for in just very few regions of the world: by the people there.
Sure "the market" came up with this prices but the same market simply ignores certain kinds of costs, which are not visible to us. One name for those is: externalities. One of those is the "total energy budget" needed to build and dispose such a machine and the power needed when running the quite often ineffective apps. I'd add a whole bunch of political / sociological costs to that.
And yes, I know that the real "power benefits" of optimizing code just don't add up into a significant number today. I think this is due to the externalities we don't pay for. I can imagine a world where it would be economically justifiable to pressure for effective code. Today only the very big fish actually feel the need to make some of their code effective. The others just consumer what is already prepared for them: the machines already running at the centers.
As of the impact of the switch: I did NOT calculate. So you are technically right (probably ;) and only as far as you chose (maybe not consciously) the boundaries of your model ;)
When a few years ago the power consumption of data-centers appeared in the world wide energy consumption overviews, I guess "we" knew, yes they are a big deal.
Cheers
But we also don't pay for the full costs of the stuff needed to have programmers optimize those programs.
A watt used to run an EC2 instance would become more expensive in the world you're suggesting - but so would the watt used to power the light while the developer worked on optimizing the code, or the watt used to power his car, etc.
So unless there's any particular reason why the costs of running EC2 instances are particularly less "priced in" than other costs, I see no reason to think the balance between the two options would change significantly.
When I say they're cheap, I mean relatively. Whether they're cheap in absolute terms (which is what you're arguing) is irrelevant to my argument.
To make an analogy, a bowling ball still weights more than a feather, even if you measure it on the Moon.
The social change in those regions would be the factories closing down and moving to other countries, as the low prices would be no longer so relevant, and so the workers would return to the famine and poverty of the 60s and 70s that they are just beginning to escape.
"Programmers waste enormous amounts of time thinking about, or worrying about, the speed of noncritical parts of their programs, and these attempts at efficiency actually have a strong negative impact when debugging and maintenance are considered. We should forget about small efficiencies, say about 97% of the time: premature optimization is the root of all evil. Yet we should not pass up our opportunities in that critical 3%."
I have the problems with many happy-coders that they only remember the part about about not optimizing early and often forget that part when you should measure performance, find bottlenecks and get rid of them.What people often miss is the measuring part.
Linked lists should be named "linked graphs" instead.
There is so much relevant science to learn about CPU caches, than there is about using a container which is based on nested pointer indirections.
Just have to follow the data and watch how its used and build your program to provide the simplest flow.
Personally, I really dislike Ruby's syntax, though I haven't spent a huge amount of time with it (because I dislike the syntax). The use of bracers and other lexical markers makes code a lot clearer and faster to decipher, imo, than a bunch of def and ends. (I know that () are optional in Ruby, can you also use {} if you desire? Again, not 100% familiar with the language features. Just know some standard rails implementations of the language).
Maybe it just hasn't 'clicked' with me yet, but bleh. The dynamic typing doesn't help its case in my book either. That's my personal preference, and why I try to avoid using ruby, even for the web backed by rails despite it's popularity. Then again, if you need to get a web project up and running quickly rails is never a bad choice (in my experience).