THAT statement is dead wrong. Or at least, in dead opposition to my approach to learning. If someone says five things, and four are wrong but the last one teaches me something new or gives me a new perspective on something I already know...
Reading the four things cost me maybe two mintes. The last one cost me thirty seconds to read, plus another couple of minutes to savour. The total investment is maybe ten minutes and I learned something.
Of course, if I skip everyone who seems to be wrong right off the bat I'll avoid some situations where all five are wrong. But I have something to help me out with that: The post is on HN and some other nice people are voting it up. So something must be right, why don't I navigate into the cave, around the pit of vipers, past the deadfall, and onto the intellectual treasure?
And in fact, I read past your first statement and liked the rest. See how that works?
;-)
Your strategy for filtering information obviously works for you, and my guess is that if you missed something good, it'll come back to you later when someone you respects says "Hey, I know Raganwald is usually a waste of time, but did you see X?"
So... Rock on! And yes, I enjoyed reading your comment and the reply to my thoughts. Thanks for taking the time to post.
The sort of operations commonly done on strings only have a small overlap with list operations, though. I'm all for orthogonality in a language when it's practical, but I'm not sure that one is a good trade-off.
Erlang does that, and aside from space issues (on 64-bit, a linked list of character code ints means 16 bytes per char, due to alignment), it's kind of awkward. While using pattern matching on individual chars is nice, in those cases it probably makes more sense to tokenize the strings into atoms/tuples via REs, and then pattern-match on those instead. Binaries can also be used as pseudo-strings, but to me it feels like two half-solutions. (And its shell's printing a list of ints as chars when they're printable 7-bit ASCII characters is a bit odd.) Having atoms (AKA symbols) helps, though.
Incidentally, Lua's strings are all immutable atoms. That's practical most of the time, but occasionally a major liability. Lua has the obvious option of just using a different string implementation via userdata when it's a problem, though, so it can make such a daring trade-off.
Of course, just saving strings as arrays with a \0 at the end is also pretty brittle.
A lot of people also complain about the performance of strings in Haskell, so this is right where the tradeoff is today. Tomorrow it won't likely be an issue.
A really smart compiler should be able to optimize most of the inefficiency away; I'd bank on compilers improving fast enough to make seemingly inefficient things cost much less.
If I'm designing a language today for the future, I'd err on the side of usefulness.
Lua does almost no optimization at compile time* (in part so it can compile data dumps rapidly) but I find it easy to predict the performance trade-offs in various implementation choices. In a roundabout way, it still ends up easy to tune.
Still, usefulness and performance are only indirectly related. I think that was PG's original point.
* Yet oddly it ends up quite a bit faster than Python, and there's also LuaJIT.
Even if it can be made configurable, adding all those knobs and dials may add a lot of complexity.
It's useful from a convenience of programming point of view. The preferred underlying implementation is now something more sophisticated than a linked list of characters. See Data.ByteString. The programmer does not need to care much, though.
You know, people love to say this, but in every company I've worked at (sample size 3), past the first year or so of operations, programming language / framework speed actually was a very significant concern for the business.
This may be self-selection; scaling and performance are two of my areas of specialization. However, whenever I'm on the job market, there doesn't seem to be any shortage of companies trying to hire people to help them handle their performance and scaling problems.
Instead of "strings represented as lists", I think it is more useful to say "strings and lists implementing a common interface." For example, in Clojure, it is perfectly fine to treat a string as a sequence.
> (doseq [ch "abc"]
(println ch))
a
b
c