All Software is Legacy
leejo.github.io
leejo.github.io
1) The perl cpan module doesn't resolve dependencies
2) The cpan module has parsing errors when passing in a list of CPAN packages
3) You have to manually grep your perl code to see what modules it depends on
4) Module installs take a long time since they can compile and unit test the code, unit tests can even make connections to the internet or try to access databases and fail, so you just have for force them to install
5) Non-interactive installs of CPAN modules requires digging in the docs and learning you need to set an env var to enable
6) CPAN modules aren't used that heavily and can have bugs that would be caught in wider used modules. (e.g. the AWS EC2::* modules don't page results from AWS so results sets can be incomplete, whereas the wider used boto lib works correctly and is better maintained.)
7) Perl devs don't think twice about shelling out to an external binary (that may or may not be installed)
8) Even if regexs are not needed, inevitably the perl dev will use them since that's the perl hammer, and it's hard to know what the intention is with regexes or what the source data even looks like
9) You have to manually include the DataDumper package to debug data structs
10) You have to manually enable warnings and strict check, it's not on by default.
Anyhow, I think we've made a lot of progress since the 1990s. :)
* It is often recommended to use cpanminus[1] instead of the CPAN.pm module. But it is up to the distribution you try to install to declare it's dependencies correctly. Not doing that is a bug.
* If you use cpanminus you can use the --notest flag to skip tests. But tests are a feature.
* Software have bugs. Reporting them when they are found is how software get less bugs.
* Cpan distributions should not[2] use external binaries (and exceptions should be clearly documented and motivated).
* The ease of use of regexes in Perl is not an argument for not documenting them (and in this case) the document format they are meant to parse.
* There are several different data dumpers. No assumption on the user's preference is made.
* If you use a newer Perl (5.12+) you get strict enabled automatically[3], and also (depending on which version your code requires) some new features. Due to backwards compatibility it is not possible for newer Perls to enable strict or warnings implicitly.
The Perl of today is also vastly improved since the 1990s, hopefully you will come across some modern perl too.
[1] https://metacpan.org/pod/App::cpanminus
[2] https://www.ietf.org/rfc/rfc2119.txt
[3] https://metacpan.org/pod/release/JESSE/perl-5.12.0/pod/perl5...
In the CPAN case, if cpanminus is the "good one", then it should be installed by default and CPAN.pm needs to tell you to use that instead or just be deprecated. I don't want 5 choices in package managers, I just want the good one. :)
Another issue is discoverability. A concrete example is that https://metacpan.org/ is a much better (imho) presentation of cpan than http://search.cpan.org/.
It is the curse of being a very stable language and ecosystem.
No, most of them do. Perl ecosystem has a killer feature called cpantesters, that allows everyone to see which modules work on which systems out of the box. You should always check cpantesters matrix before choosing a particular dependency.
> Even if regexs are not needed
They got overly complicated over the years, but they are needed. They are DSLs to make things easier when working with strings. I.e. so you wouldn't have to write 20 lines of hard to grasp code with bytes.Index(), bytes.HasSuffix(), bytes.TrimRight(), etc., like people do in Go, but a single nice regexp and therefore reduce your chances to make a mistake in that code.
Go has regexps, and a very good implementation at it.
Depending on what you do and on the specific code-path, compiling and/or executing a regexp might be slower than manually parsing the string. Go standard library is pretty concerned with performance (much more than Python's or Ruby's, for instance), so it tends to avoid regexps.
Both Perl and Go implement regexps, and neither or them compile them to native code. So I don't get your previous comment at all.
The main difference is that, in Perl, if you ever had to write manual string parsing, it would be much much slower than using regexps as Perl is an interpreted language. So regexps are needed to perform fast string parsing. In Go, you have regexps if you want, or you can go even faster if you feel it's required.
Ok, I'll try to explain.
People feel discouraged to use regexps in Go, because they are very slow for many typical parsing and validating cases and require extra step of compilation and all of the additional code complexity associated with that. So, people do parsing manually instead, with all of its problems. It's not that they need that performance, almost no one does, but the whole idea behind regular expressions is not working, parsing code is still bad most of the time.
I didn't look long enough to know if there's an easy way to convert a regular expression to Ragel syntax.
In my experience, porting code from Perl to Go, Go's regexp package is vastly inferior to Perl's, in multiple areas, speed, memory, unicode handling (eg: \b works on ascii-only in Go), etc. For example, for some large regexps handling url blacklists, reduced programmatically with Perl's awesome regexp assembly tools, I had to rely on PCRE in the end, Go just could not cope with that (not even the c++ re2). I do avoid regexps, regexps are usually best avoided, and all that, but there are areas in which they are by far the best option. In those areas, I postulate, from my own experience, that Perl's implementation is king. Speed, memory usage, Unicode.
Did you try using RE2's "set" functionality?
#1. perl -MRegexp::Assemble -E'my @list = qw< foo fo0z bar baz >; my $rx = Regexp::Assemble->new->add( @list )->re; say $rx'
(?^:(?:fo(?:0z|o)|ba[rz]))
I wonder if there's anything like that for Python and Ruby.
[1]: for example, http://cpantesters.org/author/D/DAMOG.html
^ and $ can be changed to mean beginning/end of each line in the string with the /m flag.
I'm going to disagree with this one. There's lots of things in any language where it can be hard to see, at a glance, what the intention of the programmer was. That's why we have commenting. You're supposed to comment your blocks of code so that someone else can look at it and understand what that block of code is supposed to do.
Unfortunately, as far as I can tell by looking at other people's code, I appear to be one of the only programmers on the planet who actually uses comments....
Anyway I agree, perl has made huge progress since the 1990s. I also agree there's a problem with discoverability in some parts of the cpan ecosystem. Be sure to read the Modern Perl book next time you need to do some perl work. You ought to be pleasantly surprised. Personally with the Moo(se)? family of modules, I enjoy having a multiparadigm language with reasonable optional runtime typing to keep me sane. My biggest complaint is the reference counted garbage collection.
What? CPAN absolutely does.
> 2) The cpan module has parsing errors when passing in a list of CPAN packages
Both from the commandline, and in CPAN itself can i install a list of modules as such:
cpan Data::Dumper Devel::Confess
install Data::Dumper Devel::Confess
> 3) You have to manually grep your perl code to see what modules it depends onOr you can use a CPAN module for that.
> 4) Module installs take a long time since they can compile and unit test the code
Or you just install them like this, if you're confident in your system:
install Data::Dumper Devel::Confess
> 5) Non-interactive installs of CPAN modules requires digging in the docsNon-interactive installs should be using your operating system's package manager, unless you have a special use-case, in which some doc digging is fine.
> 6) CPAN modules aren't used that heavily and can have bugs that would be caught in wider used modules.
You mean "Some CPAN modules".
> 7) Perl devs don't think twice about shelling out to an external binary (that may or may not be installed)
Again, some.
> 8) Even if regexs are not needed, inevitably the perl dev will use them since that's the perl hammer
Eh, fair enough.
> 9) You have to manually include the DataDumper package to debug data structs
Data::Dumper was first released with perl 5.005
> 10) You have to manually enable warnings and strict check, it's not on by default.Same in JS, and similar with other languages.
> Anyhow, I think we've made a lot of progress since the 1990s. :)
Not really sure, the trolling culture seems to still be the same as back then.
This is not Perl related, but I'm currently working on a developer tool that makes this part of the job easier. It's a source explorer for C/C++ named Coati that simplifies navigation within source code and thereby makes understanding the implementation faster and easier. https://www.coati.io/
Modern programming practice emphasizes tests, which is great. However, not all kinds of tests can be written so far. So we end up doing manual certification work everytime we release or publish a new version of software, for performance, fault tolerance, etc. I want to make it all automatic. Some links about my project in case you'd like to learn more: http://akkartik.name/about; http://github.com/akkartik/mu#readme. I'd love to hear your thoughts, either here or over email (address in profile).
Unit tests describe the behavior and usage of the individual systems, while integration tests describe business use cases in core workflows (usually just the happy paths). Integration tests should ideally be written in a gherkin style with feature description and acceptance criteria clearly outlined.
Tests should be optimized for readability foremost. For example, most people try and make their tests DRY (which is an abuse of DRY because it should only apply to concepts, not code, but I digress) while they should instead be making them DAMP. Each test should be able tell a story without jumping up and down around the file and outside the file. It's a lot more dangerous to misunderstand a test than it is to have a little bit of code duplication in your tests.
When I come across a project that has no documentation, I look at the tests. This isn't an excuse not to write proper documentation of course.
That's wrong on so many levels. DRY applies to code.
>while they should instead be making them DAMP. Each test should be able tell a story without jumping up and down around the file and outside the file. It's a lot more dangerous to misunderstand a test than it is to have a little bit of code duplication in your tests.
This is a problem with Gherkin. Gherkin is not particularly well suited to making tests that are both DRY and readable due to its syntax.
As a 'copyista', I would change this to "DRY should only apply to production code, not tests." It wasn't clear from your reply if you agree or disagree.
A couple of great recent links about this IMO under-discussed topic:
http://www.sandimetz.com/blog/2016/1/20/the-wrong-abstractio...
http://programmingisterrible.com/post/139222674273/write-cod...
I don't think this is true in practice. It's just that a lot of scenarios people need tests for aren't easy to replicate and most people end up not bothering and relying upon manual checking instead.
Each time the harness became capable enough to replicate a certain type of scenario it didn't take long before those kinds of bugs dried up.
The bugs then always nearly always migrated to the areas the test harness couldn't easily create scenarios for - whether that was interactions with crazy external APIs, odd timezones or weird quirky browsers or whatever.
Continually modifying the test harness to be made capable of testing bugs reported from production was often a huge amount of work but it paid off handsomely.
What can humans do, sitting at a computer, that software can't be written to do?
Before other nitpickers derail this question, I'm talking about validating a software product, not hard AI.
What can the author do, sitting at a computer, that later readers can't undo?
When phrased that way the answer is obvious: a lot. Like multiply two large numbers together, or obfuscate a program to make its meaning less clear. Both are in principle possible to undo, but the "load factor" can be arbitrarily large, making the effort for the 'reader' entirely uneconomic.
If even a human can't undo some things that the author did, how can we expect the computer to do so?
I'd change my answer to: nothing at all in the context of manual testing. And yet in practice the answer so far seems different from our in-principle argument. I'm trying to make principle and practice line up better.
My favorite bit of the article is something I've been thinking about for years:
..paradoxically the cleaner and saner your interface the more likely it is to succeed and thus more likely to become constrained by its users, to solidify.
I've been idly dreaming of a future where we decouple modules not by their interfaces but by their behavior. The major failure mode of our age is premature freezing of interfaces, and it happens because it's hard to judge how done the interface is from the outside. An interface that looks really clean can be exposed by a detailed look at the implementation that covers a corner case that hasn't been addressed yet. I'm trying instead to think about the input space of a function. I'd like in future to be able to make guarantees like this:
> The behavior of this function is fixed when argument foo is less than 5. However, for larger values we're still nailing down some details.
The argument foo might shift from being the first to the third, or its type may change in some subtle way (that still preserves previous guarantees). So upgrading will often cause some pain. However, the upside is that the pain is bounded. You never end up in some 1% situation where upgrading your OS borks up your Rails app and takes days to fix, or breaks things in subtle ways that you don't notice for months.
You know that the upgrade effort will be bounded because you know that as long as you pass in the right value into the right slot, and as long as it's less than 5, behavior won't change and all your tests will pass. That seems like it might be superior to an interface, even if it adds a little extra work up-front.
More details, in case you have any thoughts: http://akkartik.name/about http://akkartik.name/post/libraries2
a) Security issues. We need some way to allow people to quickly apply certain tiny changes without having to worry about the possibility of an interface change, while ignoring all the others.
b) Clients sometimes end up in situations where they can't control when upgrades happen. That would break things even worse with my approach than it already does. My proposal really relies on people being able to control when they upgrade.
Back to the drawing board..
Let this inspire you: http://www.bricklin.com/200yearsoftware.htm
for the kids/newbs: effectively the (co)inventor of The Spreadsheet. and consider that spreadsheets were easily the killer app for the first PC's, thus giving them much more business value, in turn causing more money to flow in, helping to create many more paying jobs, etc, in a happy snowball effect which leads to Linux, Google, AWS...
here's where you might say, "sorry, different Bricklin." oops