Perl Data Language: Scientific Computing with Perl
pdl.perl.org
pdl.perl.org
The last release wasn't in February, it was just last week! <https://metacpan.org/release/ETJ/PDL-2.050>.
I agree with many of the commenters here that Python has a lot of great libraries and is a major player for scientific computing these days. I also code in Python from time to time, but I prefer the OO modelling and language flexibility features of Perl.
Speaking for myself and not the other PDL devs, I don't think this is an issue for Perl-using scientists as Perl can actually call Python code quite easily using Inline::Python. In the future I will be working on interoperability between the two better specifically for NumPy / Pandas. This is also the path being taken by Julia and R.
Do you have a tutorial and some examples? If not, could you write one?
I sometimes deploy perl code at large scale for financial computing where only performance matters: with XS the overhead is low while gaining language flexibility.
Even in 2021, this is usually faster than alternatives by orders of magnitude.
PDL could be a good addition to our toolset for specific workloads.
I can share some examples of using PDL:
- Demos of basic usage <https://metacpan.org/release/ETJ/PDL-2.050/source/Demos/Gene...>
- Image analysis <https://nbviewer.ipython.org/github/zmughal/zmughal-iperl-no...> (I am also the author of IPerl, so if you have questions about it, let me know. My top priority with IPerl right now is to make it easy to install.)
- Physics calculations <https://github.com/wlmb/Photonic>
- Access to GSL functions for integration and statistics (with comparisons to SciPy and R): <https://gist.github.com/zmughal/fd79961a166d653a7316aef2f010...>. Note how PDL can take an array of values as input (which gets promoted into a PDL of type double) and then returns a PDL of type double of the same size. The values of that original array are processed entirely in C once they get converted to a PDL.
- Example of using Gnuplot <https://github.com/PDLPorters/PDL-Graphics-Gnuplot/blob/mast...>.
---
Just to give a summary of how PDL works relative to XS:
PDL allows for creating numeric ndarrays of any number of dimension of a specific type (e.g., byte, float, double, complex double) that can be operated on by generalized functions. These functions are compiled using a DSL called PP that generates multiple XS functions by taking a signature that defines the number of dimensions that the function operates over for each input/output variable and adding loops around it. These loops are quite flexible and can be made to work in-place so that no temporary arrays are created (also allows for doing pre-allocation). The loops will run multiple times over that same piece of memory --- this is still fast unless you have many small computations.
And if you do have many small computations, the PP DSL is available for the user to use as well so if they need to take a specific PDL computation written in Perl, they can translate the innermost loop into C and then it can do the whole computation in one loop (a faster data access pattern). There is a book for that as well called "Practical Magick with C, PDL, and PDL::PP -- a guide to compiled add-ons for PDL" <https://arxiv.org/abs/1702.07753>.
---
I'm also active on the `#pdl` IRC channel on <https://www.irc.perl.org/>, so feel free to drop by.
I did not know at the time any of the specialized languages, so intially approaching the project -- I was very concerned on how to deal with matrices, but as I got to understand the PDL better -- i was getting better and better at it.
If I may suggest someting (this is based on the old experience though) --
a) some 'built-in' way to seamlessly distribute work across processes and machines.
b) some seamless excel and libreoffice calc integration.
Meaning that I should be able to 'release' my programs as Excel/Libre Office files.
Where I code in PDL but leverage Spreadsheet as a 'UI' + calc runtime.
So that when I run my 'make' I get out a Excel/Libre office file that I can version and distribute into user or subsequent compute environments.
Where the PDL code is translated into the runtime understood by the spreadsheet engine.
I know this is a lot to ask, and may be not in the direction you are going, but wanted to mention still.
a)
A built-in way would be good. There is some work being explored in using OpenMP with Perl/PDL to get some of that. In the mean time, there is MCE which does distribute across processes and there are examples of using this with PDL <https://github.com/marioroy/mce-cookbook#sharing-perl-data-l...>, but I have not had an opportunity to use it.
b)
Output for a spreadsheet would be difficult if I understand the problem correctly. This would more about creating a mapping of PDL function names to spreadsheet function names --- not all PDL functions exist in spreadsheet languages. It might be possible to embed or do IPC with a Perl interpreter like <https://www.pyxll.com/>, but I don't know about how easy that would be to deploy when distributing to users.
Am I understanding correctly?
Interestingly enough, creating a mapping of PDL functions would be useful for other reasons, so the first part might be possible, but the code might need to be written in a certain way that makes writing the dataflow between cells easier.
As a "heavy" user of scientific computing, I must say that the name "data language" is a bit disheartening... It echoes of useless "data frames" not of cool "sparse matrices" which is what I actually need. Does PDS support large sparse matrices? I grepped around the tutorial and the book and the word "sparse" is nowhere to be found. Yet it is an essential data structure in scientific computation. Are there any plans to, e.g., provide an interface into standard libraries like suitesparse?
All the comments suggesting that this somehow shouldn't be and that people should move away from PDL, are depressing in the way that, in a split second, years of effort and thousands lines of code that are probably doing good work are dismissed just like that.
If a bank is still using Cobol, it is interesting and a testament to how a Cobol programmer can still make a good living on it. But if a scientist is using Perl to carry out calculations, this is somehow bad?
Yes. The negative externalities of doing scientific analyses in Perl are much greater than a bank having a legacy COBOL codebase. Only a handful of engineers within the bank will ever see that COBOL codebase. Science is globally collaborative; many people at many different institutions across the world would have to deal with some idiosyncratic scientist’s decision to write their analysis in Perl.
Also, the bank only has a COBOL codebase because it’s reluctant to make major changes to an extremely important system that’s been working flawlessly since the early 60s. There’s absolutely no reason to start a totally brand new project in Perl (or COBOL, for that matter), when far superior alternatives exist.
This is the same nonsense game people play where someone says "that's like taking a ship to cross the ocean instead of a plane" and someone else says "I like taking ships, I want to get fresh air for three weeks of solitude instead arriving the next morning"
Everything about UNIX is oriented toward being a programming environment in itself. There are plenty of developers who just use UNIX as their IDE. Drew DeVault, as an example, is pretty notorious for it.
Complexity distracts and makes efficiency impossible. Modern IDEs are nothing but complexity. UNIX-as-IDE simplifies.
It can learn to compensate for a missing tool, arm, eye, ...
And after a while it feels just natural.
ed is actually really useful if you're wanting to rapidly iterate on a program. There's a reason Ken Thompson advocated for it up to his retirement (which was very recently, mind).
- bc|dc/perl/little of C
- Gnuplot
- entr + make
Amazing performance with very little usage of resources.
Install cpanminus and then install local::lib.
cpanm local::lib
cat >> ~/.profile << EOF eval $(perl -I ~/perl5/lib/perl5/ -Mlocal::lib) EOF
done.
In my experience, looking at "scientist" produced code, the programming language matters very little. It is not hard to produce something completely inscrutable and non-replicable in Python and R the same way it's been done for ages using SAS, Stata, MatLab etc.
I still see people rolling out their regression computations using matrix inversion and calculating averages as `sum(x)/n`.
I really like PDL when I can use it. I have had problems building it from source on Windows in the past, but it is actually a very well thought out library.
Also worth mentioning, you can get a lot of mileage out of GSL[1].
It’s possible to write inscrutable code in any language, but some languages sure make it easier.
Syntax issues aside, the main advantage of Julia/Python/R (the latter’s syntax might even be worse than Perl’s) for scientific computing is their ecosystems. A language for a particular use case is only as good as the packages available for that use case. The scientific package ecosystems for Ju/Py/R are far richer than that of Perl, simply because their userbases are much larger. Thus, a scientist using Perl would likely be forced to roll a lot of their own functions, which makes the code idiosyncratic and more likely to contain bugs. (To use one of your examples, people might implement OLS regression by manually computing the hat matrix because no stats package exists for the language they’re using. Now imagine their language lacks something more complicated, like a robust MCMC sampler package à la PyMC3 or STAN, and they have to roll that themselves. Yikes.)
And that’s not even getting into the value of the languages for interactive scientific computing, which is how most of it gets done these days. For instance, there’s no official Jupyter notebook support for Perl (although unofficial plugins exist, they don’t support inline graphics/dataframes/other widgets), and the REPLs for Julia/Python/R are much more modern and fully-featured than PDL2.
BTW, I agree the GSL is great for building standalone tools, but it’s totally irrelevant for any interactive work.
CPAN predates all of those.
The python pystan is wrapper that ships data to/from the Stan binary and marshals it into a python-friendly form; I think Julia’s is similar.
I’m not exactly volunteering to do it, but a PerlStan would not be that hard to implement. As for scientific communication, a point you raised above, I don’t think it’d be too bad. Most readers of a paper would be interested in the model itself, and that would be written in Stan’s DSL regardless.
But tons of other numerical methods are also missing from Perl. To use another stats example, in another comment, I gave the example that PDL only supports random variable generation for common distributions (e.g. normal, gamma, Poisson). Anything beyond stats 101 level and you’re on your own.
One could do the same with Perl, and in fact, people have. If you need random variates from a Type 2 Gumbel distribution, for example, Math::GSL::Randist has you covered https://metacpan.org/pod/Math::GSL::Randist#Gumbel
Honestly, I'm not rushing to convert our stuff to PDL, but I did want to push back a little on the idea that python is The One True Way to do scientific computing. It's a fine language, but I think a lot of its specific benefits are overstated (or mixed in with the general idea of taking computing seriously).
Also note that PDL does automatic broadcasting of input variables so it does an entire C loop for an array of values being evaluated. See this example <https://gist.github.com/zmughal/fd79961a166d653a7316aef2f010...> for how that applies to all GSL functions that are available in PDL. Though I do notice that some of the distributions available at <https://docs.scipy.org/doc/scipy/reference/stats.html#contin...> are not in GSL.
Though when I do stats, I often reach for R and have done some work in the past to make PDL work with the R interpreter (it currently has some build bitrot and I need to fix that).
Not sure how official support would work in the Jupyter Project since anybody can write a kernel. I wrote the Perl one (IPerl) and that has existed since 2014 (when Jupyter was spun off from IPython). It supports graphics and has APIs for working with all other output types.
Now I do need to help make it work with Binder, but it does work.
---
The other point about MCMC samplers is valid. This is why I wrote a binding to R to access everything available in R and why I use Inline::Python sometimes. I should create a binding for Stan --- should not be hard --- at least for CmdStan at first, then Stan C++ next.
`statistics.mean` uses `_sum` which tries to avoid some basic round-off errors[1]. I think the implementation of `_sum` is needlessly baroque because the implementors are trying to handle multiple types in the same code in a not so type-aware language. Regardless, using `statistics.mean` instead of `sum(x)/len(x)` would eliminate the most common rounding error source.
As for statistical modelling handled by directly inverting matrices, there the problem is singular matrices that appear to be non-singular due to the vagaries of floating point arithmentic in addition to failing to use stable numeric techniques.
The point remains. The detriment to science are people who convert textbook formulas directly to code instead of being aware of implementation with good numerical properties.
Note:
>>> x = [1e9 + .1, 1.1] * 50000
>>> sum(x)/len(x)
500000000.60091573
whereas >>> import statistics
>>> statistics.mean(x)
500000000.6
See also my blog post "How you average numbers matters"[2].> Now, in the real world, you have programs that ingest untold amounts of data. They sum numbers, divide them, multiply them, do unspeakable things to them in the name of “big data”. Very few of the people who consider themselves C++ wizards, or F# philosophers, or C# ninjas actually know that one needs to pay attention to how you torture the data. Otherwise, by the time you add, divide, multiply, subtract, and raise to the nth power you might be reporting mush and not data.
> One saving grace of the real world is the fact that a given variable is unlikely to contain values with such an extreme range. On the other hand, in the real world, one hardly ever works with just a single variable, and one can hardly every verify the results of individual summations independently.
Correct algorithms may be slower, but I am hoping that it is easy understand why they ought to be preferred.
[1]: https://github.com/python/cpython/blob/5571cabf1b3385087aba2...
[2]: https://www.nu42.com/2015/03/how-you-average-numbers.html
As this is worth to be better known, I submitted it here: https://news.ycombinator.com/item?id=27470323
What do you mean?
This is due to two main reasons. For one, Perl’s incredible syntactical flexibility makes it easy to write “clever” one liners that are hard to comprehend. Speaking from experience, scientists tend to be the sort of people who value “clever” code over “clean” code.
Secondly, the Perl package ecosystem for numerical/scientific methods just isn’t as fully featured as the Julia/Python/R ecosystems. This leads to individuals having to reimplement methods themselves, which leads to idiosyncrasies and likely bugs. For instance, I see no way to generate Wishart random variables (or RVs from other distributions beyond the common ones) in PDL. Julia, Python (via SciPy), and R all have full featured support for many different distributions beyond the common ones.
A Perl user would thus have to implement this themselves. Someone else reading their code would have to both familiarize themselves with the custom function’s syntax (as opposed to immediately recognizing the standardized scipy.stats.wishart, which behaves like any other scipy probability distribution class) and likely check for any bugs, since a standard package is far more likely to be correct than some random one-off function. I’ve had the unfortunate experience of working with someone who refused to use off-the-shelf libraries for numerical methods, and unsurprisingly their code was not only hard to read (since there was no standardization) but also full of bugs.
Just to the other people that never have the ambition, desire or toke the time to develop new valuable skills. This is not a bad thing necessarily. Those people are time sinks.
In any case you can take the decision to document or not your code in any language. "; # this line does that" is not hard to write.
And different science teams collaborate, but also compete for money. So everybody doing the same can lead to "the more dishonest takes all" and kills all the other teams. Sometimes is useful to protect yourself from the people trying to backstab you and steal your best tools. Tools that toke you decades to develop and polish. Not ALL is freely shared in science.
And yes, my Perl scripts were a triple headache to write, but still work flawlessly after all this years.
Perl is just a tool to do something, and you should never use one (and the same) tool for everything unless your goal is to be a mediocre scientist sucking from other people's efforts all the time.
It is no harder to collaborate with Perl than with R (been there done that) or Julia (have not done that.
PDL is addressing your point about libraries.
Horses for courses. If your research involves streams of text (a lot of things are streams of text) and transforming them then Perl is a likely contender as the best choice.
1/3 of the posts here are "<old tool that already exists in a stable, mature codebase> in {rust|go|whatevernewlanguagecomesnextweek} released v0.0.1"
Perl is a great language, that does it's job for many, many things, especially with CPAN, and it has been doing so for years. You can buy a 20yo book on perl, and 99.99% of the example code from that book still works, and same goes for projects from that era (which cannot be said for python, where developers and distro mainanters seem to enjoy removing usable, mature projects, just because they're written for python2.7 and incompatible with 3+).
If I have to write a script once, that I can forget about, and just expect it to run for years, perl will always be my first choice.
Orelly's Perl books from the CD bookshelf still work.
Just declare a variable with "my =" in front of it (just once), and everything will work as usual:
old:
$num = 3;
print $num;
new: my $num = 3;
print $num;$ perl
$num = 3;
print $num;
3(perl v5.32.1, without "use strict" of course)
As I always used
use strict; use warnings;
as something like muscle memory, I didn't know this.
use strict;
no strict qw(vars);
$foo = $bar;
That's part of the reason `use strict` is recommended.---
The other major reason is `"refs"`, which disable symbolic refs. Honestly this is *THE* main reason I recommend `use strict`.
Symbolic refs are how you did arrays of arrays prior to Perl5. (Among other uses.)
use 4.0;
@a = 'b','c';
@b = (1, 2);
@c = (3, 4);
print $a[1]->[1], "\n"; # 4
That can be a security risk if an attacker can insert or change strings in `@a`.It isn't (generally) needed anymore of course.
use 5.0;
my @a = ( \[1,2], \[3,4] );
print $a[1]->[1], "\n"; # 4
(The only reason `@b` and `@c` existed was to symbolically reference them.)I am biased. I love Perl and hate Python. Makes me feel very old....
The only thing holding me back is if I want to use ${library} I probably can't do so from Perl.
No it isn't. It's a testament to how backward that bank is. You'll see upvoted contrary takes here, sure, but that's because middlebrow contrarianism is a good way to get upvoted on HN.
Perl is superior in many ways and I think using it for data-exploration still has tons of merit.
Folks “shoot themselves in the foot” when they don’t understand the language. In Perls case: list/scalar context is usually the culprit, but it is quite easy to understand. It’s more flexible and concise which in many ways makes Perl better at exploratory programming.
The question is whether Perl's PDL is better than ScyPy/Numpy/Pandas/Cython, the Spyder IDE, bunch of plotting libraries...etc.
The fact is that most people don't even have an opinion about what is Python or Perl.
Python can seem easy to share after you copy part of a python script and miss one blank space somewhere at the end of a line or start copying in the wrong line. If you use a dumb text editor the script will easily turn into a ugly mess. Perl scripts don't have this problem, so some people could say that they are easier to exchange and share in fact.
Many Perl authors will be really glad to share your code with you. This is what they built an online community of knowledge and libraries called CPAN where you can find it easily. To find authors willing to help and explain obscure parts of their own scripts also if asked politely, is not uncommon or particularly difficult.
I think we can bury this 'Python has significant whitespace' criticism for good now. I have taught python to people from hugely diverse backgrounds (including literature and law) and not once has this been an issue (on the contrary it's a massive help to readability).
Also, no sane person teaches people to code python with a 'dumb' text editor - you give them Jupyter notebooks or VSCode or PyCharm (which has an excellent educational version) or Notepad++ or something.
You realize CPAN did it first right?
Python vs Perl for science has nothing to do with c/c++/fortran. C/C++/fortran have their own place for science computations. Perl does not. For anything involving any kind of numerics Python will be faster/better tested/will have more libraries/will have better visualisation capabilities, so there is no need for perl.
Sure if you are only working with strings, you can use perl, but that's hardly scientific computing (not at least the field that I work in).
It depends on the field. RNA/DNA, Polypeptids and proteins are just long chains of text, therefore Perl can deal easily with the problems of finding things, manipulating them to build a new chain or translating the chain to a different format. This is a significant chunk of what Bioinformatics do all the time, and Bioperl can manage it.
Also a big part of astronomy is analyzing or finding stars in a tridimensional matrix of space. PDL can be useful with that. Is not dificult to extract a slice of interest in the space matrix and focus our research on it. The main problem could be the lack of experienced people available having exactly this problem to solve.
I don't know if original Perl is very good or bad for that, but there are several Math modules that could have what you want and be easily connected with the former stuff. In fact there are a lot of them to choose:
https://metacpan.org/search?q=Math
Raku at least has a Math::Model module to simulate Physics stuff. I ignore how developed is the module or how its perform would compare with Pyton's similar stuff
That’s just the surface. Working with matrices means nested data structures, and Perls syntax for nested data structures makes the list/scalar context look like child’s play.
Having said all than I have no problems using references in Perl, but see why a Pythonista can find Perl code hard to read.
I like python a lot and actually i wish less scientists used it. Perl/PDL is in my opinion much better suited to this (understandable) just getting the job done approach i often find in sciene or other areas in which writing software is not the primary goal.
Also - they can't even get their SSL certs straight.
I loved Perl - it was certainly my gateway to coding. Regex was beautiful in perl. But I shot myself in the foot so many times with that language.
https://github.com/dkogan/numpysane/
Now the core numpy has usable broadcasting, concatenation and basic linear algebra. Kudos to the PDL team for the excellent core design.
https://mxnet.apache.org/versions/1.8.0/api/perlBack then there was also an attempt at making GDL (gnu data language), but the tooling for IDL was so deep I doubt anything can directly replace it. Then again I've been out of this for twelve years.
https://www.easterbrook.ca/steve/2010/12/agu-session-on-soft...
Oh, you need to understand the underlying maths well in order to code your functions right? That's the biggest issue in data science management. Too much relying on specialized " biggies" like Numpy/CUDA with atrocious codebases where the calculations lasts months compared to a 1 hour chore with C or even Perl with Gnuplot.
http://phroxy.z3bra.org/bitreich.org:70/0/con/2020/rec/energ...
It's C, well, but C, Perl and Unix are related cousins and prototyping in Perl it's really fast.
Searching more I see a Python one appears in 2009: https://www.oreilly.com/library/view/bioinformatics-programm...
And now there are several out. Plus expanded topics like Python and Machine Learning, Python and Data Analysis, etc.