A Review of Perl 6
evanmiller.org
evanmiller.org
I usually don't program in Perl, but occasionally use it instead of sed.
Compare, Perl 5 vs. Perl 6:
$ time yes | head -n1000000 | perl -pe 's/y/n/' >/dev/null
real 0m0.945s
user 0m0.944s
sys 0m0.016s
$ time yes | head -n1000000 | perl6 -pe 's/y/n/' >/dev/null
real 2m49.881s
user 2m44.892s
sys 0m2.184s
Spending several minutes to do what can be done in under 1 second is just unacceptable.Bleed commits on VM for example, recently passed Perl 5 on speed with read-a-million-line-file-and-sum-up-number-of-chars bench ( https://twitter.com/zoffix/status/895684203550973956 ).
The “bleeding edge” is a more recent term for the same idea, emphasizing that being at the cutting edge might involve danger to yourself.
Go is on the other end of the spectrum; they have tuned their syntax to be fast to parse, and have stuck to implementing existing compiler technology. If your semantics are very far removed from how computers work, you need a very advanced compiler like GHC.
It sounds like Perl6 is somewhat centrist in its approach, and has innovated quite a bit; I think it will be possible to almost match Perl5's speed with reasonable effort. The interpreter might do well on small examples, but the VM will dominate in very large ones. Many compilers are tiered in that they use an interpreter for cold code, and spend the effort to generate code and JIT hot code.
We're only talking about on the order of a million bytes in your example, so 2 minutes to process on the order of 1 megabyte with a regex is just insane. If you applied it to 1 gigabyte of files that would be an even 2 days vs Perl 5 @ 16 minutes.[1] Sixteen minutes isn't enough to rewrite it in C, so Perl 5 is "fast enough". You can go get a coffee or think about the next thing you want to do with the data, or write the next part of your processing. But two days is prohibitive: it breaks your workflow.
[1]
This sounds like a disaster waiting to happen.
time yes | head -n1000000 | perl6 -e 'for $*IN.readchars { .subst("y", "n").print }’
Its slightly more complicated because I’m having it read the input in a chunk instead of on a per-line basis and I’m using subst on a literal instead of a regex.That's awesome. Thank you for joining the community here. I have a lot of respect for our elders. New technology gets written about using old technology.
Sir, when I grow up, I wish to have the same passion as you!
But I thought this review sounded pretty damning.
Numerical operations are opaque as you never quite know if it's a rational or a float.
The grammar engine is inferior to other languages 3rd party modules, hard to debug and seems to not be suitable for its main purpose (parsing perl6!)
Everything seems to be writable in a multitude of ways with no clear indication if these ways are functionally different.
I do a lot of NLP in python and I wish for better tools to munge and clean up my text data, and have been curious about perl6, but after reading this review I don't find Perl all that enticing at all.
Not so. When you put in floats, you get floats out. If you want always-rational numbers, put in FatRat and you get out FatRat -- with the predictable bad performance on algorithms that converge on irrational numbers.
> The grammar engine is inferior to other languages 3rd party modules
Nothing stops you from using superior 3rd party modules in Perl 6. Heck, you can use Python modules through Inline::Python in Perl 6.
> hard to debug
I haven't found it harder to debug than other parser libraries, but your mileage my vary. I'm a fan of Grammar::Tracer, which hasn't failed me yet.
> Everything seems to be writable in a multitude of ways with no clear indication if these ways are functionally different.
That is true, but I find it irksome that you mention it right next to Python, which also has a multitude of ways to do different things, even though the Zen of Python says there should be just one way. Think string formatting, for example. I believe it's a trait of every programming language that doesn't set out to babysit you.
Would you say the reviewer/author is misunderstanding Perl6, or that I am misunderstanding the review?
https://www.reddit.com/r/programming/comments/6thwnx/a_revie...
> Perl 6 is aware of graphemes and combining codepoints, and unlike Python, Swift, or Elixir, Perl 6 can can access arbitrary graphemes by position in constant time, rather than iterating over every grapheme to access a particular one. [...] Graphemes are stored internally as 32-bit integers; this address space is large enough to contain all legal codepoint combinations
It sounds like Perl6 is simply storing strings as UTF-32. AFAIK, there's nothing called a "grapheme" in Unicode; the closest I know of are "grapheme clusters", which are theoretically unbounded in size and so cannot simply be stored in a 32-bit integer. Maybe by "grapheme" the author means "Unicode scalar value"? But being able to access those in constant time isn't especially useful, AFAIK.
Synthetic codepoints get dynamically assigned as they are encountered during program execution.
I'm also sort of curious what their rationale is for doing any of this in the first place
The rational is constant-time indexing at grapheme cluster level, with the regex engine probably part of the motivation.
What they do is a bit like LZW compression: use a set of codes that is larger than the number of simple graphemes, assign each simple grapheme its 'normal' code and, whenever one encounters a complex one, assign it the first hitherto unused 'code point'.
That works fine until it breaks. They hope it never breaks in practice. I think there's a good chance it won't, but still don't think this is a good idea.
Indeed: It's not an exchange format, but the format of the internal in-memory string representation.
Because 24-bit integers are not a thing on x86?
The feature could be made to start a new mapping of synthetic characters before it runs out, and throw away the mapping when the last of the old strands gets GCed.
Basically it is a problem in theory, but might not be in practice.
https://cry.nu/YAPC-EU-2017/MoarVM-Internals/#/9
Link is where the string implementation starts. I don't see it detailed in our documentation, but it may be somewhere.
As part of my Unicode grant from the Perl Foundation I will try and add some implementation information to the MoarVM docs.
Let me know if you have anymore questions, Thanks!
Being able to access characters or sub-strings in constant time is useful for some (all?) character and sub-string processing.
[1] http://www.unicode.org/glossary/#grapheme and http://www.unicode.org/glossary/#grapheme_cluster
My feelings exactly.
Is the language for individuals, or for teams? Is it for software developers, or dabblers from other domains who have to crank out a bit of code?
My point was that it's not really a pro/con of a language on whether it favors individuals hacking, or teams developing, but merely a consideration for its appropriateness to your problem, and so to say "I prefer" needs to be qualified as to -when- it's preferable.
IMHO the issues are never with technology itself, it's the people. And whatever language or technology you are using won't make any difference if the people themselves don't want to work as a team.
I learned Python 3 because it seemed that it was the new one, but to great confusion Python 2 is alive and default in Linux distributions.
Things break sometimes.
In response to parent comment, Perl 6 is Perl. I don't think that it's difficult to understand that Perl is a family of languages, and Perl 6 is a more recent and sophisticated sibling to Perl 5.
All of this is bikeshedding to me.
I think that you are overlooking the fact that the vast majority of python 2 code wouldn't run by default with python 3. Yes, there are now conversion utilities that handle most (all? maybe?) of the scenarios, but the fact remains that both Python 3 and Perl 6 break backwards compatibility.
C and C++ have tried very hard to not do this. Your old C code and your old C++ code will almost always compile as is with a newer standard.
Depends on the code: Some of it won't compile without providing just the right set of flags.
In case of C, for me the crucial feature is not so much full backwards compatibility, but interoperability (you can link the object code regardless of the language version) and the convenience of using the same frontend executable.
In case of Perl, the former is currently possible to some degree via the Perl5 module Inline::Perl6 and the Perl6 module Inline::Perl5.
Personally, I would like to see the latter become a reality as well - this is the solution I'd perfer over trying to distance Perl6 from Perl5 via rebranding.
I may be imagining things but, peering years into the future, it seems this versioning management aspect of the P6 toolchain might end up being the foundation of a new multiple lang/compiler/module version juggling Perl package manager that can be used with P5, P6, python etc. Does that make sense to you?
One of the things I've heard is that `print` went from a keyword to a subroutine in Python, or something to that effect. In Perl 5 for most of them there is not really a difference, and what difference there is gets smaller every year.
Another is Unicode, which Perl 5 managed to improve without breaking most existing code. (some of it has to be enabled with a pragma statement)
Usually the only code that gets broken, is code that digs into the internals in some fashion, or use an unintended “feature”. (or occasionally even a feature that should have never been added in the first place)
One such unintended feature which relied on the compiler generating incorrect code was:
sub foo {
my $bar = 10 if 0;
say $bar++;
}
foo; # 10
foo; # 11
foo; # 12
Which was superseded by the `state` declaration. sub foo {
state $bar = 10;
say $bar++;
}
As far as Perl 5 vs Perl 6, it is more like C++ vs C# + Haskell + various other languagesPerl 6 breaks backwards compatibility in all the ways that needed to be broken to modernize the syntax, and reduce special cases. Even infix operators are no longer special.
From what I can see Python 3 breaks backward compatibility for things it didn't really need to if they tried harder. (It would have made the implementation more difficult to work on.)
Have a frontend that can invoke the perl5 interpreter as well as Rakudo as appropriate, supporting the perl5 command line flags, and arguably quite a bit of the problem goes away.
Branding is not the problem with Perl 6, though it is the easiest thing to bikeshed. The problem is compiler speed, language interoperability, and enough of a toolchain to start building serious apps. Make that good enough, and no one will care what our language is called.
https://www.haskell.org/onlinereport/
https://www.haskell.org/onlinereport/haskell2010/
The purpose of these documents is to specify what an implementation of Haskell 98 or 2010, respectively, should support. Just like the analogous standards published for the C and C++ languages, the Haskell reports make no reference to any specific implementation.
The most popular implementation of Haskell is the GHC compiler. (In fact, in reality it's the only compiler anyone uses, with very few exceptions. For better or worse, Haskell is a one-implementation language at present.) GHC accepts a greatly expanded language compared to what the reports specify, including a large number of language extensions that can be turned on with flags. (sort of like how Clang or modern GCC accept C++11/14). There are minor changes from one version to another, so the language accepted by GHC today is often called GHC Haskell for disambiguation.
What GP is pointing out is that almost everyone using Haskell today writes in the language that GHC accepts today instead of sticking to some official standard.
If you want to learn Haskell, you don't need to worry about this at all. All reference materials and books teach modern GHC Haskell, which is also what is used by practically all Haskell users. As noted before, very few people use non-GHC compilers (usually it's just the people implementing or testing them) and you're very unlikely to encounter these implementations.
s/by GHC today/by the newest available GHC/
The language has tons of warts, but also an amazing community. Participating in that community taught me more about the practice of software engineering than my computer science degree.
method FAILGOAL($goal) { die "cannot find $goal near position {self.pos}" }
There is also a very helpful grammar tracer and debugger when developing a grammar: use Grammar::Tracer;
or use Grammar::Debugger;
https://github.com/jnthn/grammar-debuggerI like to imagine the Geth in Mass Effect run on a hybrid Perl6/Erlang platform.
$ perl6
To exit type 'exit' or '^D'
> 0.3 == 3e-1
False
I'm really really not sure how I feel about this. Does anyone know why the decision was made to make exponential notation return a float while decimal notation returns a rational?But decimal literal numbers (that is, not using the "e" notation) are best represented with rational values, in the sense that they don't need any approximation, so there really is no reason to use floating points for them.
edit: can't get double-asterisk to get through HN's markup :(
PS. I guess there is the "f" suffix as in C++. Well, for some reason that's not what Perl 6 decided. I, for one, think it's simple enough to use the "e" notation. If you want to write a Rat with a power of ten, you can always literally write say
1.2*10**3Establishing the right dynamic tension between those features (and the parser, and the VM, and the user's mental model) definitely makes language design one of the high arts in the computer world.
I'm not even saying this was the wrong decision; just that it causing a surprising deviation from how numeric literals behave in other languages.
But Perl's goal has always seemed to me to be less about creating a rigorous language with well defined types that map to the low level; and more about creating an expressive feeling (to humans) language, where fundamental differences in type were merely an implementation detail.
So I'm not surprised that it deviates from the norm in some interesting ways; I'm just not sure if I personally like the feel when I try to speak it :)
... but I might think to close to the metal to be the target audience.
C:\>python
Python 2.7.3 (default, May 4 2017, 16:44:54) [MSC v.1600
32 bit (Intel)] on win32
Type "help", "copyright", "credits" or "license" for more
information.
>>> 0.1+0.2==0.3
False [python]
>>> from decimal import Decimal
>>> Decimal("0.1") + Decimal("0.2") == Decimal("0.3")
True
or [python]
>>> 0.3 == 3e-1
True
What the perl6 "0.3 != 3e-1" comparison is really doing is [python]
>>> from decimal import Decimal
>>> Decimal("0.3") == 3e-1
False ArgumentList:
AssignmentExpression
ArgumentList , AssignmentExpression
I've never managed to do something like that with Perl 6's grammar. Basically every time I try to use recursion, it doesn't quite work unless I do some tricks.Maybe I ask too much, I don't know.
1. http://www.ecma-international.org/ecma-262/7.0/index.html#se...
grammar LOL {
rule TOP { <string> }
rule string { <[a..z]><string>? }
}
my @lulz = LOL.parse("abcdef");
say @lulz;
--- [「abcdef」
string => 「abcdef」
string => 「bcdef」
string => 「cdef」
string => 「def」
string => 「ef」
string => 「f」]Had you written it like this:
grammar {
rule TOP { <string> }
rule string { <[a..z]> | <string><[a..z]> }
}
as inspired by the rule mentioned above in the ECMAscript reference, it would fail.TL;DR Don't forget that 'rule' does <.ws> magic. This may or may not actually be part of the problem...
I have seen a similar sort of thing and only half understand the answer i got.. Check your implicit <.ws> and '|' LMF.
On a more serious note, your comment is pretty much what I've seen most people say. I will probably poke at it, at some point. I won't have a good reason to do so, except just to learn.
This review calls out the ease of FFI and a reasonable set of concurrency primitives so, in conjunction with the grammars and the data-mangling heritage, I conjecture that Perl6 might be good for IoT programming or SCADA. Of course it's a long trek from "might be good for" to "look at this compelling thing we built with it".
NOT a fun place to learn the many many random subtleties of $ vs @ variable dereferencing behavior or any of the other samples from the long list of surprising perl behaviors ...
[1] http://unattended.sourceforge.net/index.php
I have a lifelong preference for languages which don't actively prefer obfuscation over clarity... If your expectation for your developer's workflow is such that calling a directory 't' for 'tests' is deemed a useful abbreviation -- then I just simply do not want anything to do with that workflow.
> me want to try Perl6 out
Given your own understanding of desirabilities, I'd recommend you to steer away from this community in particular. :-)
> (Male|Female) (?:[Cc][Aa][Tt]|[Dd][Oo][Gg])
> would be translated to Perl 6 as:
> (Male || Female) ' ' [:i cat || dog]
Perl 5 does support this feature:
(Male|Female) ((?i)cat|dog)
or:
(Male|Female) (?i:cat|dog)
( Male | Female )
in Perl 6, as it tries to match both in parallel.The reason the version with || is in P5-to-p6 is that is how | works in Perl 5.
The idea is, at least as of a decade ago, it was straightforward to massage the then ECMA spec into Perl 6 code. The spec says "Go to step 4 ... 4. ..."? Well, "goto Step_4;... Step_4: ..." is code. And so on. Perl 6, at least as envisioned then, is a very nice language for language implementation. Implementation often goes like this: most of the language is easy, nicely compatible with the target; some of it is hard, requiring some struggle to impedance match; and tiny bit of it is impossible, or nearly so, or requires rewriting all the rest, because of "gotcha" conflicts with the target semantics. With a rich kitchen-sink target, almost everything becomes easier, and there's likely some hook you can use to get that impossible bit done. And when writing a compiler, (then, almost) general grammars, and multiple dispatch with type sets, is a nice place to be.
But the corollary is, when I don't see a compliant ES2017 transliterated to Perl 6, then I assume something is wrong - something isn't working yet - be it unfortunate spec, or "doesn't quite work yet at scale" implementation, or community.
http://www.wolfram.com/language/elementary-introduction/2nd-...
Essentially, WL uses # instead of * and requires a "&" to indicate when you're done defining your function. So the WL equivalent to Perl's
(1, 2, 3).map: * * 2
is Map[ #*2&, {1,2,3}]
(idiomatically you would use /@ as shorthand for Map[], so it would be) #*2&/@{1,2,3}
Except WL is far more general with what you can do with #, and you can do things like this: Outer[ #1 * #2 &, {1,2,3}, {4,5,6}]
which returns {{4, 5, 6}, {8, 10, 12}, {12, 15, 18}}
because #1 * #2 & defines a function of two arguments, and "Outer" feeds the lists into that function. You can even use ## to refer to all the arguments passed into that function (regardless of how many arguments there are), so Outer[## &, {1, 2, 3}, {4, 5, 6}]
returns {{1, 4, 1, 5, 1, 6}, {2, 4, 2, 5, 2, 6}, {3, 4, 3, 5, 3, 6}}
And you can even do stuff like: f[x_] := x^2
Outer[ #1[#2] &, {e, f, g}, {4, 5, 6}]
to get {{e[4], e[5], e[6]}, {16, 25, 36}, {g[4], g[5], g[6]}} sub ($a, $b) { $a + $b, [1, 2, 3], [4, 5, 6] }
Or even sub { $^a + $^b, [1, 2, 3], [4, 5, 6] }
So yes, the whatever-star isn't the most powerful thing imaginable, but it's good enough for very simple cases, and more powerful mechanism are available too.It's a very useful shorthand, although it becomes less necessary with properly first class functions and/or currying by default. E.g. I don't know if Haskell has it at all, but it doesn't need it.
Plus there's good old q's implicit argument naming of x, y, z when you didn't give the args explicit names:
each[{2*x}] 1 2 3
1 2 3 {x*y}/: 4 5 6At the bottom of the page, it says:
>Introduced in 1988 (1.0)
So WL had this syntax several years before the JVM was even invented.
「Man is amazing, but he is not a masterpiece」
Is different than the other two, it can be thought of as short for Q 「Man is amazing, but he is not a masterpiece」
While the others are short for qq “Man is amazing, but he is not a masterpiece”
Which is short for Q :qq “Man is amazing, but he is not a masterpiece”
Q :double “Man is amazing, but he is not a masterpiece”
Which is also short for Q :s :a :h :f :b :c “Man is amazing, but he is not a masterpiece”
Q :scalar :array :hash :function :backslash :closure “Man is amazing, but he is not a masterpiece”
For more information see the [Quoting Constructs](https://docs.perl6.org/language/quoting) documentation.---
There are modules for [debugging and tracing](https://github.com/jnthn/grammar-debugger) Grammars.
A regex is just a special kind of method by the way.
say Regex.^mro;
# ((Regex) (Method) (Routine) (Block) (Code) (Any) (Mu))
Which is part of the reason it doesn't have some of the niceties of parser generators built-in yet. The main reason is it got to a working state, then other features needed more designing at that point.---
# Perl 5
/(Male|Female) (?:[Cc][Aa][Tt]|[Dd][Oo][Gg])/
# is better written as
/(Male|Female) (?i:cat|dog)/
In Perl 6 you can turn on and off sigspace mode /:s (:!s Male | Female ) [:!s:i cat | dog ]/
# or fully spelled out
/:sigspace (:!sigspace Male | Female ) [:!sigspace :ignorecase cat | dog ]/
# or just use spaces more sparingly
/:s ( Male| Female) [:i cat| dog]/
Note that in this case it is more like /( Male | Female ) \s+ [:i cat | dog ]/
In other contexts it could be slightly different. Basically it ignores insignificant whitespace.Note that `( a || b )` is more like the Perl 5 behaviour, but `( a | b )` tries both in parallel with longest literal match.
Regular expressions are also the reason `:35minutes` is in the language by the way
say 'a ab abc' ~~ m:3rd/ \S+ /;
# 「abc」
Rather than make it a special syntax, it was generalized so it can be used everywhere.---
The asterix in a Term position turns into a Whatever, when that is part of an expression it turns into a positional parameter of a WhateverCode.
$deck.pick(*); # randomized deck of cards
$deck.pick(Whatever.new); # ditto
$dice.roll(*); # infinite list of random rolls of a die
$dice.roll(Whatever.new); # ditto
%a.sort( *.value ); # sort the Pairs by value (rather than key then value)
%a.sort( -> $_ { .value } ); # ditto
Note that the last asterix was part of an expression, while the others weren't. my &shuffle = *.pick(*); # only the first * represents a parameter to the lambda
# the other is an argument to the pick method
The main reason I think for its addition to the language is for indexing in an array @a[ * - 1 ];
Rather than make it a special syntax exclusively for index operations, it was made a general lambda creation syntax.I will agree that it takes some getting used to, but it is not intractable. WhateverCode lambdas should also only be used for very short code, as it can get difficult to understand in a hurry.
---
A `$_` inside of `{ }` creates a bare block lambda, basically this removes the specialness of Perl 5's `grep` and `map` keywords.
There is a similar feature of placeholder parameters `{ $^a <=> $^b }` to remove the specialness of Perl 5's `sort` keyword.
Another feature is pointy block, which removes the specialness of a `for` loop iterator value syntax.
# Perl 5 (this is the only place where this is valid)
for my $line (<>) { say $line }
# Perl 6
for lines() -> $line { say $line }
# not special
lines().map( -> $line { say $line } );
# really not special
if $a.method-call -> $result { say $result }
---There is more to NativeCall that you haven't discovered yet.
For example, you can directly declare an external C function as a method in a class, and expose it with a different name. (if the first parameter is the instance)
Also it doesn't matter what you put in the code block for a NativeCall sub, as long as it parses. That is why it doesn't matter if you put a Whatever star (asterix) there or a stub-code ellipsis `...` in it. (you can also leave it empty)
use NativeCall;
sub double ( --> size_t ) is native() is symbol<fork> { say 'Hello World' }
say double;
say '$*PID == ',$*PID;
# 1555
# 0
# $*PID == 1552
# $*PID == 1555
---Supplies can be pattern matched, just use `grep` on them as if they were a list. It in turn returns a Supply. You can also call `map`, `first`, `head`, and `tail` on them. Basically every List method is also on a Supply, along with special methods like `delayed`.
---
A lot about what you talked about with lists, and itemization is something that does take some time to get used to. It does get easier, but is always something you have to be cognizant of. Sort of like returning lists from subroutines are from Perl 5. It allows control over flattening that isn't available in Perl 5.
perl6 is doing the worse-is-better approach, by focusing on all thinkable features (sans structural matching and proper destructuring), but with a performance making it unsuitable for production.
there are some perl11 projects (5+6): cperl and rperl which solve both problems. suitable for production, proper performance and proper features.
So… Perl 6… has Perl-incompatible regular expressions?? Amazing.
Perl 6: The Ruby on Rails of command-line tools.
If you want to get hold of the latest stable full release you can use chocolatey which is kept up to date https://chocolatey.org/packages/rakudostar or use a .msi installer from the official Rakudo site.