Haskell RecordDotSyntax language extension proposal (Accepted)
github.com
github.com
https://github.com/ghc-proposals/ghc-proposals/blob/master/p...
data Grade = A | B | C | D | E | F
data Quarter = Fall | Winter | Spring
data Status = Passed | Failed | Incomplete | Withdrawn
data Taken =
Taken { year :: Int
, term :: Quarter
}
data Class =
Class { hours :: Int
, units :: Int
, grade :: Grade
, result :: Status
, taken :: Taken
}
getResult :: Class -> Status
getResult c = c.result -- get
setResult :: Class -> Status -> Class
setResult c r = c{result = r} -- update
setYearTaken :: Class -> Int -> Class
setYearTaken c y = c{taken.year = y} -- nested update
getResults :: [Class] -> [Status]
getResults = map (.result) -- selector
getTerms :: [Class] -> [Quarter]
getTerms = map (.taken.term) -- nested selectorWhen/ how can I use this? Do I have to wait for a new GHC release?
I agree that it's very nice, and a big improvement!
I am so elated to see this change, even if it is 15 years overdo. Nested selectors in particular are beautiful.
Maybe it's time again to do me some haskell for great good.
E.g. if both data types were to contain an "id" field then "taken.id" and "class.id" wouldn't result in a compilation error as it does now.
That code is copy-pasted by the way, from the proposal.
So what will this mean now? Will function composition require a space between the dot, or how will a “record dot” be disambiguated?
.g is a field selector.
f.g is field selection of a record.
f . g is function composition.
f. g is function composition.
While I wouldn't write f. g, if I see it, it's not hugging the field name so it's composition. I don't think I'd have trouble reading it.
But... it will probably trip up newbies and people with odd spacing styles. Hopefully GHC can give a useful error like "In the field selection expression f.g, g is a function defined on XX. Did you mean f . g?"
> :t (. id)
> (. id) :: (b -> c) -> b -> c
(:t prints the type of an expression, and id is the "identity function" that's b -> b here.)(EDIT: Maybe you know the above and you're asking if it's being changed in the proposal. It doesn't seem so to me.)
import Data.Set (fromList)
import qualified Data.Set as Set
data Set a = Set (Set.Set a)
f :: Ord a => [a] -> Set a
f = Set . fromList
g :: Ord a => [a] -> Set a
g = Set .fromList
h :: Ord a => [a] -> Set a
h = Set. fromList
i :: Ord a => [a] -> Set.Set a
i = Set.fromList
In other words, the . is interpreted as a "module dot" if and only if there is no whitespace either preceding or following the dot.It looks like the "record dot" follows the same parsing rules as the "module dot": the presence of whitespace turns the . into function composition.
Both pages link to each other, but the discussion is unlikely to be of very much note unless one has already read the proposal being discussed.
1. https://github.com/ghc-proposals/ghc-proposals/blob/master/p...
Records were my number one gripe in the language that otherwise does so many advanced things so uniquely well.
The whitespace requirements, though, are going to lead to some confusing error messages for newbies who want to type myRecord . someField
data Student = Student { name :: Text, school :: School }
we end up with two functions 'name' and 'school' in scope that extract the values of their respective fields: name :: Student -> Text
school :: Student -> School
This means it is common for field names to be composed with functions. I write code like this all the time: schools = map (schoolId . school) students
This is compounded by the fact that a lot of newtypes use record syntax to define functions in this style that are even named like functions: newtype Reader r a = Reader { runReader :: r -> a }
Importing runReader from a module, you might not even know that it's implemented as a record field! This is pretty handy since you can change the implementation of runReader from being a record field to a normal function without breaking most of your callers. (In fact, I think this might have happened with runReader being implemented in terms of runReaderT.)This kind of pattern means that people write lots of code that uses record fields as functions, sometimes without even realizing that they are using a record field!
The practical upshot is that myRecord . someField is used with . as function composition all the time, so any proposal that cares about backwards compatibility can't change the meaning of . in that context.
The change appears to just be syntactic sugar and so would not conflate anything.
Rereading now, it seems you're asking how interpreting a . b as record access only if b is a record field in scope would break existing code.
The new record system is implemented on top of a HasField typeclass under the hood. Thanks to the way Haskell's typeclasses work, this means that the instances that define what fields a record has will be in scope when you import the module containing that record—directly or indirectly—even if you don't import the record itself.
This means that even if we do not have a record with a field b in scope, we might still have the instance for some other record in scope—leading to ambiguity in previously unambiguous code. Moreover, even if a . b is unambiguous right now, a record with a field called b might be added somewhere deep in the codebase and cause unexpected ambiguities in the future.
I suppose you could say that a . b is composition whenever a function b is in scope, even if there is also a record with a field called b. That rule would not break existing code, but it also seems inconsistent and confusing—a.b and a . b are sometimes the same and sometimes different, depending entirely on whether b is in scope.
Making the meaning of . purely lexical rather than depending on what's in scope seems a lot more consistent and easier to follow. It's already the case with Haskell anyway, since . can mean function composition or a qualified name depending on the spaces around it. This isn't an ideal situation, but I don't think we can aim for "ideal" in a living language that's 30 years old.
(Hopefully britanny will get an update to handle this correctly!)
They could, In Theory, rename the composition operator selectively when this language extension is enabled. That'd mean Prelude has to change what operators it exports depending on what extensions are enabled, which would be horribly magical.
foo()
.filter(...)
.map(...)
then you have to allow foo() .filter(...) .map(...)
as well, as a matter of principle. (1) foo()
.bar()
is equivalent to (2) foo() .bar()
there is a difference between (3) foo()
bar()
and (4) foo() bar()
(the (3) results in valid code, whereas (4) would not).But trying to argue that (2) should be syntactically invalid would be a very hard ask indeed.
>>> 2020 . to_bytes(4, 'little')
b'\xe4\x07\x00\x00'
>>> 2020.to_bytes(4, 'little')
File "<stdin>", line 1
2020.to_bytes(4, 'little')
^
SyntaxError: invalid syntax
JavaScript also requires the spaces: > 2020 . toFixed(2)
'2020.00'
> 2020.toFixed(2)
2020.toFixed(2)
^^^^^
Uncaught SyntaxError: Invalid or unexpected tokenBut for decimals there is no ambiguity:
>>> 2020.50.is_integer()
FalseSo basically, you could easily write a parser that allowed '0.toString()', but you'd either have to piece numeric literals together in the parser or add nasty hacks to the lexer.
This is actually not hacky. It's just a rule that the "." cannot be followed by [ \t]∗\w, which is a simple negative lookahead assertion. Replace \w with whatever you use at the start of identifiers.
It is extremely common for languages to have corner cases like this in the lexer to make the language more usable. For example, consider the rules in JavaScript or Go concerning where you can put line breaks. Or the rules for JavaScript concerning regular expression literals, which must be disambiguated from division.
> So basically, you could easily write a parser that allowed '0.toString()', but you'd either have to piece numeric literals together in the parser or add nasty hacks to the lexer.
This is factually incorrect. As I explained, you would only need one character of lookahead. There is no need to parse "0. toString()" successfully. If you wanted to parse "0. toString()" correctly, you could use unbounded lookahead, which is fairly simple in practice (speaking as a sometimes parser writer). I don't get why you say it is hacky, this is all just a bunch of regular expression stuff (in the traditional sense of "regular").
Right, which is what I said. If you agree that unbounded lookahead is required then we don't really disagree, except on the somewhat subjective question of how 'hacky' that is.
If I understand correctly, you suggest that unbounded lookahead could be avoid by allowing '0.toString()' but not '0. toString()', while still allowing both '(0).toString()' and '(0). toString()' and both 'foo.bar' and 'foo. bar'. That would produce highly counterintuitive results in some instances:
Parsed as one expression:
{}.
foo
Parsed as two statements:
0.
toString()
But again, it is really a subjective judgment. Obviously you could modify Javascript in this way, and on that point there is no disagreement.But I just noticed C# support this, so it is not the same for all languages. Java doesn't, but then again you cant call methods on numbers in Java anyway.
I even use it sometimes, like this:
local field = require "some/module" . field
When I want exactly one field from a module which returns a table. I think the spaces look nicer next to the string.
EDIT: I mean, lenses and RecordDotSyntax both provide a way to access nested fields in records, so the new syntax would take away some of the reasons to use lenses. But are there any other relations among the two, or do lenses and RecordDotSyntax just provide two different solutions to the same problem?
- Mutation of nested records,
- Iteration of collections,
- Folding of collections,
- Viewing discriminated cases that may or may not exist,
- Building a structure around a value, and so on and so forth.
[1] https://github.com/ghc-proposals/ghc-proposals/pull/282#issu...
I would say in my 5 years of production Haskell, I've seen that sort of thing deployed in production _maybe_ once, and I've run into it during development once or twice.
One time it was in integration testing, and a simple heap profile made it clear it was in a library dependency. And a quick look at the library source was enough to diagnose it. It was like an hour or two tops from "we have a space leak somewhere" to "here's a PR to a third-party library fixing the leak."
A bit of a problem when you first start using Haskell seriously is that the IO mechanisms built into the prelude are very slow once you get to even only a few MiB.
That's easily remedied by switching to eg ByteStrings, but it's still a bit annoying that what the language presents as the 'default' is such a toy implementation.
My only worry is . being function composition operator so it might not be crazy amazing for readability only you get
a . b . c . d.e
So maybe it should have been something like "->".
(.a.b.c) . (.d.e.f) = (.d.e.f.a.b.c)
If you use -> both for the member function and for reverse function composition you'd get:
(->a->b->c) -> (->d->e->f) = (->a->b->c->d->e->f)
which is somehow nicer, especially because accessing a member and composing functions become the same thing when you view a value as the unique function sending the initial object to that value.
But unfortunately we've spent centuries composing functions the wrong way round, so now we're forced to deal with the consequences.
Given functions f, g , h,
h >>> g >>> f is equivalent to f . g . h
Its from Control.Category, so it generalizes to other things, but for functions specifically its left to right composition.
( theres also (<<<), which in this case is identical to (.) )
( Also present in Control.Arrow . Why? I dont know this topic well enough to explain. My understanding is limited to what these operators mean in the specific of functions. Functions(->) are a type as well, so they can be instances of typeclasses. for example
instance Functor ((->) r) where
fmap = (.)
)The problem is that "->" has an even more special meaning than "." in Haskell, since "->" separates the arguments of an unnamed function from its definition.
So, for example
addOne = \x -> x + 1
and addOne x = x + 1
are equivalent.This means that "->" is about as special as "=". Which leads me to think it would require some very substantial changes to the parser, and perhaps even defeat the purpose of increased readability. For example, how readable is the following?
getY = \x -> x -> y
(which reads as "getY = \x -> x.y" in the current proposal) (<.) :: (b -> c) -> (a -> b) -> a -> c
and (.>) :: (a -> b) -> (b -> c) -> a -> c e{lbl = val}
Though this has been a part of haskell for a while it seems:> Note: e{lbl = val} is the syntax of a standard H98 record update
FSharp, OCaml and Elm have not got great solution here either:
{ e with lbl = value } // ocaml/fsharp
{ e | lbl = value }
Do any functional languages have anything as nice as: e.lbl = value new_e = e.lbl = value
at which point "nice" is not how I would describe it. e & lbl .~ value
It's not an in-place update like `e.lbl = value`. a.b.c.d
but a write is: let a = {a with b = { a.b with c = { a.b.c with d = 5 } } } }
which is much much more verbose than the OO/imperative: a.b.c.d = 5 a & b.c.d .~ 5
To construct a new version of `a` where those nested fields have the new value 5.Or for an in-place update in a context where that is allowed,
a.b.c.d .= 5
The lens concept is incredibly nice.syntax, and Ramda has
e2 = assoc(lbl, value, e)
which is curried and data-last, so you can keep that setter around for future instances of e. Usually not too bad to translate Ramda into your language of choice, and a valuable exercise. I did it with Python and it helped a TON
Once JS gets immutables, web devs will have more tools for FP, and we can look forward to that expanding the community of functional programmers a lot: (https://www.infoworld.com/article/3569118/ecma-proposal-woul...)
Python’s lack of an immutable map is annoying but there are nice pip packages for it
e = e.lbl <- value
(I didn't reuse the '=' because two '=' like that looks weird.The main thing I want to address is that a.b.c.d is so much work to write to.
%{existing_map_or_struct | some_key_1: :some_val, some_key_2: :some_val}
Then there is also `put_in/2`, `put_in/3`, `update_in/2`, `update_in/3`, these are nice because they work in pipelines some_val
|> put_in([:foo, :bar, :baz], 3)
some_val
|> update_in([:foo, :bar, :baz], fn x -> x + 1 end)
The macro forms are cool too put_in(foo.bar.baz, 2)
update_in(foo[:some]["nested"][:path], & &1 + 1)
The cool thing about the `/2` forms of `put_in` and `update_in` is that they are macros they rewrite the ast to return the entire object, and not just the changed value. iex(1)> foo = %{bar: %{baz: 2}}
%{bar: %{baz: 2}}
iex(2)> update_in(foo.bar.baz, & &1 + 1)
%{bar: %{baz: 3}}
There's even more powerful macros like `get_and_update_in` which can use "selectors" to extend their functionality (https://hexdocs.pm/elixir/master/Kernel.html#get_and_update_...)https://hexdocs.pm/elixir/master/Kernel.html#put_in/2
https://hexdocs.pm/elixir/master/Kernel.html#put_in/3
let a = {a with b = { a.b with c = { a.b.c with d = 5 } } } } e{lbl = val, lbl2 = val2}I agree that's quite nice, however updating nested values is very painful.
I can't find a subscribe button in gutlab but I will set a reminder
The old classic languages just create so-called selectors in the global namespace, like defstruct macro does. This is not the best possible solution.
However, the . is used for function composition, so the better alternative could be the -> syntax, reflecting that in Haskell everything is a pointer.