Crystal in Production: Diploid
crystal-lang.org
crystal-lang.org
I've been using Crystal for more than 2 years have some projects in production. Can't wait for 1.0 :)
* immaturity in terms of the ecosystem (libs, frameworks, best practices etc) which will be solved in 1-2 years
* coming from Ruby I am wondering how powerful its metaprogramming capabilities are. In Ruby, using metaprogramming alcohol I can save tons of time by cutting off code and creating human friendly, clean interfaces that otherwise would require a lot of code. How powerful are macros compared to Ruby's meta capabilities where you can do pretty much anything?
On the bright side (in regards with elixir comparison) it has great type system that would save us from tons of bugs, it's more speedy and transition from Ruby would be slightly easier.
You can see a bit of what the macro system enables here: https://robots.thoughtbot.com/lucky-an-experimental-new-web-...
It generates a bunch of query methods and type specific querying. Without the macro system this would have been impossible.
The first thing that struck me about this post is that it is a good illustration for why dynamic-typing people and static-typing people often butt heads:
Most static typing we see is shallow like this, where the concern is to shove data into an unstructured object like a String, and the problem is something trivial.
In this case, String? solves the immediate problem of whether or not the value is really nil, but it fails to consider other aspects:
- Are you really intending to remove the country code from just US numbers? Or all North American ones? Quite possibly you're ok with stripping it from all North American ones rather than just US ones, but if not (e.g. because you later use the presence or absence of a country code to differentiate on billing), it does the wrong thing
- Does the code otherwise enforce a format where "+" can not appear elsewhere? (because someone e.g. decided to use it as a non-standard separator) Does it ensure nobody has input country codes without a "+"? (I've lost count of the number of times I've seen just "1" or (1)). Does it ensure nobody has used spaces? ("+ 1"). What about the number after that - it does nothing to prevent the returned local number from having consistent formatting.
Maybe this is all ok in the code you looked at and the data is guaranteed to only ever have "+1" and be nicely formatted when you strip it.
But it is a really bad way of selling static typing as a feature of the framework, as my first reaction is "but it gives me nothing, as there are all kinds of other checks I also need to do".
Which means either modelling the type of the data much more precisely - which I can do with or without static typing - and/or building a test suite that includes testing any places where data can get put into that table in the first place.
In both cases the kind of problems above tends to fall away with little to no effort.
I'm all for typing used for data validation, but an example showing what more precise modelling would look like would be a lot more convincing.
E.g. I detest Haskell syntax, but one of the strengths of static typing as Haskell developers tends to apply it, is that there tends to be a lot more focus on the power that comes from more precisely modelling the data.
I totally agree with your sentiment and you have a lot of great points. This post was not meant to get too deep into things right off the bat, so it left off a lot of this stuff.
In this case the phone number always has a country code with +1 because we have a validation that ensures it won't make it into the database without it. We also only accept US country codes in the validation.
I get your point though that just having a `String` doesn't really guarantee that those validations took place, luckily, LuckyRecord has the idea of custom types that can define their own casting and validation behavior. So we could have done this:
``` field phone : UnitedStatesPhoneNumber? ```
And it would validate that the string param is a valid US number before saving. It would also do that for queries and you could add whatever other type specific methods you want to it. so you could do
``` def formatted_fax_number phone.try { |number| number.without_country_code } end ```
But like I said, I think this is fitting for a whole separate post, rather than an intro style post :D
I will go more into depth about leveraging the type system with Lucky for even better data modeling in one of the Lucky guides
Even more so because the "try" syntax is exaggerated. This works fine:
fax_number.try(:gsub,"+1", "")
And with a new enough version of Ruby, this works too: fax_number&.gsub("+1","")
You still need to remember to do it of course, but using the "old" syntax makes the problem seem exaggerated.I had a chance to see static typing benefits in action and I will always agree they are a bit ahead of the dynamic typing.
That being said, I feel static typing is a bit overrated in web dev and outside of enterprise systems overall. F.ex. when I tried to quickly immerse myself in Elm, I figured that obsessiveness with static typing can make a dev's life very miserable and can basically require doing cartesian product structs for many scenarios -- especially if you have to work with data coming from several API providers and provide API yourself to users whose requirements are periodically changing.
I might be a bit in a fanboy mode here I admit, but the compromises that Erlang / Elixir do seem very adequate -- they do forgo some of the benefits of the static typing in return for a bit more productivity / less friction. Granted, if you are irresponsible then you can easily shoot yourself in the foot with them as well. No magic bullets.
(Conversely, if you apply some discipline -- which is still probably two orders of magnitude less than the discipline you need in C/Cpp -- then Erlang / Elixir's dynamic typing is almost like static typing.)
This isn't a blind hate towards static typing; having exposed myself for educational purposes for limited amounts of time to Elm, Pony, Crystal and just an hour of Haskell -- and having worked with Go professionally for several months -- I am not as impressed as I expected to be. I clearly see the benefits but again, I feel they are a bit overrated. A language with a reasonable compromise between type safety and programmer productivity seems to be my cup of tea. And I am not claiming that this "reasonable compromise" -- which is a very subjective term -- is an universal truth. Not at all.
Lastly, Elixir's macro system is not hugely powerful (at least compared to what I know about Clojure) but has served all semi-arcane tasks I tried -- and succeeded -- to achieve with it. But that's of course very specific to one's work and it too can't be claimed as an absolute truth.
Not willing to derail here. Your mention of Haskell triggered a few associations. Apologies if the comment is out of place / topic.
I like the String? sugar and believe there is probably some opportunity in Rails to address it. For example make a new NillableObject class that wraps the real object or delegates method_missing.
The idea would be to throw a warning, raise an error, etc. when accessing a nillable attribute. This could get obviously more robust than the below example, hook into read_attribute, leverage validations, log a warning instead of raise a runtime error, etc... but hopefully the point is made.
class NillableObject
def initialize(obj = nil);@obj = obj;end
def method_missing(name, *args, &block)
raise "I'm nillable, don't call methods against me directly"
end
def try(&block)
return if @obj.nil?
block.call(@obj)
end
end
ns = NillableObject.new
ns.gsub("f","g") # RuntimeError: I'm nillable, don't call methods against me directly
ns.try{ |o| o.gsub("f","g") } # => nil
s = NillableObject.new("food")
s.gsub("f","g") # RuntimeError: I'm nillable, don't call methods against me directly
s.try{ |o| o.gsub("f","g") } # => "good"My colleagues wrote a library that does something like this. You may want to check it out: https://github.com/thoughtbot/wrapped
Yeah the runtime thing is a disadvantage over type systems for sure. It might be possible with a NillablObject to do some kind of sanity checks at Rails initialization time. Also some other safety nets could help for example use of a ViewModel object to dictate some constraints on Views (type constraints and others).
I like Lucky’s table definition in the model class. A similar construct in AR might be useful in implementing some sanity checks. I often wish there was a way to get a table definition into AR while keeping the goodness of Migrations.
With promises and things like array.map being widely accepted in e.g. JS community, I'd hazard to say that mainstream industrial programming finally starts to embrace the use of monads. It would be great to embrace the most badly missing of them all, Option / Maybe, instead of the "billion dollar mistake" of null. ("Nullable" is a half-step in the right direction.)
It is the "happy path", represented compactly but without a way to forget and step on a nil.
obj&.is&.nil? can replace the use of try obj.try(:is).try(:nil?)
obj = nil
obj&.gsub("f","g") # => nil
obj = "food"
obj&.gsub("f","g") # => "good"
Unfortunately this is not likely second nature and will continue to bite us, especially where our expectations about AR models are not well thought out. A wrapping class for nillables might force safe access of AR attributes.Otherwise I just throw these adjectives away. I argue a new compiler/interpreter will always lose against the JVM, which has thousands of man hours of optimization built-in.
The JVM in turn will always lose against a clever memory-conscious lowlevel-implementation in Rust or C or assembler.
Please don‘t advertise speed without any studies or comparison to back up that claim.
We don't pretend to be a mature language with even so much as predictable performance characteristics but the "blazing fast" statement is there to indicate we're very much closer to C performance than even go is typically.
Fran Allen is the opinion that the adoption of C delayed the field of compiler optimizations research back to pre-history (Coders at Work).
It took 40 years of optimization research and clever use of UB defined in the standard, for C compilers to achieve the code quality generation they have nowadys.
Yes! The entire book is wonderful, but as a compiler writer myself, Fran's interview really stuck with me.
The relevant passage, for the curious:
———
Seibel: When do you think was the last time that you programmed?
Allen: Oh, it was quite a while ago. I kind of stopped when C came out. That was a big blow. We were making so much good progress on optimizations and transformations. We were getting rid of just one nice problem after another. When C came out, at one of the SIGPLAN compiler conferences, there was a debate between Steve Johnson from Bell Labs, who was supporting C, and one of our people, Bill Harrison, who was working on a project that I had at that time supporting automatic optimization.
The nubbin of the debate was Steve's defense of not having to build optimizers anymore because the programmer would take care of it. That it was really a programmer's issue. The motivation for the design of C was three problems they couldn't solve in the high-level languages: One of them was interrupt handling. Another was scheduling resources, taking over the machine and scheduling a process that was in the queue. And a third one was allocating memory. And you couldn't do that from a high-level language. So that was the excuse for C.
Seibel: Do you think C is a reasonable language if they had restricted its use to operating-system kernels?
Allen: Oh, yeah. That would have been fine. And, in fact, you need to have something like that, something where experts can really fine-tune without big bottlenecks because those are key problems to solve.
By 1960, we had a long list of amazing languages: Lisp, APL, Fortran, COBOL, Algol 60. These are higher-level than C. We have seriously regressed, since C developed. C has destroyed our ability to advance the state of the art in automatic optimization, automatic parallelization, automatic mapping of a high-level language to the machine. This is one of the reasons compilers are …[sic] basically not taught much anymore in the colleges and universities.
Seibel: Surely there are still courses on building a compiler?
Allen: Not in lots of schools. It's shocking. There are still conferences going on, and people doing good algorithms, good work, but the payoff for that is, in my opinion, quite minimal. Because languages like C totally overspecify the solution of problems. Those kinds of languages are what is destroying computer science as a study.
———
(pp. 501-502)
I recommend that any programmer who hasn't read this book give it a read. In fact, I think I might give it another read this week :)
It is a tremendously nice looking language - and I say this as a pythonesque guy who never wrote one line of Ruby. The feeling I get from community and projects is that it's a very up-and-coming thing, about to take off in a major way. There is just too much enthusiasm and too many things done right for Crystal not to earn some solid share within a very foreseeable future.
Thank you for this awesome feedback. It's really great to see that you like so much about Nim. Of course, I am more concerned about the negatives and I would love to do something about them.
In general it would be great if you could elaborate on your remarks. Some questions that I have: Are there other complexities in the language that you dislike? Is there something you think Nim would be better without?
Some specific remarks and questions:
> a lot of the implicit (or so it seems) stuff (I'm not comfortable exporting my procs by placing an asterisk somewhere, and even somewhere so illogical I can never remember where to put it)
An easy rule to remember this is "the asterisk goes after the identifier", for example:
proc life*() =
echo 42
proc identity*[T](x: T): T =
return x
var ident* = "hello"
type
World* = object
field*: int
> not at ease with the weird variable naming and casing conventionsWhat's weird about them?
> Don't even mention the arrays and the seqs...
Please elaborate on this as well.
By the way, you might have done this already but in the future please consider sharing this kind of feedback with myself or somebody else on Nim's team directly (via IRC/Twitter/Gitter/email). It's easy to miss something like this on HN. Please let me know if there is something I could do to make giving feedback easier for everyone, I understand that some people may not feel comfortable coming into Nim's IRC channel and giving criticism that way. It's very valuable to us though, and we don't want it going missing.
Personally, I still have symbol-overload PTSD from my Perl 4 days. When I see a flurry of that many symbols crunched together, followed by a dev saying "what's weird about them", I stay far far away from that language. Could just be me.
You will actually find that Nim typically avoids operators. For example, boolean expressions use `and`, `or` instead of `&&`, `||`. I personally like that a lot.
Having to write `public` before every procedure would, on the other hand, be a massive PITA :)
Anyway, let me try to address your points:
> Is there something you think Nim would be better without? Definitely, but these are entrenched features, and clearly not to be fiddled with at this stage: As hinted above, the sheer size of the thing, and the multiple ways there always seem to be to do whatever. Not necessarily a problem when writing code, but a real burden when trying to read it: "So what's a nice @ like you doing in a declaration like this?" "Oh, you're another way of declaring - or is it creating? - a seq. Which is sort of another way of declaring an array. Which has a couple of different implementations. Excuse me while I go and bang my head". Way too much of that stuff. Beginning with the alternative ways of seemingly each and every proc: Standalone and object-like, with a dot. I know this is meant to bridge the gap between us procedural folks and the object crowd. But it's still an obfuscation. It leads to disappointment, not least because at first look, Nim looks so deceptively clean and simple. In reality, I find the reading of Nim-code a lot harder than it needed to be. For my too heavy beginner project, I am porting some old Delphi code. And I mean old! I wrote it twenty years ago, and haven't looked at it since. Despite which, I have no trouble whatsoever understanding what it all means. But I do have trouble working out how best to represent my simple packed array[] of byte, what with the multitude of choices I have, though not a single one of them as concise and expressive as the Pascal, and all harder to comprehend after two weeks (as I discovered) than the original after two decades. I appreciate the enourmous dedication, knowledge, skill, and talent brought to Nim by Andreas and yourself, and whichever other core developers are out there, but damn!, I wish some useability freak had been consulted, and a lot of the heavyweight CS stuff had been kept solidly under lid and out of sight. There simply is too much unnessecary choice. It's overwhelming. And the overloading! This is a purely personal point of view; please ignore. But I don't like it, I never do. I see there may be use cases - that '+' might need to have a new meaning assigned in some new context, but to my uncomprehending eyes, that should be a hardwired lexer/compiler thing, never a userspace toy. And yes, I know just about every language does this nowadays. It's just that Nim-code seems to abound in it. And again, it hampers readability. A '+' should mean a '+', always, everywhere, no need to ever backtrack in order to understand.
>An easy rule to remember this is "the asterisk goes after the identifier" Yes, I sort of remember that now. But it's easily forgettable. And those poor characters - asterisk [how to escape in an HN comment?], @, #, !, etc. - don't carry any implicit meaning. Their use work against the ethos of Nim, at least as seen on the glittering surface. "public" or "export", or even just "exp" would do the trick, never mind the typing. If I wish to compress meaning down to arbitrary squiggles, I should go and code in C or Perl or some such nonsense.
>What's weird about them? [weird variable naming and casing conventions] That point has been beaten to death in threads all over HN, Reddit, and StackExchange. All the "this variable name equals that variable name, except when it doesn't, and give or take an underscore or two". It makes for uncertainty. The rules are sufficiently strange and arbitrary that I am never quite sure I remember them correctly. Like so many others, I should vastly, VASTLY prefer the way it's done almost everywhere else. With or without case sensitivity, but clear and unmistakable. It is not a huge problem, but it's a constant irritant.
>Please elaborate on this as well [arrays and seqs] I touched upon it above. I'm sure there is loads of clever, competent, unassailable reasoning behind the way things are done here. Me, a rather naive and superficial user, I just don't grasp it. Why do I need two very distinct types performing essentially the same function? And the array type in different flavors to boot, the fixed size and the open variety, each with its own incompatible set of behaviors? Again, like my other points, it's not a huge problem in itself, but all these double (and triple) options do add up, and [apparently] needlessly add to the burden of extra stuff I have to carry in my head. Clear and simple explanations of the need and rationale would go a long way towards making this acceptable. As it stands, it merely annoys.
>in the future please consider sharing this kind of feedback with myself or somebody else on Nim's team You are absolutely right, I am grossly remiss. I spew out my somewhat unqualified opinions in a place like this, but hang back from entering the lion's den - somehow thinking that as noob as I am, I really have no useful input to give. Right at the moment, I am insanely pressed for time and resources, but the instant I am rid of my present day-job, which I hate, I shall make an effort to participate much more actively in the community.
>Why do I need two very distinct types performing essentially the same function?
Arrays are very simple static blocks of memory and can be allocated on the stack, which is really fast, or at compile time. They can't ever be resized.
Sequences have overhead due to being pointers to pointers, however this means that they can be safely resized.
It sounds a bit like you're overwhelmed. Maybe the language needs a really paired down version of nim-by-example to explain these sort of things.
On the other hand I've read reports that sufficiently large codebases begin to hit machine memory limits for the compiler because they have to scan and construct every possible union type in the codebase. Not sure what there is to do about that other than bisecting the codebase into shared libraries over time.
And if it that were the worst part of meta programming, that might be ok, but it's not because the existence of meta programming puts ceiling on just how fast the Ruby runtime can be.
It's true that it shouldn't be abused and it most often is, but it does have its place.
It makes me sad for the reason you said, it’s most often abused.
I'll never ever touch upon these areas, so I still hold Ruby to a high regard. It simply is fun.
I'd be interested why a compiler would even need to do so in the first place --- constructing every possible one instead of every one it encounters as clearly needed right now.
Edit: OK I see it now, you can have stuff like Int32|String. Still though.. TS can "handle" that too?
Here is one of the creators of the language talking about the global type inference of Crystal:
def foo(bar)
bar + bar
end
bar1 : Int32
foo(bar1)
bar2 : String
foo(bar1)
bar2 : Int32 | String
foo(bar1)
The above code gives you three methods to generate code for: foo(Int32), foo(String), foo(Int32 | String).This isn't really a problem for small projects, but with large projects (60k+ loc), you end up with some methods such as puts, Array(T)#<<, etc with thousands of instantiations.
We have some edge-case semantics (we don't cast arguments to their restriction types, which we plan to remove before 1.0) which also prevent us from allowing only instantiating foo(Int32 | String) when we change the definition to `def foo(bar : Int32 | String)`. Hopefully in the future, we can allow the compiler to take advantage of method argument type restrictions to reduce the number of instantiated methods and speed up the compiler.
That doesn't mean the compiler will be super slow if you don't annotate your methods, it just means that the stdlib and shards' public API will start annotating it's method arguments more (which gives you better docs and error messages anyway).
For more info: https://github.com/crystal-lang/crystal/issues/4864
(I've got no real point here other than to say I had a strong feeling of deja vu while reading about Crystal.)
[1] https://en.wikipedia.org/wiki/Mirah_(programming_language)
There are plenty of people who dislike the JVM for any number of reasons (performance often being cited as one) and prefer to target specific platforms individually. Like many other things, it's a tradeoff, and people may arrive at different decisions depending on their priorities.
Java is not only OpenJDK.
It’s not much worse than a standard dynamically linked binary at that point. You just need a JVM on the system. Combined with the AOT compiler (as it stabilizes) makes Java applications more competitive with standard binaries.
All that aside, performance is great and has a massive ecosystem of mature libraries.
I’m not saying it’s always appropriate, but it’s not a bad platform.
The requirement for a JVM to be installed on a system isn't that dramatic in my opinion, it may change for others' use cases though (e.g. if you are shipping code to embedded systems etc., it's not a good option probably). Also, if you are OK with heavy app sizes, you can ship with OpenJDK although I acknowledge that it is as bad taste as shipping with electron.
[1]: JIT compilers can help with performance even for pretty static languages (e.g. Java is not so much dynamic than C++) because you have runtime statistics for that specific run and you can speculate using that. Although, with smarter branch predictors etc. in CPUs, some of the benefits of the JITs are slowly disappearing.
These are literally the top reasons why I love Go. To each their own?
That said, I'm eagerly waiting for true paralellism support since the use case we have in mind would greatly benefit from that. Some of the testing tools are also not quite as polished as Rspec (yet).
P.S: I'm the author of Kemal :)
The State of Crystal at v0.21 article (https://crystal-lang.org/2017/02/24/state-of-crystal-at-0.21...) stated that multithreading with work stealing was coming "soon" and that you already managed to run kemal in parallel. Can you share anything about the current state of that project?
You can check the wiki for more info https://github.com/crystal-lang/crystal/wiki/Threads-support
Was a cinch to swap over to it from Ruby. It is a little fussier with type definitions, but I guess that is to be understandable with a compiled system.
The third party ecosystem is still a little thin on the ground, or immature, and I hope it will grow. The Crystal community has also been really friendly and responsive. The couple of questions I have asked on Reddit or SO have been answered quickly and with lots of useful info.
Big if true.
Are there any benchmarks to back up this claim? There is some mention of experiments, but I'm not seeing any numbers or code.
If OOP is so important to them, why did they bother including Elixir in their list of possibilities?
Also, given their problem domain, I'm doubtful of the fit of OOP (vs functional). But regarding Crystal, it is nice to see a potential performant Ruby replacement.
Part of their problem domain was to easily port their code from Ruby, which would presumably be OOP.
There is another strategy using macros, which is much better. I wasn't familiar enough with macros when I first tried it, which is why it was a bit hacky.
Crystal's core team has plans for a DSL to do this, so there will surely be something nice in the future. It's definitely feasible, and the performance was decent.
If you're interested (just remember mine's broken with current Crystal), my work is https://github.com/phoffer/crystalized_ruby and a macro approach is demonstrated at https://github.com/spalladino/crystal-ruby_exts
Waiting for parallelism & Windows support +1
fh = File.open("logs1.txt")
fh.each_line {|x| puts x if (/\b\w{15}\b/.match x) }
fh.closeOn my machine, parsing a 374MB file took:
- 25.63 using Ruby
- 19.97 using Crystal (run w/o optimizations)
- 19.01 using Crystal (compiled w/o optimizations)
- 18.68 using Crystal (run w/ `--release` flag)
- 11.54 using Crystal (compiled w/ `--release` flag)