Scala at scale at Databricks
databricks.com
databricks.com
I work at a small company that has been using Scala for 7 years. Some of the prior employees enjoyed playing with advanced language features and writing libraries for the most general possible case even when that made it hard to understand how they were used for the 2 actual cases we needed to handle. Akka, Cats, and Shapeless were all over the place.
Those earlier employees have churned off to other places and I have successively simplified the code they wrote that is still useful, while encouraging the use of no more language power than necessary in new development. Hands-On Scala Programming is the book I give new hires as a language introduction that shows the sort of style to be preferred. It's much more like super-powered Python than like Haskell.
I have written C, Fortran, JavaScript, Python, and Scala for money. When I started on Scala I had never written Java nor used any JVM language. I have come to really appreciate the rich ecosystem of JVM libraries, the instrumentation and profiling tools I get, and many aspects of the Scala language and standard library. I love Scala's collections and miss their power and ease when I'm writing Python. (Which I still do for certain scripting tasks and for accessing Python-ecosystem libraries.)
This hurts me a lot
Scala the language is quite fine, but it attracts this kind of programmers. The damage they do to a codebase has to be undone by someone who both understands those advanced concepts and is not a FP fanatic.
Oxymoron.
It does. I call them "code scientists": engineers want to learn a language to build things, scientists want to build things to learn the language. I've interviewed >100 Scala engineers and if the panel thought they were the kind of person that just wanted to understand FP instead of solve problems, they were rejected.
I am now retired, but Scala is still my goto choice for engineering software. It checks all the boxes on concurrency, bulk I/O, ecosystem. In the hands of a sane and pragmatic practitioner, it's hugely productive.
I think this is a bit contradictory as it is phrased. Learning the language is essential to use it effectively. Otherwise one is better off with a language already known. And it is very possible to build things to learn the language and still keep it sane.
If we went back to let's say 2014 "when the cake pattern was the solution of all problems in world" I think I'd agree with your statement, even back then when I was a bit of a Scala fanatic.
Alternatively you can work in Go where braindead simplicity is the mandated norm.
1. How easy it is for an average programmer to understand the code.
2. How understandable the code actually is if the reader is fluent with that prior knowledge.
If one finds the latter is much better than code written in conventional techniques, but people are scared off by the former, it can get pretty frustrating, and "people are stupid" may be a conclusion drawn.
---
But my experience is that those techniques do not yield better code.
And "Educate yourself" made me barf.
I hate how “best tool for the job” really just means “what we can have a large hiring pool of people who’ve only ever learned one programming style and possibly only one language of” or “for legacy reasons this is the only choice”.
Yes I’m familiar with reality, I just hate it.
In other words, do you solve problems that actually require category theory or are you just navel gazing?
You don't need more than highschool maths to understand the category theory behind common typeclasses.
I know, and have been told, to have written hard to read code. In my case it's usually vectorized code in numpy or a C or cython extension for really hot code. But I always have a good reason, usually performance when it matters.
It's usually said that premature optimization is the root of all evil, but it's nothing compared against premature abstraction. Optimization at least is local in its nature while abstraction tends to expand all over a code base, and when the assumptions made for the abstraction no longer hold people is still forced to keep dancing for a music that no longer plays.
"Educate yourself" assumes that your interests are everyone's interests but this is a vast field and not everyone has interest nor time to learn about category theory. Most probably your coworkers are quite intelligent. If they are not interested in your idea it may be that it's not appropriate or maybe not correctly framed, or not mature.
Most computer science graduates are well schooled in algorithms, complexity, linear algebra etc. However, the vast majority of their time commercially will not be spent implementing some core algorithm, but instead gluing systems together and consuming/writing APIs. Any mathematics that helps them better accomplish this would be good to add to their toolbox.
Category theory gives you the precise language to talk about such stuff and helps to communicate and talk about more complex concepts and that's always a plus.
I'm not saying FP is bad, it is actually really powerful when it is in the hands of those who understand when and how to use it. The problem with the attitude that you demonstrate in your reply is also not limited to FP: in the era of OOO, there were those that spread the evangelism of design patterns to all places regardless if it made any sense; before that in the 90's there was a group of programmers that liked to generate code until no colleague understood anymore what was happening. And when you tried to tell them that abstractions don't come for free, but with a cost because they make code harder to understand, the answer for the last 25 years has always be the same, although rarely said straight: that it was only hard for -you- to understand, and not for the enlightened master himself.
I have seen more projects fail because of too much unneeded abstraction than by all other causes together. I even have seen companies go almost bankrupt because of projects engineered by lone wolfs where nobody understood the abstractions anymore except for the designer himself, and at some point he himself not really anymore either.
And the problem is that this mindset continues to be cultivated by CS books and conferences. Few people want to read a book that tells them that the secret of being productive as a team depends on mostly the culture in the team and the simplicity of the code, and not on the latest hyped framework, language or paradigm.
Agreed, and what did we get? dogmatic decrying of how OOP is completely useless and objectively bad, not too dissimilar from some of the comments on FP and Scala on here.
And just who were the ones pushing this? The people peddling FP.
The people in this thread are not pushing anything. They're just sharing their experiences with Scala and how unproductive it is to deal with FP zealots.
Ah! Just write a quine and be done with it!
> I have seen more projects fail because of too much unneeded abstraction than by all other causes together. I even have seen companies go almost bankrupt because of projects engineered by lone wolfs where nobody understood the abstractions anymore except for the designer himself, and at some point he himself not really anymore either.
"cake pattern" on one side and scalatest on the other.
As for
> I’m constantly surprised to see this type anti-intellectualism.
You hardly need any Category Theory to theory to understand type theory or mostly important Scala's type system. This the exact knowledge and allows you use it effectively. Yeah, the way he puts it makes those who like to understand the theory behind everything look bad.
I learnt the basics of FP a long time ago and understood the data-centric view of computation, the benefits of immutability etc - those things actually do help me write better code. The advanced FP, not so much. The nightmarish encoding of advanced FP shoe-horned into a language not built to support it that compiles into Java definitely not built to support it actively hurts. There are so many dirty macro shenanigans and hacks in the guts of the FP-purist libraries to force the language to implement things it just isn't particularly well suited to implement, which makes reading or debugging library code a nightmare. Like if you want Haskell, just use Haskell FFS and let us code reasonable Scala in peace.
There are so many parts of math that are much more useful in various fields - basic optimization in solving sudoku puzzles[1], linear programming in Z3, gradient descent in ML, calculus & vectors in graphics programming etc.
Instead of bringing the juniors up, we drag the more experienced guys down to their level.
The result is really shit basic code everywhere. or you end up using a language like go where there's basically only for, if and arrays and writing code in it is miserable and tedious.
There are plenty of companies where everyone has a very mature understanding of monads and where new joiners are coached to reach that level of understanding.
And there are plenty of places that aren't using source control, CI and unit tests.
- Engineers had written higher abstractions seemingly just because they could. When I audited how internal libraries were used across our services, calling applications weren't making use of the advanced abstractions. I'm talking about things like using Cats to abstract across AWS S3 error handling. Cool, except that it wasn't needed because we never actually encountered the exotic compositions of failures anticipated by the libraries.
- The abstractions written for our own business logic were worse than abstracting over S3. They were premature. Our business logic had to change frequently because the end user experience was still evolving rapidly. Changes that violated previous assumptions and their corresponding abstractions took longer than they should have and/or led to very awkward code.
- At least at the time, tooling had more problems with the "advanced" code. The IntelliJ IDEA Scala plugin could not yet show how implicits were used. It couldn't find senders sending to an Actor the way it can easily find plain callers of a function. You would need to manually force a "clean" in certain modules before code changes would compile as expected. IDEs would also fail to flag code that couldn't compile, and incorrectly flag code that would compile, at a higher rate compared to plainer Scala.
I'm still glad that I have access to Cats, Akka, and other advanced parts of the Scala ecosystem. They're still used in a few places where their value is greater than their cost. Even in the plain code, I'm still very glad I have pattern matching, immutability-by-default, rich collections, map, flatMap, filter, fold, scan, find, etc. I have no plans to transition our company off Scala internally. If I were starting a greenfield project with myself as the sole developer, I'd probably be using Scala for that too. But I prefer to write a bunch of simple repetitive code first, then develop abstractions after it's clear what the commonalities are.
“We like your solution but did you seriously just use Maven?”
What did they say exactly?
If I'm being honest here, a quite significant part of day-to-day backend JVM software development positions is dealing with the build tool. While I wouldn't outright reject a candidate for not being familiar with the latest standard toolchain (Gradle), I would certainly consider it as one of many factors - Gradle is pretty complicated, yes, but also quite powerful in what it is capable of achieving. It's not as eyebrow-raising as, say, if you sent in a coding challenge compiled using Ant (which would, in fact, be quite a fun discussion to have with a candidate) but it would certainly be something considered.
perhaps, because your work has over complicated it [1].
[1] If you are doing multi-distributions or multi-repo stuff perhaps it's warranted, but those complex build features better be providing serious value to the business to be worth losing 10% of the development workforce's time on build tooling.
Sure, sometimes progress is slower than you'd like, but that's the case in every large monorepo.
Just my counterpoint to all the negativity that people invariably feel compelled to pour out in any thread about Scala. Tons of devs use Scala productively without deep dives into monadic code golf, you know.
I agree that Scala works well in some skilled hands (see Li's libs), but the people that write the most terrible Scala programs actually know a lot about Scala. Scala has a unique way of encouraging genius programmers to write crazy code. People that know little Scala can just use it as a "better Python" language and actually write amazing Scala code.
I agree that Scala codebases that follow the Principle of Least Power can be amazing: https://www.lihaoyi.com/post/StrategicScalaStylePrincipleofL...
Scala programmers generally frown on programmers that want to use it as a better Python and I think it should be the opposite.
Scala is a popular platform to explore FP concepts without fully understanding their caveats or having enough respect for the rest of the team's ability and willingness to follow, I'll give you that. But that's only because it's a great platform to explore FP concepts in general!
You have to reign in the "geniuses", that's what engineering culture is for. I really don't like that label for people writing crazy code, btw. They shouldn't be on a pedestal just because they lack the empathy to write maintainable code.
The pythonic-Scala approach is certainly gaining steam in the ecosystem. The problem is, it's hard to see because it's boring and doesn't require many special tools – Scala has much of that built-in. So the popularity of this approach is not apparent if you just look at open source activity. But internally companies who work with Scala use it in very benign ways for the most part, in my experience at least.
Scala is almost as much joy to write data munging tasks as Python (unlike say C#, C++, Java, and I have to say Golang).
Chaining multiple maps and filters is so much more pleasurable than writing 4 nested loops with multiple ifs inside.
The better part is that you can reliably deploy Scala unlike Python.
We just had yet another discussion on the mess that Python deployment is: https://news.ycombinator.com/item?id=29439149
It's hard to make generalizations like this that are meaningful. There are a number of different Scala ecosystems, as you probably know. Spark is really a world of its own. Lightbend's libraries form their own ecosystem, as do Typelevel's and ZIO's.
My own opinion is that using algebraic data types (sealed traits, final case classes) with a sprinkling of type classes (implicits) where needed is pretty easy to work with, and that's the style that Typelevel is using. It is mostly inline with Li's "Least Power" post, though using IO pervasively instead of Future.
Is Typelevel's approach a better Python? I don't know. I don't really know what that means, and I'm not convinced that Python is a great model to follow. I see you posting along similar lines in many Scala threads on HN, so you seem quite invested in the language. I think it would help make your point clearer if wrote up a bit more of your thoughts, perhaps as a blog post.
The only Typelevel library I've used is frameless, which I blogged about here: https://mungingdata.com/apache-spark/frameless-typed-dataset.... Don't think the lib is really usable. Typelevel stuff is not what I'm referring to when I say better Python. I'll try to make another blog to clarify.
I think Scala is a better language than Python. But for some reason it just seems like Scala is full of mind bending code that I can't understand. Full disclosure: I am way less smart than lots of the Scala ppl I've worked with.
I find what most people mean when they say "simple" is really "familiar". When I first started teaching most programmers weren't familiar with first class functions and they were "complex". Now almost everyone has been exposed to them. Similarly I'm seeing more students are learning FP at university and thus find FP techniques "simple".
Yes, but, in my experience, the productivity gains are not that big. Strong developpers can and _will_ fare well in any technical environments, therefore choosing Scala over anything else won't give your team any advantage over a "simpler" technology. It's actually riskier to pick Scala, because you will increase your recruitment difficulty by a factor of 100-1000 (less developers available in the pool, even less when you only consider skilled developers, and huge barriers to teach and bring up to speed anyone that haven't worked in Scala before). What do you think will happen when the "good" developers are gone ? A rewrite in Python.
Save yourself, don't use Scala, pick Rust.
That's an hilariously bad advice if you think hiring competent developers and onboarding juniors is hard in Scala.
Most things companies do on the JVM, especially at scale, are going to be a lot harder in Rust and its ecosystem.
Absolutely not! I'm serious, and I stand by it. Rust ecosystem is like a better Python. You have to consider why someone is willing to move away from Python and into a foreign language: Developer supply, performance, safety, ecosystem, political reasons.
I can't cover all cases so I'm going to focus on the one that makes the most sense _from my personal point of view_.
I do AI Software Enginering and Devops when I'm not project managing. So Python is the only and primary choice for all things related to Data Science, dashboards, APIs, pipelines and transformations (on data).
Why and where would I choose Rust ?
In data pipelines, for performances and memory efficiency. This also ensure we have proper boundaries and documentation (APIs, file formats, schemas) between Data Science-related code and pipelines.
In Dashboards and APIs, for improving type safety, and handling more concurrency on background computations than python without resorting to other frameworks (Celery, Airflow, ...).
In data versioning and supporting infrastructure, for additional type safety and mostly making sure that _nobody_ who is not a software engineer won't cross component boundaries by accident. These components are often very delicate and essentials for a working platform. Changing them can have ripples into other parts of the project, and I only want trained engineers on these.
In the opposite, direction, I could pick Python over Rust when it's most appropriate: Restricted developper availability, improving coordination with Data Scientists, single container deployments, existing experience with framework or ecosystem...
The main blocker for not using Rust isn't because the language is new or more complicated, it's because Python offers more advantages to a situation. At the difference from Scala, I can pick _any_ skilled developper and put him from Python to Rust without any side effects (appart from the training part). This is because writing Rust code is much more consistent at all skill level than Scala. There no "flavor" of Rust. There no bickering about "coding style". The Rust tooling is best in class (Scala's SBT is actually worst-in-class, I believe). You can use the same IDE in Rust an Python (VSCode). This leads to a very similar coding experience between writing type-hinted-Python and Rust. This is not the case between Python and Scala.
While the criticism that Scala's tooling support might have been true some six or seven years ago, nowadays it is top notch.
There seems to be an experience in this thread that training in Scala is difficult. For anyone who is considering and wants another perspective - I have found that junior engineers tend to get up to speed more quickly in Scala than in other languages and that I can give them more freedom (because I know their programs won't compile if they don't line their types up and our use of the type system very immediately forces them to deal with the possibility that an operation can fail, e.g.). I can simply look at types and give them a nudge in the right direction (heyoo -fp either joke-). Yes they suffer a little in the beginning, but are usually shipping code within a few weeks then their trajectory really begins to rocket. It's very cool to watch them grow to building complex systems in a complex domain so quickly. As they grow the ease of refactoring means that it is easy to change their prior work as they gain more understanding, so for our management the penalty of allowing them to learn in the wild about something that's suboptimal is pretty low - they will come back and fix it. Of course what I'm saying above is really about FP and type systems and an implicit (heyoo -scala joke-) presence of thought about how our systems work top to bottom and how we train people. I think that is also required no matter what tools you use or what you're doing in any organization?
In what meaningful way can you give them more freedom? When i think about how i spend my time with junior folks, it tends to be on things like:
1. Do they understand enough of our problem domain that i can begin to describe a problem that needs to be solved?
2. Have they understood the problem i described? How confident am i that they're not going to implement a solution to the wrong problem?
We don't spend time on "hey, you passed the parameters in the wrong order here".
To put it another way based on my experience, what's more useful to a junior developer?
Option A: Hey, here's a test case, here's how to run it. I need you to change the behaviour of this component to do X instead.
Or
Option B: Hey, here's a component with well designed types. I need you to change the behaviour of this component to do X instead.
My experience has been that Option A - play with the component hands on, try things, run it, tweak the data supplied to it and observe the results - these encourage engagement. The developer will be excited even.
And for Option B, they'll be quite dejected, not that invested in the project. I can't even compile the stupid thing because of some cryptic compile error.
And in terms of morale, I view almost the opposite of what you're saying. Because types and the compiler enforce correctness in the program and compatibility within our systems I can have a relatively junior person thinking more about the domain and knowing that they can get something "real" into prod much more quickly with almost no risk to the org. They can move on to bigger blocks of work more quickly. Solving real, big problems and making impact gets engineers jazzed up. Regarding compilation, I find the "this won't compile" phase is almost non-existent in practice. People quickly adapt from "this thing doesn't compile [insert sad longing for python]" to "this thing won't compile - somewhere there's a misalignment with what I said I'm going to do (via types) and what I'm actually doing, [insert silent acknowledgement that they would have shipped an error in python and not known it]."
I'm not seeing where the gap is between A and B except that in option A it appears that there is no compiler involved, which per the above, means that I have to ask them to focus more on exercising mental discipline about coding conventions (boring, low value add work) and also I can't allow them to contribute as quickly because the risk they threw an exception somewhere they shouldn't have (e.g.) is higher.
What are the most common risks to an org (originating from its code) and how do types get you down to almost no risk?
When i think of common risks, most of them don't even originate inside the process's address space so the type checker is completely disconnected from the risks (e.g. hey - you misunderstood what we wanted here...). For the risks that do originate directly from the code, it's stuff like hey - your solution is quadratic time, sure your code compiles and your tests all pass but i'm telling you this won't work under prod data volumes.
I'm trying to steelman your argument and the best case i can come up with is an IO monad.
Here's a blog post I wrote on reading / writing JSON using one of Li's libs: https://mungingdata.com/scala/read-write-json/ There's now a completely sensible option for JSON in Scala.
But there are also 16 other Scala JSON lib options: https://stackoverflow.com/questions/8054018/what-json-librar...
It's so annoying working on a Scala project with dependencies that all use different JSON libs.
Even new, awesome projects like Delta Lake still use Scala dependencies like Scalatest and SBT that should be replaced with Li's libs in my opinion (utest & Mill).
Scala could be an amazing ecosystem if the community unified around Li's libs. It just doesn't seem like that's going to happen. Li wrote the "whats wrong with SBT" article back in 2017 and it seems like everyone is still using SBT.
I wrote an article on the maintenance challenges with Scala projects that trended on HN a while back: https://mungingdata.com/scala/maintenance-nightmare-upgrade/. Li's libs would solve a lot of these issues cause most of them are intentionally dependency free.
It was frustrating for me to see how Li basically singlehandedly fixed the Scala library ecosystem, but how none of the projects I actually used depended on his libs. I ended up leaving Scala after several years even though I like the language when it's used a certain way.
Scala is such a powerful programming language that has so much potential to be mainstream, but it's just so abused by "genius programmers".
The hilarious thing is/was that Scala tutorials/books would show writing JSON parser as one examples. Yeah Scala makes it easy to write a JSON parser but you still do not want to roll your own.
Case in point Odersky's Programming in Scala.
For example, its very fast and easy to prototype code with ADTs and scale that code to big codebase. I rarely have to run my code and if I do - it most of the time works as intended because of the type system. Its like magic. In other languages if you want to write code without bugs you are stuck with TDD or a lot of tests. With Scala usually few tests are enough, everything else encoded in its expressive type system.
I would say the only big negative I see is using Scala as Haskell. With tagless final, Kleisli and so on. If you do - it becomes very hard to onboard new people and I am not sure the positives outweigh the negatives.
Tagless final is a good way to enforce the principle of least power, which has many advantages when onboarding inexperienced developers. It can become unwieldy if abused, but that's true for any style.
And if it really bothers you, then you can use your concrete "super monad" everywhere. It's the only way with ZIO, but you can also do that with other effects.
Regarding Kleisli, it's such a broad abstraction that you'll find examples everywhere, whether explicit or not. In practice, even with heavy usage of libraries like fs2 or http4s, it's mostly an internal implementation detail, in most cases you'll only see type aliases.
There are pros and cons in every tech. Maintaining a large Scala code base is going to be much easier, less to write, less to maintain, easy to refactor.
Don’t use a library if you have no need for it but don’t dismiss it as fanatical unless you have actually compared what the same functionality in your fav tech looks like.
Choose the right tool for the job, don’t aim it at your foot. Get the job done and have a side project for “fun” coding.
Scala is a damn good option for targeting JVM. Don’t blame the tool.
1. Write the code like a better-Java/typesafe-Python (that is, imperative style).
2. Glance through the code. For collections, usually there's a more concise approach.
3. Read the code again. If you think what you did in #2 is making the code _too_ concise, you might need to think twice (or add comments).
Are you sure you're not confusing Kotlin with Scala?
> For example, Kotlin has null safety and it lets you write code using errors-as-values style "either" types - but it has two completely separate syntaxes for these things, and so it's impossible to interoperate or reuse code between those two approaches
And that is a problem how? Stick to one style.
> In practice Kotlin codebases still use magical incomprehensible reflection (Spring Boot)
https://github.com/spring-projects-experimental/spring-fu/tr...
> and magical compile-time manipulation (Kapt)
There's nothing magical about it.
Switching an API from "a result or nothing" to "a result or an error message" happens all the time, and switching in the other direction is only slightly less frequent. And of course most programs have some APIs where one is appropriate and some where the other is. So consistency is valuable.
> https://github.com/spring-projects-experimental/spring-fu/tr...
Still reflection-based.
> There's nothing magical about it.
It's magical to anyone thinking in the language - it breaks the rules of the language, so you can't reason about what it does.
Yikes. Just switch to Java already.
This is interesting as in taking a different approach to most other Scala shops I’ve experienced so far where die-hard monoid/functors fanatics take over and make the code base completely incomprehensible to outsiders / new hires. This seems to be a much more sensible approach.
They didnt want to explain / help other devs.
How to deprecate this at scale?
The extent to which i like scala is as a dsl for describing etl jobs with spark and i prefer the RDD api.
Engineers are consistently crippled by it, we dont even use shapeless or scalaz. Spec2 is a terrible dsl that offers a shitty vocabular to write unit tests and makes the common path incredible difficult compared to junit.
unless you like ridiculously long compilation times and having libs tied to what version of the compiler you're using (serious WTF)
It's true that some features of Scala have made it to Java, but I am not sure the gap between the two languages has shrunk much. For example while Java has added first-class functions and closures, it is still clumsy to attempt any serious functional programming approach with Java; Scala's type system is still way ahead; Scala compile to the JVM, JavaScript, and now native via LLVM.
> having libs tied to what version of the compiler you're using
It's never been that. However libraries used to be tied to the major version of Scala (for example 2.11, 2.12, 2.13) as major versions were not binary-compatible with each other.
Scala 3 addresses this thanks to the TASTY format which creates a strong foundation for binary compatibility. [1] You can use Scala 3 libraries with Scala 2.13 and, more importantly, Scala 2.13 libraries with Scala 3.
In short, Scala is not standing still.
[1] https://docs.scala-lang.org/scala3/guides/tasty-overview.htm...
Yes it is. Deprecate using it and migrate away from it. I hate that these shitty overengineered Cucumber wannabees are what "Scala" means to most people.
It's a great language IMO, but you need to make sure fancy frameworks pay their weight. IMO Shapeless and Scalaz-or-equivalent do - replacing reflection or AOP with plain old compiled code, even complex code, is a huge win. But SBT and the mess of Scala testing libraries are just huge complexity for marginal benefit.
I got interested in the Scala as a language and decided to pursue PhD in computer science focusing on compilers. Long story short, I am finishing my PhD on Monday :)
Don't be negative about technology. Be positive.
It's funny that if you tell Cat-purist that Spark is functional-style, they would laugh at your face :)
Got rid of 1000s of lines of Java and SQL for 100s of lines of scala, scalding and Spark.
The codebase size on the first migration was 10% of the original Java+SQL. overall, a team of 6 worked for 4 years to produce 10K lines of high quality Scala code for a very complex app and platform after the first transition from 10K lines of java+sql to 1K lines of scala.
Never experienced something like this before or after.
At the next place I wrote 400 lines of C++ mapreduce over 3 days, which could have been 40 lines of sql or 10 lines of scalding/spark.
Jarring transition.
Edit: Wait, did you imply that the SQL you replaced with Scala ultimately gets transpiled into SQL (spark or other dataframe engine) and "pushed down" to the database? Now you're just maintaining Scala and not a rats nest of SQL strings concatenated together? If so, that's pretty awesome.
Spark jobs were native RDDs before the SQL Spark era.
The SQL was HiveQL. Spark didn't really support SQL at the time and was native JVM Scala RDDs.
HiveQL with UDFs compiled to JVM bytecode typically ended up performing worse than scalding that corresponds to straight raw jvm map reduce jobs that are the real native interface of Hadoop.
The real performance drag from HQL came from developers writing inefficient map reduce jobs with extra joins to avoid writing Java UDFs because they were a huge PITA.
In scalding, developers wrote map reduce jobs with as few group bys as possible. Our hadoop infra team actually complimented us on how great the execution profile of our teams jobs were.
Also I agree with everyone that so many people write the worsttt Scala.
It’s like, let’s say it requires 2x more intelligence budget to read/debug than it is to write code. So if we all have more or less an intelligence budget of like 10, maybe some it’s 12 or others even 15-20, and you write crazy intelligent Scala, like at a 6, then you’ve already maxed out most peoples ability to read and debug the code, as we all more or less have a 10.
Scala let’s you write code that maxes out these budgets if you’re not careful. Sometimes, very rarely, it’s necessary, but most of the time there is a simpler, dumber way to do the same thing that is actually much better.
For most of my career in Scala, when I had conflict with other engineers, it was over this kind of stuff. They insisted on using cats or scalaZ or writing these crazy complicated, overly generic abstractions, and it’s like… I get what you’re doing and even I think it’s stupid and could just be done more directly and more attainably for other people.
While I kinda agree that you have to be careful with your abstractions, I see no harm getting familiar with some of the concepts that these libraries use.
Being able to just implement a Monoid and then use it to do nested aggregation in maps is pretty neat, for example.
It's a fairly accurate predictor of a company downturn after getting sunk in total complexity and ensuing certain employee churn.
The only reason for Scala adoption was Spark 10 years ago. Scala/Spark engineers were pulling the top buck. With pyspark and other cloud big data options on the scene, no one sane is using scala for this anymore.
Plus, the cost of teaching Scala to your new hires means they won't be productive for literally _months_
I agree that other big data tooling providers are eating into spark but I'm surprised people are choosing python over Scala for Spark.
And when I say Scala, I mean libraries like akka http, spray-json and slick. I like FP parts of Scala overall.
You would think that Scala 3 released in 2021 would have first class net/http and json that we can get rid of spray json. But no.
But I'm keen to learn more :) I always prefer compile time to run time errors.
https://circe.github.io/circe/codecs/auto-derivation.html is one popular library that does this - note that the `.asJson` will fail at compile time if you add a member to Greeting or Person whose type can't be serialized/deserialized (e.g. File or Lock or something) i.e. one that doesn't have an Encoder typeclass instance available.
can never convince the FP purists to let you use a non FP library.
I think at some stage specialists from some Akka pro shop had to be air dropped in to try fix it all, but it didn't help much.
It always made me laugh how well all the services written in Python and Go worked and how easy they were to extend, operate and deploy comparatively.
Is the business problem not suited to the actor model in the first place?
Cats et al are frequently abused - 90% of the time rather than using a complex tool to solve a complex problem you want to think about it a bit and simplify it. But when you need them they do things that you can't do any other way, and they're vastly better than the reflection/AOP/bytecode-manipulation that people use to solve the same problem in Java or Kotlin.
Can’t speak to category theory. Not familiar with it, but I have done my fair share of “reflection metaprogramming” that might get me me shot (dynamically constricting and using generic types through reflection gets pretty nasty).
I've inherited and built some actor based applications and libraries, both in Erlang and .NET (via Akka.NET). Usually these systems are:
- heavily Domain/event driven
- have 'funnel points' (example; bids on specific cars)
- have strong (but not extreme) requirements around consistent, fast response.
- be reasonably maintainable without a whole lot of gotchas
In all cases the actor model was the easiest way to solve the problems presented. That last point may be a bit contentious, but what every shop I've been at has found that while sometimes domain evolution results in extra boilerplate code to handle different versions of events, it is usually not the same sort of nightmare to do larger redesigns of the system as it ages and needs evolve.
> Every time I saw Akka used, replacing it with straightforward Futures-and-for/yield code was an improvement.
I'll admit I've over-used Akka.NET from time to time. I'd say the most overkill case is 'I need background workers running on timers with automatic recovery.'
How frequently do these need to change?
The problem with the library landscape is that one library will be promoted by the purists and another by the company with the money. Suddenly you have to learn both to understand Scala codebases.
If anything, you should be saying things like Shapeless provide the extra complexity.
Now that says something about the Scala ecosystem. Doesn't it?
I'm comparing this to how easily I can do the same task in nodejs and Go.
This isn't about the type system. This is about complicating things by making everything FP while getting nothing in return.
I think F# is doing much better in this category.
Scala can surprise you many unexpected ways. https://www.youtube.com/watch?v=v1wrWQcqLpo
But other than that, Scala is miles ahead. And if you want to do pure FP, then Scala is the only option.
Also about that video, almost all of the points mentioned are either not available in other languages like F# anyways, or they already fixed in the latest Scala version.
I don't think that Scala is especially good for "pure FP".
I also do not want to do "pure FP".
I meant when compared to F#. Of course there are other languages for FP as well.
> I also do not want to do "pure FP".
That's your choice!
Could you elaborate? Because that's not been my experience at all. I have found Scala to provide many tools to wrap around the idiosyncrasies of the underlying imperative that they are barely noticeable anymore.
In the video Paul Phillips elaborates on many subject much better than I could ever do.
I also found the .net ecosystem libraries better, even though there are fewer libraries available.
Some links that I found great on the subject:
https://www.reddit.com/r/scala/comments/4246qc/heres_why_sca...
https://techblog.bozho.net/i-dont-like-scala/
To be honest with you my social circle of software engineers gave up Scala long time ago and we do not even consider it to try it ever again. Most of the projects we do we simple use Java for JVM (since version 8 there are many useful features that you can use and have relatively terse code that is also easy to read) for the .net ecosystem we use C# of F# depending on how relaxed the engineering team and the stakeholders about it.
For data engineering (ETLs, glue code, etc.) Python works. Since Spark has a Python interface it is also not needed to use Scala there.
For frontend obviously Scala was probably never a serious option.
There is no use for Scala for me or for the companies I work for. The last time I tried Scala it was horrendous. SBT? Seriously? Binary incompatibility between minor versions? I could go on an on of the collection of anti-patterns I have encountered. Since I have limited time and there are a ton of other software environments that we can use and have much better experience why on Earth would I want to waste time on Scala ever again?
SBT is awful, don't use it.
"Minor versions" is overstating things - 2.11 was released in 2014, 2.12 was released in 2016, 2.13 in 2019, and 3.0 onwards are backwards compatible with 2.11. So that's one compatibility break every 2-3 years, which is better than most languages.
Scala is the only vaguely mainstream language I've found where you can do a whole enterprise-scale project in the language itself - no reflection, no AOP, no decorators, no macros, no monkeypatching, no nothing. So until I find another language where I can do custom context-like types (e.g. a type to represent "this must happen in a database transaction", and libraries already offer me functions like "take this list of database actions and combine them into a single database action" that work with my custom type) and "walk the object graph" style tasks (e.g. JSON serialization) in plain old type-safe code, I'm sticking with Scala.
Thanks for elaborating your point of view.
My company also though about doing so, because we have a few typescript projects that share a common module with business logic and everything lags behind when the common module was upgraded, but I don't know if I should sell them to switch over to a mono repo as well. My biggest pro argument would be the speed up in development and review time, because everything is in the same PR/MR and not split into each project Are there any huge downsides?
If you follow the design principles outlined in this doc (https://www.lihaoyi.com/post/StrategicScalaStylePrincipleofL...), Scala projects are really refactorable, which is nice.
But I've also worked on lots of large Scala codebases that are completely incomprehensible and painful to change. Scala allows for devs to make beautiful monoliths, but that's not usually what happens.
Still searching for pros/cons with current toolings etc, as most articles about monorepos are 3-4y old and therefore outdated in some points
Then there's the ecosystem. Sure, you have the whole of the Java ecosystem, but then there are also tons and tons of Scala libraries that are so well designed (Quick shout-out to Lihaoyi here; your libraries are amazing!) and work together so well is astounding.
I was very worried with the migration path to Scala 3 but that turned to be out pretty much painless. Scala-2.13-backward-compatiblity is amazing and together with scalafix the transition is as smooth as anyone could reasonably expect.
Anyone getting into Scala or worried about complexity with Scala should have a good look at https://www.handsonscala.com/
the reality is, we are in this spaghetti together, so pick your noodle and move on
its just computers, dont take it so seriously :)
This is the key to their success.
For Akka, do you really need it? You have to be building a pretty specific kind of system to need it over simpler constructs built into the language.