GoPlus – The Go+ language for data science
github.com
github.com
I hope this succeeds and forces some changes.
Comprehensions become unreadable quite quickly - even simple ones have to be read inside-out, and complex (especially nested) ones are much worse.
Most modern languages seem to prefer functional-style methods like `map`, `filter` etc, which are more readable in all but the simplest cases.
Except in practice, the vast majority of comprehensions are simple.
I've come to dislike list comprehensions. Simple ones are ok but they do not handle incremental complexity well. Invariably it means code slowly becomes really unreadable as time goes by because nobody wants to rewrite the list comprehension when one more little tweak stuffed in there will do the job.
I much prefer object-functional chaining style ala Scala, Groovy, Ruby etc:
some_giant_list
.findAll { it.foo > 20 }
.groupBy { it.bar }
.countBy { it.value.size() > 5 }
They scale much better over time as new constructs get inserted as functional additions in the sequential pipeline.I feel the nice thing in Python is generator expressions, so for lots of code you can decrease the memory footprint and speed some code up through easy lazy evaluation.
I do appreciate the clarity of function chaining, and I think Rust made particularly good decisions around iteration in that way.
Function chaining ideally needs good ways to break iteration, skip and things, and JS isn't very feature rich beyond the basic map, reduce, forEach. It feels very incomplete compared to a lot of languages that offer that stuff. Hence the popularity of underscore and then lodash.
The more you get experience in Python, and the less you use specifically lists.
In fact, you can spot people getting confortable with the language when they start importing itertools, unless they are data scientists, of course.
Also I don't think it's possible in Python using await since await is a prefix keyword so in your example .groupBy is not a method on Future so it won't work.
I love ReactiveX as much as anyone but these issues cheese me off.
Rust has await as a postfix operator so this pattern works flawlessly when mixing async and sync code.
Looking at that code, I don't even know what the intermediary objects are, but given the names of the operations, they can't be all flat lists.
Besides, you may not put your list comprehension inline. It is often more readable to make it span on several lines, espacially if it's a complex one:
banned_ip = {
con.ip
for con in connexions
if con.ip in blacklist
or
con.type == "internal"
}The real challenge I've encountered is typically you can only use one expression. If I need to write a slightly more complex mapping I am forced to write a function, which I normally define just before the comprehension. Even though this works, it introduces boilerplate I'd rather not write.
I will admit this doesn't happen often, but it happens enough to bother me.
> This assumes some_giant_list is an object that has "findAll", that "findAll" returns and object that has "groupBy", which returns an object that has "countBy".
> Looking at that code, I don't even know what the intermediary objects are, but given the names of the operations, they can't be all flat lists.
In regards to this, I would say sure, but does that really matter. Most IDEs will give you enough inference to list the operations you can perform and the return types they give. If anything the "super collections" make a developers life far easier, I think Kotlin does a particularly good job of this with a vast set of extension functions. Have a scroll of this documentation https://kotlinlang.org/api/latest/jvm/stdlib/kotlin.collecti... mapIndexedNotNull is a great example
[x for y in z for x in zz if y in x]
Then you begin to wonder how great they really are.
for y in z:
for x in zz:
if y in x:
yield x
It’s extremely easy to sight read. If the lines become long, just split them up and it’s even more obvious. [x
for y in z
for x in zz
if y in x]
There’s nothing tricky about multiple loops and conditions in comprehensions in most languages that support them. Just go left to right and it mirrors outer to inner loops/conditionals.{x : y ∈ z, x ∈ zz | y ∈ x}
If you've read enough statistics/ML papers, this becomes second nature.
As someone who breathes that style of notation, there is excellent support in vim and emacs for custom display for readability. Conceal syntax in vim¹(builtin for rust/help/$few_others and external such as vim-cute-python²), and pretty-mode³ in emacs. I'm sure similar things exist in other editors too, but I don't use those ;)
You don't get the exact representation in your example with the things I've mentioned, but if you're used to a more "mathy"-style they really are quite nice. YM[and taste]MV.
1. http://vimdoc.sourceforge.net/htmldoc/syntax.html#conceal
2. https://github.com/ehamberg/vim-cute-python (moresymbols branch for er… more symbols)
In f(g(x)) you can directly access or patch f and g as top level names.
In x.g().f() you have to patch methods internal to other structures, and this can lead to problems if the patching should only happen in a certain local scope.
I encountered this recently with differences between pathlib.Path and os in Python.
Consider
def my_mkdir(pathname):
pathlib.Path(pathname).mkdir()
# or
os.mkdir(pathname)
From the pov of testing and decomposition, the second option is much nicer, because I can patch os.mkdir directly, and not patch pathlib.Path.mkdir (and also worry about controlling when instance creation happens to use the patch when I need it, but let other possible pathlib.Path objects my tests interacts with be constructed normally). mkdir is even a very simple example since pathlib.Path.mkdir is an instance method but only relies on the string data the instance has. Imagine how much harder if pathlib.Path.mkdir has complex interaction with the internal object structure or other instance methods.Obviously you _can_ solve it either way, but the fluent interface does nothing except require more code.
On this balance I think list comprehensions (or just fmap, which is all comprehensions are) are much, much better than chaining.
Also if you want to operate on data structures just operate on them with module functions.
I think pandas really messes up on this.
agg(groupby(df, cols), funcs)
is way better than df.groupby(cols).agg(funcs)Neugram [1] aimed to make Go a better scripting tool, but unfortunately it seems the project is dead. Although Go+ says its focus is on data science, I think it could fill this niche too.
By the way: does Go+ have shebang support?
What do you mean by a main-less script? And what would be "simple and fast" about it? If I want to try out something fast, I just do everything in a main.go and run it with "go run main.go", and that works well as a scripting language.
probably something that's common in "scripting" languages, where you don't have to wrap your code in main(), it just executes from the top:
myscript.py
-----------
print('hi')
vs myscript.go
-----------
func main() {
fmt.PrintLn('hi')
} frodo ~ $ cat t.go
//opt/go/bin/go run $0 $@ ; exit
package main
import( "fmt" )
func main() {
fmt.Printf("Hello, world\n" );
}
frodo ~ $ ./t.go
Hello, world
Otherwise there are a bunch of go-interpreters out there, which can be used for adhoc scripting. I'm sure that in some circumstances they can be useful, but I've only used them for providing extensions / scripting to host-applications, rather than trying to use them interactively.The workarounds mentioned by you and the other commenter’s linked article are sub-optional in that they require a system-wide modification, require wrapper scripts, or are non-standard hacks.
That's why Python+PyData has had so much success. There are packages to support data science, but the language itself can also be used to implement a system, so integration is rather seamless. That's not true for, say, R.
RStudio lists dozens of example clients here: https://rstudio.com/about/customer-stories/.
Use cases include collaborative model development, EDA tools, dashboarding, printed report generation (PDFs and HTML), public facing websites, etc.
If you’re trying to create ETL pipelines that integrate with BigQuery, Mongo, or whatever other database, I think it’s fair to say that the Python packages are generally better documented than their R counterparts.
For most other things, IMO it’s hard to really separate the two languages. Is standing up a Flask API really easier than in plumber?
For dashboarding, it’s is as quick (if not much quicker) to create a decent prototype with Shiny vs Plotly Dash or bokeh.
For simple linear and logistic model training, R’s built-in stats package has much more interpretable outputs vs sklearn, and directly inspired statsmodel. Wes McKinney has acknowledged that pandas draws heavily from R’s native dataframe. And so on and so on.
EDIT:
Also forgot to mention that with R packages like reticulate, you can also directly run Python code within an R environment now. So if there happens to be some Python package that doesn’t have an R equivalent, you can still work in R (though I’ve found the opposite situation to be far more common).
[1]: https://github.com/qiniu/goplus/graphs/contributors
Nice! Coming from Python, I miss being able to use the normal arithmetic operators on bigints in Go.
This solves some of the pain points but is still fatally flawed as any other Go ML tool in that it can’t accelerate due to the Go<->C FFI.
Until that issue is resolved Go simply won’t be broadly accepted in data science.
This isn't too much of a problem if your doing all supervised learning with batch operations because the speedup of a GPU over a bigger operation outweighs the FFI latency.
However, it's a problem that doesn't appear to have a solution due to Go's memory management choices, and will hamper it ever being used for accelerated computing problems. This is one of the reasons Rust moved to using ownership rules.
You can read a bit more at https://dave.cheney.net/2016/01/18/cgo-is-not-go
foo()? // for goplus
try(foo()) // the go proposal
https://github.com/golang/go/issues/32437Do any notebooks support a compiled language?
https://github.com/jupyter/jupyter/wiki/Jupyter-kernels will give you a more detailed list
We already have Python and Scala