HNHacker News
TopNewBestAskShowJobs

cbkeller

1,804 karma · joined October 22, 2018

geologist - brenhinkeller.github.io
submissionscomments
cbkeller··on Julia 1.6 Highlights
See also Lyndon’s blog post [1] about what all has changed since 1.0, for anyone who’s been away for a while.

[1] https://www.oxinabox.net/2021/02/13/Julia-1.6-what-has-chang...

cbkeller··on Release v1.6.0 of Julia
Parallel precompilation!

See also: https://www.oxinabox.net/2021/02/13/Julia-1.6-what-has-chang...

cbkeller··on Microsoft in talks to buy Discord for more than $10B
Potentially yes: the supernodes were just other "regular users who had sufficient bandwidth" which could be quite efficient if you were all in, say a remote region with good local bandwidth but little connection back to the rest of the world. Obviously though, this P2P structure would make interception of message content by Skype corporate rather implausible! There's a pretty good figure and explanation in the ars article above.
cbkeller··on Crystal 1.0 – What to expect
On one hand you're absolutely correct in that Julia is dispatch-oriented, and doesn't have ton in common with what is commonly considered OOP today.

However, as far as I understand it there is also an argument to be made that multiple dispatch is at least as close to Alan Kay’s original intent for OO (e.g. [1]) than is the currently-ubiquitous class-based version of OO that leads to things like `RequestProcessorFactoryFactory` s with `RequestProcessorFactoryFactory.getRequestProcessorFactory(Class)` methods.

[1] https://medium.com/javascript-scene/the-forgotten-history-of...

cbkeller··on Matrix multiplication inches closer to mythic goal
I was annoyed by that at first, but eventually came to appreciate how Julia's semantics for when broadcasting is or is not required for operations on scalars, vectors, and matrices map to the underlying math. I.e., for an N x N matrix `A`

  5 * A
does not require broadcasting because multiplication of a matrix by a scalar is mathematically well-defined, whereas

  5 + A
(assuming you want element-wise addition) does require broadcasting, i.e.,

  5 .+ A
because addition of a vector and a matrix is not mathematically well-defined.
cbkeller··on Microsoft in talks to buy Discord for more than $10B
Well that's probably at least in part because a major motivation for that acquisition was not to actually improve the underlying technology. To quote verbatim one of my favorite hn posts of all time (https://news.ycombinator.com/item?id=18411779):

  A series of headlines:

  2009 - NSA offering 'billions' for Skype eavesdrop solution [1]
  2011 - Microsoft buys Skype for $8.5 billion. Why, exactly? [2]
  2012 - Skype replaces P2P supernodes with Linux boxes hosted by Microsoft [3]
  2013 - Microsoft handed the NSA access to encrypted messages. Subhead: Skype worked to enable Prism collection of video calls [4]
  
  1. https://www.theregister.co.uk/2009/02/12/nsa_offers_billions_for_skype_pwnage/
  2. https://www.wired.com/2011/05/microsoft-buys-skype-2/
  3. https://arstechnica.com/information-technology/2012/05/skype-replaces-p2p-supernodes-with-linux-boxes-hosted-by-microsoft/
  4. https://www.theguardian.com/world/2013/jul/11/microsoft-nsa-collaboration-user-data
cbkeller··on Opticsim.jl: Optical Simulation Software
They do what they can, including things like [1] recently discussed here, but they're not exactly a Fortune-500 company yet.

[1] https://news.ycombinator.com/item?id=26425659

cbkeller··on Opticsim.jl: Optical Simulation Software
I think people often underestimate (or just plain don't know about) the degree to which a multiple-dispatch-based programming language like Julia effectively implies its whole own dispatch-oriented programming paradigm, with both some amazing advantages (composability [1], and an IMO excellent balance of speed and interactivity when combined with JAOT compilation), but also some entirely new pitfalls to watch out for (particularly, type-instability [2,3]). Meanwhile, some habits and code patterns that may be seen as "best practices" in Python, Matlab can be detrimental and lead to excess allocations in Julia [4], so it may almost be easier to switch to Julia (and get good performance from day 1) if you are coming from a language like C where you are used to thinking about allocations, in-place methods, and loops being fast.

Things are definitely stabilizing a bit post-1.0, but it's still a young language, so it'll take a while for documentation to fully catch up; in the meanwhile, the best option in my experience has been to lurk the various chat forums (slack/zulip/etc. [5]) and pick up best-practices from the folks on the cutting edge by osmosis.

[1] https://www.youtube.com/watch?v=kc9HwsxE1OY

[2] https://www.johnmyleswhite.com/notebook/2013/12/06/writing-t...

[3] https://docs.julialang.org/en/v1.5/manual/performance-tips/#...

[4] https://github.com/brenhinkeller/JuliaAdviceForMatlabProgram...

[5] https://julialang.org/community/#official_channels

cbkeller··on Data Science in Julia for Hackers
Little-known fact: there is a Julia interpreter, if anyone really wants to go that way: https://github.com/JuliaDebug/JuliaInterpreter.jl
cbkeller··on Data Science in Julia for Hackers
I think it really depends on what subset of data science you have in mind. For me it solved my personal version of the "two-language" problem, but YMMV. Longer-term I think the most interesting thing about the Julia ecosystem is the composability that comes from dispatch-oriented programming [1]

[1] https://www.youtube.com/watch?v=kc9HwsxE1OY

cbkeller··on Data Science in Julia for Hackers
I don't mind Julia's Unicode support, but definitely share the gripe about how the Unicode standard doesn't yet support a complete a-z set of sub/superscripts!
cbkeller··on What is the most underestimated programming language?
While normal Julia syntax is indeed arguably not quite homoiconic in the same way that Lisp s-expressions are (a good discussion of the issue by Stefan Karpinski here: [1]), it turns out there is a little-known built-in s-expression syntax for Julia (which can be more-or-less enabled as a REPL mode in only a handful of lines [2]).

[1] https://stackoverflow.com/questions/31733766/in-what-sense-a...

[2] https://gist.github.com/brenhinkeller/44051118c2f9d18b26dc76...

cbkeller··on Data Science in Julia for Hackers
Adding Plots (or for that matter, Pluto!) to your sysimg with PackageCompiler.jl is a potentially significant QoL improvement on that front. Parallel auto-precompile in 1.6 is also awesome.
cbkeller··on Data Science in Julia for Hackers
Can you give a code example in R of what you would like to do in Julia?

There are quite a lot of different ways of mapping in Julia, including functions like `map`, `mapreduce`, `replace`, and various syntaxes for broadcasting and array comprehensions.

For maximum SIMD performance there are also things like `vmap` and `vmapreduce` from LoopVectorization.jl.

cbkeller··on Why are most climate models in Fortran?
Well, I'd say I'm generally at the shallow end of the HPC pool, and also in a field that's relatively new to HPC in general so there aren't a ton of established community codes to draw upon. Consequently, I'm willing to make a few mild performance tradeoffs if it means my grad students and I can iterate faster.

If you wanted to use Julia at petascale or above (like the Celeste folks), then you'd probably want to be doing that in a case where you see fundamental algorithmic improvements you could readily make over the current SOTA with a higher level, dispatch-oriented language - or else a case where you really need, say, certain types of AD or the DiffEq + ML capabilities of the SciML ecosystem (the latter of which very much depends on the level of composability that follows from the dispatch-oriented programming paradigm).

In general, my two cents on what it takes to get "good enough" performance from Julia for my sort of HPC are that you (1) embrace the dispatch-oriented paradigm and take type-stability seriously, and (2) either disable the GC and manage memory manually or else (my usual approach) allocate everything you need heap-allocated up-front, and subsequently restrict yourself to in-place methods and the stack.

MPI.jl pretty much "just works" in my usage so far (just have to point it towards your cluster's OpenMPI/MPICH) and things like LoopVectorization.jl are great for properly filling up your vector registers with minimal programmer effort.

cbkeller··on Why are most climate models in Fortran?
C++ certainly, and I would not be surprised to see Rust in the near future.

I have not seen anyone use Nim or Zig yet. There are also some special-purpose languages like Fortress (apparently now defunct), Coarray-Fortran, and Chapel, though none seems to have achieved too much market-share.

Personally I have almost entirely switched to Julia (from mostly C), which lets me do my everyday plotting / analysis / interpretation and my HPC (via MPI.jl) in the same language. Fortran definitely still has some appeal as well though.

cbkeller··on Why are most climate models in Fortran?
Well Fortran was, notably, one of the first languages to have proper source-to-source autodiff (TAPENADE) [1-3], so it’s probably not impossible, though my choice for a fully differentiable climate model would personally be Julia, like the CliMA folks at Caltech [4].

[1] https://doi.org/10.1145/2450153.2450158

[2] http://www-tapenade.inria.fr:8080/tapenade/index.jsp

[3] http://www-sop.inria.fr/ecuador/tapenade/distrib/README.html

[4] https://clima.caltech.edu/

cbkeller··on Why are most climate models in Fortran?
You can actually switch starting indices fairly easily in Fortran [1], or for that matter Julia, but for some reason the option doesn’t seem to be particularly widely used. People seem to give a lot of weight to the default, even when the effort to change is minimal. Added complexity/ambiguity?

[1] https://stackoverflow.com/questions/48562873/zero-indexed-ar...

cbkeller··on Why are most climate models in Fortran?
That is a great question, and while I don’t know for sure, I think a relevant anecdotal observation that stands out to me is that many languages which do have first-class multidimensional arrays also have one-based indexing. And I would speculate this in turn is because most people who wanted matrices badly enough to make them a language feature wanted to do linear algebra with them, where you generally also want one-based notation since that is how all the equations in textbooks and papers are written.

So a language with multidimensional arrays is in a lose-lose position of having to choose to either satisfy the linear algebraists at the cost of alienating general-purpose programmers who want to do pointer arithmetic, or else satisfy the latter while alienating the core demographic for multidimensional numeric arrays.

Personally, I’m fine with (or even slightly prefer) one-based for my own scientific computing, despite starting with C, since it really is more elegant for linear algebra, and I have never found myself needing or wanting to do pointer arithmetic in a language that does have good multidimensional arrays — but clearly it is still a major turn-off to many others.

cbkeller··on Why are most climate models in Fortran?
Switching to any sort of commercial grid or cloud computing setup would be rather complicated by the fact that climate models are critically dependent on the fast, low-latency interconnects (e.g., infiniband) of a proper HPC system to achieve good performance at scale. This is usually coordinated with hand-written message passing via MPI directly in the relevant top-level Fortran (or C/++) program.

There are some other (i.e, “embarrassingly parallel”) scientific computing problems where a higher-latency distributed setup would be fine, but in climate models, as in any finite-element model, each grid cell needs to be able to “talk to” its neighbors at each timestep, leading to quite a lot of inter-process communication.

cbkeller··on Julia in RMarkdown
I really like both the Julia and R communities, so always glad to see good interoperability between them!
cbkeller··on Why are most climate models in Fortran?
High performance scientific computing is very much still a FORTRAN, C, and C++ game. And of those, FORTRAN has some compelling advantages in terms of first-class built-in support for multidimensional arrays, and quite excellent compilers. And, as others have noted, until the `restrict` keyword in C99, there were optimizations in FORTRAN that were not even possible in C.

I mostly used C for my (small-scale) HPC work in grad school because it’s what I knew best, but at several points I wished I had learned Fortran instead.

Probably one of the only “higher level” languages that’s ever been used for serious petascale scientific computing is Julia (first with Celeste on the astro side, possibly soon with CliMA for climate modeling), which not coincidentally follows similar array indexing conventions as FORTRAN. And while that’s what I mostly use now, I don’t see Fortran going away any time soon.

If anything, with cool new projects like LFortran [1] making it possible to use Fortran more interactively, it’s probably quite a good time to learn modern Fortran!

[1] https://lfortran.org/

cbkeller··on Ultra-weak gravitational field detected
In that case, go for it ;)
cbkeller··on Ultra-weak gravitational field detected
Quite true! I thought about saying “kg of positrons and kg of electrons” or such so that the mass/charge ratio would be more equivalent.
cbkeller··on Ultra-weak gravitational field detected
Not exactly: if you had a kg of free protons and a kg of free electrons, the electromagnetic forces between them would dwarf the gravitational forces between them even at macroscopic distances. By a lot, as Feynman liked to note in his Lectures [0,1]. It’s just hard to get a kg of free protons exactly because the forces involved are so large. So charges tend to pair up and screen [2], which just camouflages the true strength of the force.

For the strong force the situation is the same in a sense, if you had a kg each of quarks of different colors — except you literally can’t have that, because the energy involved in separating a pair of quarks is so large that it actually causes two new quarks to snap into existence and pair off with the two you were trying to separate before you can separate them on anything approaching macroscopic scales [3]. So the strong force is obligatorily screened (camouflaged) at long distances, but only because it is so strong.

The kg-of-protons experiment on the other hand you could do in principle, it would just probably be better if you didn’t.

[0] https://www.feynmanlectures.caltech.edu/II_01.html

[1] https://sieste.wordpress.com/2012/04/22/feynmans-electric-fo...

[2] https://en.wikipedia.org/wiki/Electric-field_screening

[3] https://en.wikipedia.org/wiki/Color_confinement

cbkeller··on Julia: A Post-Mortem
Arguably it may also be because without JAOT compilation as in Julia, multiple dispatch normally comes with a significant runtime performance cost, as far as I understand.
cbkeller··on Julia: A Post-Mortem
Examples being out of date is definitely an issue since a lot of things are still changing quickly in the ecosystem as a whole. I think it’s getting better as more packages go 1.0 and internals stabilize a bit, but nonetheless.

As far as your comment about “awkwardness”, the key point for me was when I actually started to understand/embrace multiple dispatch as a programming paradigm. If you try to write Julia like in a an imperative or oo or etc. style, it may work, but will be clunky and quite possibly full of type instabilities. But if you write Julia in a “dispatch-oriented” paradigm then it really is just as performant and elegant as advertised, IMO.

cbkeller··on Symbolics.jl: A Modern Computer Algebra System for a Modern Language
Currently the best you can do is to build a somewhat unwieldy relocatable “bundle” with PackageCompiler.jl [1], but there are apparently plans for actual static compilation using an approach more similar to that used in GPUCompiler.jl [2], and the latter approach should as far as I understand allow for the creation of proper .so/.dlls.

[1] https://github.com/JuliaLang/PackageCompiler.jl

[2] https://github.com/JuliaGPU/GPUCompiler.jl

cbkeller··on Reversible Computing
I’ve been wanting to try that package out! Seeing NiLang discussed in the context of automatic differentiation in the Julia community was a bit mind bending, and the first time I ever heard about reversible computing.
cbkeller··on Julia 1.6: what has changed since Julia 1.0?
Then write a function that does everything your script would do in the clean local scope of that function, and call it many times as needed. I mean heck, that’s a more elegant solution even if script latency wasn’t in the equation.
← PreviousPage 5 of 14Next →