Regent: A Language for Implicit Dataflow Parallelism
regent-lang.org
regent-lang.org
So far I've focused on HPC users. I think focusing on a niche carries some risk that the language might be too narrow (and this is particularly true of HPC). But one thing I like about HPC is that HPC users really know what they're talking about when it comes to performance.
It gets me a server error. Is it still online ?
But did you find it was hard to implement enough of a language to run useful benchmarks?
In my experience, porting a ~5k line application is usually a week or two worth of work, and usually most of that is spent trying to understand what the original code is doing, rather than writing a lot of new code. Regent code is usually fairly compact compared to the sorts of things we benchmark against (C++ and MPI codes).
The advantage that Regent has over a C++ library like Legion is that the language can embody the abstractions we want to provide. With C++, there is always an impedance mismatch between how we want people to think about programming, and what C++ forces us to do. C++ code is more verbose, and the C++ type system doesn't capture all the invariants that the programming model requires.
Regent code is also a lot more amenable to static analysis. In Legion, if the system needs to know some property of the program, the user needs to supply that information. (And we're forced to trust whatever the user says---there's no reasonable way to statically analyze arbitrary C++ to validate what they've given us. At best, we can check those properties at runtime, but that isn't always feasible either.) In Regent, many of these sorts of things can be analyzed automatically, reducing the number of moving parts users have to think about. This is only possible because we've thought very carefully about what abstractions we need in the language---the type system is already on the edge of what is decidable, so adding anything more would make analysis intractable.
To be sure, C++ has its advantages too. Some users will simply never accept a new language. So, that's why we're doing both. And for those who prefer to stick with C++, that's always going to be an option.
I agree that static analyses and invariant enforcement/extraction is a good reason to have a new language (I think most of the time you have to "get around" with "convention" when using C++ EDSL).
Also, do you think the type system / syntax could be implemented in something like Julia by a combination of macros and types?
But I'm also curious about the parallelization costs, and other costs incurred by the programming model. The best way to explore those costs are to compare against an optimized sequential version that solves the same problem. Do you have any experiments which do that?
Understanding what it is useful for will probably help programmers apply it somewhere.
Here are a couple representative examples of what we've done so far:
* A hydrodynamics simulation on a 2D unstructured mesh (a pointer data structure): https://github.com/StanfordLegion/legion/blob/master/languag...
* A simulation of an electric circuit (a graph): https://github.com/StanfordLegion/legion/blob/master/languag...
(We have smaller toy examples that work with linked lists, trees, matrices, etc.: https://github.com/StanfordLegion/legion/tree/master/languag...)
var grid = ispace(int2d, { x = 4, y = 4 })
var matrix = region(grid, float)
But beyond that, most things are left to the programmer. E.g., if you want to build a mesh data structure, it's up to you to decide exactly how to build it. (Do you build a graph with pointers? Do you make it an N-dimensional grid? etc.) But my impression talking to application programmers is that most of the time, they want to make those decisions themselves anyway. A standard library might be nice, but it's never going to fit all use cases. So the first job of any language design is to make sure the abstractions are general enough to describe what users want.