To me this seems like poor design just like locales for languages are. As the computing devices a program runs on is dynamic it shouldn't me modeled as global state. Five tasks are run and should be scheduled on N cpus and M gpus.
That said, other communities may obviously prefer different approaches due to differing needs and constraints.
Once the data is created, computations can be executed with affinity to a specific variable in a data-driven manner using patterns like `on myVar do foo(myVar, anotherVar)`. Alternatively, an abstraction can abstract such details away from a user's concern and control the affinity within its implementation, as the parallel iterator implementing `forall elem in MyDistributedArray` does.
var HostArr: [1..10] int; // allocated on the host memory
on here.gpus[0] {
// now we are on a GPU sublocale...
var DevArr:[1..10] int; // allocated on the device memory
...
}
In the near term, we are planning to publish our 2nd GPU blog post where we will discuss how to move data between device and host.