146 karma · joined October 10, 2011
[0] http://www.phoronix.com/scan.php?page=news_item&px=MTY2OTg
Many of Bartosz' discussions focus on concurrency. It is one of the big and interesting problems we have at the moment. The criticism repeated here that "No one will ever be able to debug it" and the like is exactly what Bartosz sees in our future if we continue to program concurrent systems the way we do. He favours adoption of functional programming concepts to manage the complexity of concurrent systems. And I agree.
If code similar to the code in his blog is not acceptable in production code, a code review should make sure it never reaches production. An informed discussion about the pros and cons of such code could ensue during the code review, new insights are gained and everyone walks away a better, more knowledgeable programmer. I suspect this is one of Bartosz main goals.
This approach makes a lot of sense: abstract the annoying low level architecture details. They have a lot of targets which is fantastic: x86/SSE, ARM v7/NEON, CUDA, Native Client, and OpenCL. Let the architecture specialist worry about the architecture specifics. The disadvantage: the achieved performance then depends on quality and wisdom of the compiler. But once certain things are optimized for a specific architecture, every user will benefit.
How they do it on the compiler end of things, I'm not sure. There are a number of techniques. Among the simpler is auto-tuning. There is also a new term: "copious-parallelism" [0]. It acknowledges that to achieve performance portability across platforms, algorithms must offer explicit ways of parametrization and tuning to adapt to different platforms. I think this is the right concept but believe that it could be implemented within the compiler. The domain specialist should not have to think about those things.
[0] http://www.hpcwire.com/2014/01/09/future-accelerator-program...
"Hueman is free and anonymous, and to create that seamless experience, it needs to store data on your phone. If it was a web page on mobile safari, that data is more ephemeral."
If the Hueman developers read this: have you considered something like unhosted [0] for your application? It should be easy enough to implement.
After publishing my own rant about the same topic [0], Dave Winer commented via Twitter and argued that RSS is a very successful tool that decentralizes the web. I didn't understand his point at the time but today I think he is absolutely right.
His point was that when trying to fix the problem we should look at what already exists and works. Why did RSS succeed? Probably because it solved a problem for many people. It did not try to reinvent old things in a decentralized way but offered real benefits. Another good example for this git vs svn (centralized vs decentralized revision system).
[0] http://www.soa-world.de/echelon/2011/09/the-decentralized-we...
[0] http://en.wikipedia.org/wiki/Foundations_of_Geopolitics [1] http://www.hoover.org/publications/hoover-digest/article/807...
I am trying to see if I can put a) a portable interface together (for both writing kernels and the coordination of cards, contexts, memory, queues/streams etc.) and b) if the performance is portable. I can already see that performance for trivial kernels is portable from AMD to NVIDIA but as soon as I go to the Intel PHI things are suddenly very different.
For me, the most important thing in both CUDA and OpenCL is the programming model. It allows us to describe data parallel problems and related data (in)dependence explicitly. Compilers should be (and already are) able to generate efficient code from this. It is not as nice as it could be, we have to write kernels by hand etc. But there are libraries that make our lives easier (we discussed them earlier). And there is also C++ AMP which tries to integrate better. Yet still, while we have all these options and we can solve most of our problems with more or less effort and elegance, I believe there must be something better out there: the right way to describe data parallel and task parallel problems as well as concurrency etc. Maybe the FP guys are on to something, I don't know. I'll be on the lookout.
[0] http://www.pds.ewi.tudelft.nl/fileadmin/pds/homepages/shenji... [1] http://comparch.gatech.edu/hparch/papers/lee_plc2013.pdf
CPU is supported through OpenCL by these libraries. Remember, OpenCL code can run on CPUs. As for Aura, it is pre-alpha, not a lot of functionality there yet, I'm still figuring out the interface. So not usable yet, but keep an eye on it, it will be.
In my own library Aura, I focus on maximizing performance (over developer convenience). The target audience are developers of real-time applications that need every last drop of performance from their hardware, while still maintaining a sane and cross-platform API. I strive for a Boost.Asio for accelerator developers. Aura has a already a rudimentary wrapper for clFFT, clBLAS is in the works. So the idea is to, for each platform, utilize optimal vendor-supplied library functions and combine them in a coherent interface.
[0] https://github.com/kylelutz/compute
[1] https://github.com/ddemidov/vexcl
Taking a potent drug like heroine or cocaine for the first time is described by some users as "pushing open a door to a very enjoyable place you did not know existed". Once this door is open, you will always know that this place exists, you can not "unknow" the experience.
While some people might be able to handle the knowledge of this enjoyable place and are able to choose when and if to go back, some might not be able to. Therefore I believe it makes sense to tell people not to open the door at all.
I also can imagine that the first time with these potent drugs is a positive experience for many, unlike the first cigarette or the first beer a person ever tries. The second cigarette is not smoked because the first was so good.
What excites me about their project is a very simple file-system based interface they are planning to implement for their upcoming version. Plug the key in, it looks like a regular USB key, put a file in a specific place, get an encrypted/signed/whatever file from the file system from another place. No driver or software required. I hope their project takes off and people start buying these keys. The price is a little high right now (59,00 €) and it is currently not available.
compute a() in parallel with b() and
handle mouse events concurrently (could be a()) with keyboard events (could be b())
differ and how does
compute a(0), a(1), ..., a(N) in parallel differ from
handle events from 1000 sockets concurrently which would be something like a(event0) a(event1)?
Don't we require the same tools for "task parallel" and "task concurrent" as well as "data parallel" and "data concurrent"? To me they don't seem orthogonal but I believe they are very much the same thing.
[0] http://www.nature.com/nature/journal/v473/n7346/full/nature1...
Task parallelism and data parallelism are two important concepts and the tools we need to implement them are very different. The two words used by the author to convey the difference between those two concepts are not chosen optimal in my opinion. There seems to be a hierarchy between concurrency and parallelism whereas task parallelism and data parallelism are directly comparable.
Too often is the inventor the Faustian idealist who tries to change the world but fails due to the harsh reality. If he wants to implement his idea, he has to get involved with powers, whose sense for reality is sharper and more pronounced than his. Today such powers are, without expressing a value judgment, military and management. [...] In my experience, chances are slim to defend oneself against such pacts.
I don't think this is true for everyone. In fact I believe recently we have seen a renaissance of native code. C++ is the number one language for that. Those higher level languages waste a lot of cycles. You pay the warm and fuzzy feeling during development with performance and/or energy.
C++ challenges you to think in terms of types instead of functions. This is in my opinion often a more natural way to approach and solve problems and thus produces more elegant code. Because of that, C++ is a very powerful tool for library developers - you can design libraries that allow domain-specific developers to write expressive code that performs as if you had written more verbose C code. And this C++ code is not only more expressive for your fellow co-workers to read and understand but also in some cases for the compiler. This is the reason why C++ code can sometimes be compiled to faster machine code: you can give the compiler more information about what you actually try to accomplish (keyword: template meta-programming).
As with most things there is no definite answer. C code can be quite elegant and appropriate as well.
Also people working with those devices are often scientists that are eager to share what they found out (if only to say "You're doing it wrong!"). See for example Vasily Volkov's work here http://www.cs.berkeley.edu/~volkov/