64 bits ought to be enough for anybody
blog.trailofbits.com
blog.trailofbits.com
OpenCL needs more examples, tutorials, and documentation. I really wanted to use it because it is cross platform, but quickly gave up. The tutorials and documentation wasn’t there, and what I did find was confusing. For CUDA there is an immense body of examples, support, and documentation that makes creating your first CUDA program dead simple.
I went with CUDA.
Isn't it dead simple only if you have compatible hardware? OpenCL works everywhere, even in the worst case on the CPU through pocl.
If you want to use C++ then it definitely doesn't work everywhere. Precompiled kernels (SPIR-V) also do not work (anywhere?).
Writing a GPU C program that only gets compiled when you run your program is extra overhead on productivity.
Same applies to Vulkan versus other next-gen APIs.
That’s unfortunate because it would be interesting to compare virtualized V100 GPU with real hardware 1080Ti. On paper, the two are pretty close. If that’s true in practice, you can use https://vast.ai to reduce the cost by a factor of ~5, i.e. it will only cost about $350.
I don’t need their services, but I think the prices are reasonable.
Nvidia is way too greedy, they charged $5000 for server equivalent of a $700 consumer GPU. The major difference between the two GPUs is legal, not technical.
Cloud providers like Google or Amazon have little choice but to pass the cost to users.
1080Ti / P100:
Chip: GP102 / GP100
Cores: 3584 / 3584
Base clock: 1480 / 1328
Single precision TFlops: 10.6 / 10.6
TDP, Watts: 250 / 300
The only major difference of the chip is double precision performance, probably irrelevant for the OP's task. And I'm not convinced the difference is in the silicon as opposed to firmware or drivers, games don't need doubles so the NV could cripple the consumer model without consequences, just to differentiate the two products.
nVidia also made Tesla P40. It has 100% identical chip to the 1080Ti, same GP102, same frequencies, 100% same specs. The price is around $5000. The GPU is not 100% same, it has more VRAM, but still, that's 8x price difference for almost identical products.
If Teslas would be better, nVidia wouldn't need to put that ridiculous paragraph in the EULA of the driver, about data center usage. That paragraph is the only reason why we don't have affordable GPGPUs in AWS.
From techpowerup:
GP100: 610 mm^2
GP102: 471 mm^2
E.g. to create an SSE register with all 64 bit lanes set to the same value, use _mm_set1_epi64x, to different values _mm_set_epi64x (don't forget to flip the order). No need to use GCC's proprietary __attribute__ ((aligned (16))).
Another portable option is C++ and alignas keyword.
The impetus for doing this was though a real discussion when we were debating whether to do an analysis to identify some constants, or just to brute force comparisons. At what point was the analysis faster? One thing led to another and next thing I know I was learning CUDA...
I also don't understand what type of real-world application is being simulated here. I do understand that it is a simplified example, but of what?
The takeaway is that it is feasible to perform a single comparison operation for each 64-bit bitstring. But we don't just do a single operation per item in the real world.
But any kind of substantial operation like building data structures and spilling to main memory can trivially leave you with a program that takes weeks.
But these tests shows minimal difference between AVX2 and AVX512. So what’s the big deal?
The clock speed reduction also affects other cores, which may not be executing AVX512 instructions at that time. Code that is nearly 100% AVX512 will get an overall speed boost, but mixed multi-threaded workloads can actually regress in performance. On virtualised or multi-user systems such as Citrix or some database engines this is a serious issue at the moment.
This will improve or go away entirely once they're on 10nm or some smaller process.
Secondly, the rarity of AVX512 means that few applications take advantage of this instruction set, and those that do have not had anywhere the same level of fine-tuning as the more common AVX2 code.
For example, SQL Server recently got vector instruction set support, called "Batch Mode Processing", but I can't find any references to indicate that it uses AVX512, and it probably doesn't.
Once AMD supports AVX512, it trickles down to the mainstream CPUs, and the clock speeds are maintained it'll have a significant advantage over AVX2.
In this particular use case those did not really matter, but when you can use them they really help.
It's highly likely there's a market for it, however I'm doubtful they'd be willing to spec lower-density, higher-performance machines for the few people who would actually take advantage of them. Considering high core count CPUs are becoming the new norm, re-writing or finding software that can take advantage of multiple cores/threads is certainly the way to for maximum performance (where possible, of course)
I half remember an article about stock trading companies overclocking servers and running them single threaded at max clock speed. The aim was shaving a few more microseconds from their response time. They accepted the reduced lifetime of the hardware, just replaced it sooner. I got the impression it was a niche market though.
The initial cost is what makes this insane, not the time, and the end result (GPUs) is undebatable the most cost efficient method (in this case very likely by a factor of 50-100x).
A situation like this is not premature optimization. Developers sometimes need to understand that in the scheme of things, their time is not worth very much compared to the cost of the infrastructure required to run their code. Throwing the corporate credit card at optimization issues is far too common today in tech.
We could have gifted them all $5000 machines and spent less. Plus, when you remember that the point of spending on developers is to earn back many times that cost in sales, the opportunity cost of that 3 months of very senior development effort was massive.
Just putting a problem on a ton of cores is an absolutely valid strategy for problems that are a bad fit for GPUs.
so we'll still laugh at this comment even though the addressable space turned out to be "good enough", it was just an antiquated framework to begin with
The takeaway here is that cracking a 64bit password only cost $1700 and is totally doable in 3 weeks.
If you need things to be really safe I guess you need to start lengthening those passwords.
Not quite - you need to run the hash function that the password is stored with, and that might be a factor of 1,000,000+ slowdown relative to "are these numbers the same".
Those $1700 of computational costs are likely to be run on hacked infrastructure, so definitely doable.
1.7 billion in computation
>The takeaway here is that cracking a 64bit password only cost $1700 and is totally doable in 3 weeks.
...assuming that each "guess" only requires a comparison. If you factor in hashing, it's much slower. According to gpuhashcat benchmarks[1], a 1080 ti crack md5 passwords at ~25 GH/s (1GH/s = 1 billion guesses per second), phpass at ~6.9 MH/s, and bcrypt passwords at 13 KH/s.
[1] https://gist.github.com/epixoip/a83d38f412b4737e99bbef804a27...