“CPUs are optimized for video games”
moderncrypto.org
moderncrypto.org
Sometimes I can't help but wonder how the world where there is no need to spend endless billions on "cybersecurity", "infosec" would look like. Perhaps these billions would be used to create more value for the people. I find it insane that so much money and manpower is spent on scrambling the data to "secure" it from vandal-ish script kiddies (sometimes hired by governments), there is definitely something unhealthy about it.
I don't spend a lot of money on physical security. I leave my car and front door unlocked usually, and don't bother with security systems.
If you find yourself having to lock and bolt everything under the sun lest it get damaged/stolen, then yes, I think it is an indication that the current state of things is wrong. There is something wrong with the economy/community/etc. in your area.
I realize that "the internet" doesn't really have boundaries like physical communities do, but I too wish for a world where security was not an endless abyss sucking money into it and requiring security updates until the end of time. In other words - a world where you could leave the front door unlocked online without having to worry about malicious actors. It will never happen, of course (at least not until the Second Coming ;)
I also assume you don't live in a dense city. Even in a utopia you couldn't have perfect safety without investing in security: some people are going to commit crimes and violate property just for fun, and when there's a lot of people in one area, that becomes a concern.
There has never in history been a time when people can be secure without investing in security. To claim that investing in security is in indication of something "wrong" is therefore at best a sentiment far ahead of its time.
Why are we still assuming that as a given, instead of trying to change it at its roots?
My upbringing has had this "golden rule of ethics" (don't do to others what you don't want done upon yourself) ingrained in me pretty well, so I would never feel excitement about this.
(Hypothesis: This ridiculous objectivism has had people forget how much cooperation is ingrained into humans by nature.)
That said, I am aware that such thoughts exist in my subconsciousness, because the subc is always exploring all possible paths, but I have never experienced that as an even remotely acceptable (in terms of my own moral) path of action.
This only works for people who sufficiently appreciate their own possessions. Others might develop a very relaxed "stuff comes, stuff goes, who cares it's just materialistic crap anyways" attitude and and that equips them to have surprisingly little remorse in regards to theft. This is probably exemplified best in the extremely high rate of theft affecting near-zero-value bikes in many Dutch cities. That kind of thief might even rationalize by inverting guilt, "if he feels bad about the loss it's his own fault that he is not as cool about shitty old piles of rust as I am"
You've missed my point for the opportunity to moralize and posture. Consider this: You know cheating is wrong, you probably consider yourself to be someone that would never do it, yet you can still acknowledge that it would most likely be enjoyable for at least a very short while, right? That's all I'm saying, not (lmao) some kind of faux-objectivist philosophy on the right to take other people's property or whatever it is you seem to be imagining.
There, now the philosophical twaddle is excised without ruining the spirit of the point
Also, thanks for the unjustified rudeness, you've made a great contribution today.
You mean change human nature?
Actually it's the exact opposite: very pragmatic and empirically verified.
It's exactly what people, in statistical quantities, tend to do over what they tend not to do.
Plus, most of it is shared with our animal siblings.
Except if you think that an Elephant or a Cat doesn't have certain natural characteristics (insticts, behaviors, traits) based on their species.
That's because it's a casual online conversation. I don't stuff my responses with citations when I'm not writing a paper.
>Do people ever?
Yes. There are tons of research on human psychology, cognition, physiology, evolutionary traits, instincts and other aspects of what colloquially is called "human nature".
Do you mean chemical lobotomies or a post scarcity economy? Because there are some wants that are impossible to fulfill, for example: there are a bunch of murderers in Syria motivated by the desire to fly a winged horse to a sky mansion where 72 infinity virgins await. I don't see that happening, even post scarcity.
We have a way to go, but we have made at least some progress. Maybe in another few hundred years humans won't be so 'Us and Them'. Maybe materialism will be the next thing we work on. Who knows, maybe by the time we reach the next galaxy we'll have it sorted out.
If the most powerful institution in the western world couldn't "fix" this aspect of human nature over millennia, it is sheer arrogance to think we can do better today.
Something like aggressive behavior though, is something that can be selected against.
Because eugenics went out of fashion a long time ago. The "roots" is free will, so you can't really change that without horrific consequences. Using birth control (voluntary and involuntary) to bias the population was a popular idea in the 1920s. China's latest gamifaction of "social credit" is another biasing attempt. I'm not aware of any "roots" solution that doesn't require total surveillance or head measurements...
And yes, people here (mainly the elderly) do keep vast sums of cash at home http://www.dailymail.co.uk/news/article-2027129/Honest-Japan...
Japan also has some of the densest cities on earth.
Of course, it's far from perfect (there's plenty of crime, people still lock their doors, and sexual violence seems to be off the charts), but it shows that paranoia doesn't have to be the default state. If they made it this far, how much further is possible within human nature?
All the apartments have thick security doors[1], lower windows of buildings have iron bars and all gardens have tall fences. If you leave anything that looks even remotely valuable in your car, someone will break a window and grab it. There is tonnes of vandalism, and even if you leave a motorbike locked with a secure lock it's somewhat likely it won't be there the next day.
It sounds the complete opposite of Japan!
[1] A few weeks after I arrived I was at a friends house and there was a disturbance of some sort in the apartment next door. The police couldn't get in, so called the fire service who had to work for 20 minutes to break the door down. They considered going in through the outside, but it was the third floor and all the windows had metal shutters.
People here are basically honest. There aren't any turnstiles on the transit system --- they just assume you'll have bought a ticket. Children roam the streets on their own and think nothing of talking to random adults.
Of course, they back up this trust with enforcement; there are occasional spot checks on the transit system with big fines; and if you don't pay a driving ticket the police will come to your house and remove your license plates...
Now I guess someone would call a bomb disposal team and cut the traffic 3 blocks around it. (and I would have been very sad, because to this day I still love that bag, leather, light, durable, practical, it still fits a modern 13" laptop perfectly)
Anyhow, this thread is all anecdotes.
That's true. And crime statistics aren't much more useful because willingness to report and how well the police handle filing differ a lot between cultures.
The crime statistics themselves are very useful, there's no need to chuck them about because some level of crime goes unreported. Statisticians have long ago realised that we can measure this too.
1) the fact that the rest of us lock our doors, so most ne'er-do-wells have accepted the notion that closed doors tend to be locked. It's a form of herd immunity that protects you so long as only a small number of people behave like you.
2) the shift from cash to electronic money, meaning that physical theft is less lucrative than it once was (and meaning that most money is protected by entities that invest meaningfully in security.)
You're a free rider, benefiting from effort made by others, while being a (small) net negative to society. You're welcome for the security we've provided for you.
You passively pay for security at a bar when you buy a drink - the price of the bouncer is part of how they derive the cost of your beer. Similarly you pay taxes (local, state, federal) that pay for a massive amount of physical security that you utilize unknowingly on a day to day basis.
Just because you don't find it necessary to lock up your stuff in your driveway doesn't mean that you can claim that you don't use or pay for physical security.
I think it's a density thing. In a small town, there's nobody to steal from you, and even if there was you'd eventually figure out who it was since everybody knows everyone. In a big city, someone can be in and out without you ever knowing, and you would never find them again. Indeed, most of the city residents I met in NZ still locked their doors (the people mentioned above were in smaller villages), while when my family went on vacation to a rural vacation home (in the US) my mom felt comfortable enough to leave the back door unlocked.
There is some social engineering going on there. If you live in an area where everybody locks his door and don't have a huge problem (like bands of bored teens trying every house, or lots of junkies, ...) the chances are very small that anyone would try to open your door, and if they do, they most likely come prepared. I remember that in my previous flat in the center of London, we left the kitchen window opened for years without even realising it. It would have been trivial to break into our place, but nobody cared. They entered 2 times over the same period at my neighbour place, likely because they saw something they wanted to steal through the window.
[1] https://www.buzzfeed.com/copyranter/police-blotter-reports-f...
I was once walking through Atherton and a police officer stopped me to make sure I wasn't doing anything suspicious. I suppose the number of "walking while white" infractions a community tallies might be a good (if unfortunate) indication of how wealthy they really are.
The same police officer later bought me a coffee for no particular reason. What a country.
- Cheapest home for sale in Atherton on Zillow is listed for $3.2 million
- The public schools shown to be in the same area as this house have unbelievably low scores (as measured by GreatSchools.org):
-- Selby Lane Elementary (assigned) .... 4/10
-- Garfield Elementary ..................2/10
What gives?
Before I moved into my second college apartment, the previous tenants warned that they got robbed a few times, and then learned never to open the blinds, and didn't have a problem after that. We therefore never opened the blinds. We pretty quickly broke off a key in the lock and ended up leaving it unlocked except when everyone was gone for multiple days (I won't bother explaining how this worked, or why we didn't get it fixed for a year.)
So our door was always unlocked, but since you couldn't see inside, theft was never the goal of the folks who wandered in. It was always those who were friends and customers of other people in the building who were told they could hang out with us until their friends got back from class or those who had to use the bathroom late at night.
One of my roommates was up until 4 am in the living room every night, and one time someone wandered in while no one was there, and left a note thanking us for letting them use the bathroom, but nothing was ever taken except xbox controllers, and we thought we knew the culprit.
It's not even really a poverty issue - there are places where people are poor as churchmice, but the social stigma around theft is strong enough to keep people in line.
in most suburban places you can also leave the garage door open all day long just fine.
it's primarily dense cities that are the issue. just like online, when there's transience and anonymity, there's a high incentive for people to behave poorly.
Sometimes people tend to behave even more poorly if there is no anonymity. This happens when they think that they are morally right in an important area although they are totally wrong. Many racist "this needs to be said" comments are written using real names.
I went to college in Atlanta, where leaving anything slightly valuable inside your vehicle would lead to a smashed window. So I've developed a lifelong paranoia about never leaving valuable items in sight in a locked car.
Also, HN crew if someone is going to rob your house they don't just try to enter in a locked door. The robber is going to knock, if you answer, make up some story how they are looking for Alice & Bob, and go on their way. If nobody answers, then they try to see if your doors are locked. I used to frequently get strangers "looking for someone" at my door in Atlanta, too. Few things are scarier than hearing a knock at the door when you are in bed, ignoring it, then hearing them try to open your door.
I guess the long commute is worth it.
I once did a Habitat for Humanity build out the Bankhead (not Buckhead, big difference) highway. The had to have Atlanta PD watch the lot DURING THE DAY because multiple cars had been stolen from Habitat Volunteers. I still think that was crazy cars were stolen that brazenly in broad daylight.
I don't have a pulse how things are these days. But don't get scared off by these stories, there are lots of good things about living in the city too. Go Jackets. :)
What did you study btw?
The epitome of this is most of Switzerland. I went on a Tinder date with a girl recently whose apartment door did not even have a lock.
Sure it's different when you're in the city, but out here it's mostly older folks.
If you're suggesting I should move, then sorry. I would love to live in the woods on a lake and live out my days. I could do that easily. But I think I have a responsibility to help solve the problems the economy that pays my rent is causing. I choose to live in the ghetto where I can see what's going on so that I can help.
If that is true, I very much respect that ...
Anyway, I suppose you know, that people usually adopt to their enviroment. So if you see not so nice changes in your ethics and livestyle, you should probably reconsider your home ...
I think it's amazing we manage to spend so little on security.
Try living in a bad neighborhood then.
If you already live somewhere nice and safe, the spendings and effort you don't do on security has been done by the state (police), buying or paying rent for the house, etc.
And perhaps more than it saves (like policing, it's a continuous effort too, you can just make the nice and safe and leave them at that).
In my city you can't use a wooden door, because thief can easily break it with good kick, so everyone uses heavy iron doors with good locks. You can't leave your automobile with factory locks, it'll be stolen, sooner or later. Everyone uses additional security systems. You can't even leave your bag in the closed car, if your windows are not tinted. Someone will see it, break window and stole your bag. You must use heavily tinted window, so no one from the street could see anything inside. Even the thought that I could leave my car or home unlocked is foreign for me. It's my property, so it's my work to defend it.
You would have spent more money on personal entertainment than on your personal security. Your TV, tablet and books would have cost you a lot more than your locks.
Add in the cost of tickets to concerts and plays and your entertainment expenses dwarf any security expenses that you have had.
The current state of things is wrong, and unhealthy. We may one day achieve a healthier reality after recognizing this is so.
The reason for this is simple -- it's twice as slow. The vector width is half the size, and so you can do half as many operations at a time.
This would also explain for instance why many programming languages drop single precision floats altogether: they don't plan to vectorize in the first place.
OTOH rendering of geometry only needs about as much precision as the display offers, which often means 8 bits on older hardware.
"CPU designers see vastly more benefit to spending area on, e.g.,vectorized floating-point multipliers."
followed by "Intel is doubling the vector size in its newest CPUs---again!---this time from 256 bits to 512 bits."
Seems like they are being optimized to be better at vector math, and games just happen to highly use these pieces of HW.
More likely the other way around ...
It may in fact be that desktop/mobile CPUs are being optimized for their contemporary benchmark suites, thus targeting the ensuing benefits in marketing. The benchmarks themselves were, for a fair amount of their existence, focused on games-related performance, at least from what I recall.
In short, I think "optimizing for games" means nothing to a CPU designer at say Intel, and anyways qualitative differences (expanding the ISA, or integrating new features on the IC) are more expensive to develop and sometimes tricky to market. Instead, marketing quantitative differences is much easier (hence optimize for benchmarks) - though arguably no easier to develop. Witness Intel's first rocky attempt at targeting the benchmarks in the early 2000s: Netburst [0]. Of course, in the past 5 years, things have changed (the rise of smartphones, meteoric rise in GPU performance with expanding markets & new software, CPU "per core" performance stagnation), so Intel is in the process of re-positioning itself.
[0] https://en.wikipedia.org/wiki/NetBurst_(microarchitecture)
The integer vector extensions they're almost certainly using I'd argue are primarily designed to optimize various codecs not games; unsurprisingly, crypto algorithms tend to have similar CPU loads to compression algorithms.
Now, put real money on the line and you get the gambling industry which is much more serious about security.
Yes, I get it makes sense that there's more of a market for fun stuff than serious stuff, but putting aside security theatre, false flag operations, etc. - the idea that there aren't really threats in the world seems like it needs a bit more explanation.
In fairness, some of this design philosophy --- which Bernstein has held for a very long time! --- gets a lot clearer with the advent of modern ARX ciphers, which dispense with a lot of black magic in earlier block cipher designs.
A really good paper to read here is the original Salsa20 paper, which explains operation-by-operation why Bernstein chose each component of the cipher.
At some point the inequality gets too far out of whack and it's time to reconsider your values. Same reason we keep switching back and forth between dominant network architectures every 8 years. Local storage gets too fast or too slow relative to network overhead and everyone wants to move the data. Then we catch up and they move it back.
I worked on AAA engines that completely disregarded caches and branch prediction and while that worked great 10 years ago the same architecture became crippled on modern CPUs. Its so very easy to trash CPU resources that at this point I'm convinced most programs easily waste 90% of their CPU time.
For the last AAA title I shipped we spent weeks optimizing the threads scheduling, priorities and affinities along with profiling and whatnot; it was still an incredible challenge to use the main core above 80% and the other cores above 60%. If your architecture doesn't take the hardware into account from the ground up, you're not going to fully use the hardware :)
Hah, in my experience that's hopelessly optimistic. The C I write is probably 10-100x slower than it should be (leaving SSE and threading aside entirely), but the Python I write most days is a hundred times slower still.
* pure Python: 1.0x (baseline)
* NumPy (numeric python): 120x
* Google-optimized pure C: 2,500x
* optimized Python + Cython + BLAS: 8,700x
This is optimizing the same numeric algorithm (word2vec) using different languages and tricks, from naive pure-Python down to compiled Cython with optimized CPU-specific assembly/Fortran (BLAS) for the hotspots [0].
[0] http://rare-technologies.com/word2vec-in-python-part-two-opt...
Fortran yields some of the fastest code out of the box, yes, but C/C++ can yield equivalent performance by adhering to certain basic rules.
I did just that in C# last years and got at least two orders of magnitude of performance gain from it. Still, I constantly wished I was working in C instead :)
I don't know if that's strictly true, but I do feel games aren't as easy to split across threads/cores as many other types of software. Ultimately, a game boils down to one big giant blob of highly mutable and intertwined state (the world) that is modified very frequently with very low latency and where all reads should show a consistent view.
In other domains, avoiding interaction between different parts of your application state is a good thing that leads to easier maintenance and often better behavior. In a game, the whole point is having a rich world where entities in it interact with each other in interesting ways.
That's a common trope; it mostly applies to the rendering path. AI / Physic engines are nicely scalable with the number of cores (AI in particular, because a) it's oftern agent based, which naturally work concurrently and b) with AI, slighly relaxed synchronization constraints sometimes generate some desirable degrees of fuzziness and unpredictability ... although this is probably a very slippery slope ...)
Now it takes a change in perspective on how to structure code and data to properly exploit multiple cores to their full potential. It is certainly possible to have interaction in the game world and do it in a multi-core multi-threaded way, it just needs smarter structures and better organisation of data.
Game developers have and will adapt to this.
Games are still just an update and render pass in a loop. Nothing will change that, they might be decoupled and overlapped but they're still present. They're just much more complex than they used to be. Some engines will fork/join at every step (sometimes hundreds of times per frame), others will build task graphs and reduce them and a 3rd team will use fibers to build implicit graphs. No matter what you do you're still executing code with very clean sequential dependencies.
Game engines started going multi-core at least a decade ago. First by using one thread per system, then moving towards tasks. Today I don't know of any AAA game engine not already parallelizing its render path and unit updates massively.
You don't need interactions between entities at every point in time. There are very specific sync points during a frame and there isn't many of them (maybe 5 or 6). Your units can also be organized in a table where each row can be updated in parallel before moving on to the next.
Many types of software can get away using only concurrency or parallelism. Game engines have to master both. You want parallelism to crunch your unit updates and you want concurrency to overlap your render pass on top of the next frame's update pass.
See http://www.gamedev.net/topic/666419-what-are-your-opinions-o...
On consoles we get access to the hardware specs and profilers for everything - even if behind huge paywalls and confidentiality contracts. Yet with all that information available its still insanely hard to optimize the codebase.
There's absolutely no incentive to push these tools/specs to hobbyists because even professionals have a hard time getting them. Its also worth mentioning that having access to the tools and specs doesn't mean you'll understand how to properly use them.
Nature doesn't have any inherent ethical system, so unless you can fault the universe for existing incorrectly, I don't buy your argument.
A goofy example, but I suspect that the aliens in the Independence Day film lived in such a world; their system didn't seem to have too much security - and why would you need any, in a telepathic society?
That brings up an interesting hypothetical... Is it possible to have a password that you don't even know? Sure biometrics is one method, but are there any password schemes that work on things like word associations or unconscious behaviors found during the typing process, like statistical analysis of the time between key presses?
I remember doing an online course and they had me type several paragraphs in order to determine my typing "signature" but it seems doubtful to me that it would be precise enough to be used for authentication.
https://www.technologyreview.com/s/515726/a-password-so-secr...
Certainly, that's why the "forget password" link is so common ;)
Trying to recall a password that I commited only to muscle memory is damn near impossible without some kind of keyboard in front of me.
(edit: in fact, I went ahead and posted the question: https://www.reddit.com/r/AskScienceFiction/comments/4x7yin/i...)
Seemingly, yes, according to https://www.usenix.org/conference/usenixsecurity12/technical... , which has a paper I read a while ago and a video I never got around to watching.
Abstract:
Cryptographic systems often rely on the secrecy of cryptographic keys given to users. Many schemes, however, cannot resist coercion attacks where the user is forcibly asked by an attacker to reveal the key. These attacks, known as rubber hose cryptanalysis, are often the easiest way to defeat cryptography. We present a defense against coercion attacks using the concept of implicit learning from cognitive psychology. Implicit learning refers to learning of patterns without any conscious knowledge of the learned pattern. We use a carefully crafted computer game to plant a secret password in the participant’s brain without the participant having any conscious knowledge of the trained password. While the planted secret can be used for authentication, the participant cannot be coerced into revealing it since he or she has no conscious knowledge of it. We performed a number of user studies using Amazon’s Mechanical Turk to verify that participants can successfully re-authenticate over time and that they are unable to reconstruct or even recognize short fragments of the planted secret.
Well, aside from the cult of Mac... They spend the same amount of money as gamers, but for bog-standard commodity hardware in a pretty case.
Gamers are willing to pay three times the price for hardware just to get a few percent more performance. I would claim that drives innovation.
I imagine professional workstations for industries such as software development, visual arts, industrial design, film production, music production etc. are large users of high end CPUs.
None of them uses overclocked CPUs (like Intel's K-series) that's fairly standard for high-end gaming
However, I remember running it just fine on an old XP machine with a 32bit, single-core pentium III and 2GB of RAM. In addition, this game has been available on smartphones quite some (phone model) generations ago.
Actually, Minecraft is widely derided as being extremely wasteful with resources. And yet yeah, it still runs on my kids' 9-year-old laptop.
Similarly with war, there is likely some optimal level above zero that increases our overall safety as humans by honing our abilities in force.
The real problem is, for lack of a better term, the mentality of individuals and businesses. It doesn't matter how uncrackable and fast your encryption is if nobody uses it because it's fundamentally inconvenient or hard to justify on a balance-sheet.
Someone should write a song about it.
People would justify getting a better machine for playing games by saying that it would make their work more productive too.
In more recent times, look at all of the video games that made sure they could run well enough for casual players on a reasonably recent vintage laptop. I know I bought a new laptop at least once specifically so that a video game would be playable. My code compiled a lot faster on it (SSD) but I accept that I really bought it for playing games.
edit: found it http://draginol.joeuser.com/article/303512/Piracy_PC_Gaming
Now what do we have? Ssl protecting the packets that we broadcast to everyone on ad delivery platforms.
Now what do we have? Now we're spending more money time and effort, getting poor and limited solutions, trying to patch up email sender verification, shove SELinux around things, wrap every browser request in "are you sure?" modal dialogs and "Request refused for your protection" trip-ups, replace IRC-plain-text-chat with "text-chat in a browser" (Slack, Discord) or walled gardens (iMessage, Google, Skype), fight hard to replace C with languages where buffer-overflows aren't one mistake away and everything isn't de-facto built around rudimentary-types and string-concatenation, and do it all without compromising backwards compatibility or established user experience too hard.
'Awesome'.
'Every time I use a key in a lock, I am reminded that we live in a fallen world'
It's kind of sad isn't it.
My sample size is only the set of projects that I've personally worked on, but in the projects where I have intentionally profiled where the time is being spent, I have only ever run into UX-impacting CPU bottlenecks when working on games. This is spanning code I've written from assembly all the way up to Python. Not counting JS because the DOM is a bit of a black box to me. In almost every other case, the bottleneck was not being able to pull data out of physical storage or some network store fast enough. In some cases it was not being able to pull data from RAM fast enough. In a few cases it was not being able to pull data from a database fast enough due to needing to cover too big of a table, which I lump into the "waiting for I/O" category under the lightly investigated hypothesis that it's not because the CPU is having trouble iterating over the indices, but because the database can't keep the whole index in memory.
Try out any "enterprise" app that thinks its a good idea to load 25,000 widgets on one page. That'll show you the meaning of "CPU bottleneck".
Depending on the implementation it might also be a GPU bottleneck.
Edit those known params were all used as a param for the same column
Sounds like you and I mostly agree. Just that when you do a root cause analysis, "more efficient CPU usage" seems to rarely be the best place to optimize, compared to "do less I/O" or "stop putting 25,000 widgets on one page", or "smaller RAM footprint", and so on.
Agreed, in most cases. That's why I'm not sure what I described is really a problem :)
(I do think that mobile apps and the client-side portion of Web apps, for example, are often CPU/GPU-bottlenecked, though.)
I was just picking at a nuance of what you said about wasted CPU cycles. Yes, some languages and toolchains these days are less efficient to trade for ease of development, but from my experience it seems like I/O to RAM/disk/network is still 80% of the problem and "less work per CPU instruction" is 20% of the problem.
No idea whether mainstream platforms like Ruby or Python do this...it wouldn't surprise me if there's relatively low hanging fruit for speeding up almost every webapp on the planet.
This is too obvious an issue, so there must be a solid reason. What is it?
The compiler should just provide good inlining support, so that if eg. you include the short-string optimization in your stdlib, the compiler can optimize it down to a couple bit operations, a test, and a word copy. If the test fails and your string is more than 7 bytes, it's perfectly fine to call a function - the function call overhead is usually dwarfed by the copy loop for large strings. And then if new hardware comes out and you vectorize it differently, you can get away with replacing that one function in the stdlib instead of recompiling every single program in existence.
So instead, this information/knowledge about high level data types is encapsulated by standard libraries and then the compiler below that. Most CPUs have single instructions to copy a chunk of data from somewhere to somewhere else and a nice basic way to repeat this process efficiently, and it's up to the compiler to use this.
Somewhat related is e.g. the sendfile() syscall that's used by web servers/frameworks to pass a file directly from your fcgi application to the outgoing socket.
Summary: 'Why doesn't "hardware support" automatically translate to "low cost"/"efficiency"? The short answer is, hardware is an electric circuit and you can't do magic with that, there are rules.'
Why don't they just patch the memcmp/memcpy in their libc?
(The string utilities I was referring to actually are separate APIs, focused around manipulating string pieces that are backed by buffers owned by other objects. It's like the slice concept in Go or Rust. With the growth in Google's engineering department, it got very difficult to ensure that everybody knew about them and used them correctly; this is probably easier if they're part of the stdlib. Indeed, they're in boost as string_ref, but most of Google's codebase predates boost - indeed, they were added by a Googler.)
Python 2.7.3 concatenates strings in string_concat using Py_MEMCPY, which is a macro defined at Include/pyport.h:292. That invokes memcpy, except for very short strings, where it just uses a loop, because on some platforms memcpying three bytes is a lot slower than just copying them. In http://canonical.org/~kragen/sw/dev3/propfont.c I got a substantial speedup from writing a short_memcpy function that does this kind of nonsense:
if (nbytes == 4) {
memcpy(dest, src, 4);
} else if (nbytes < 4) {
if (nbytes == 2) {
memcpy(dest, src, 2);
} else if (nbytes < 2) {
The main case in eglibc 2.13 memcpy, which is what Python is invoking on my machine, is as follows: /* Copy just a few bytes to make DSTP aligned. */
len -= (-dstp) % OPSIZ;
BYTE_COPY_FWD (dstp, srcp, (-dstp) % OPSIZ);
/* Copy whole pages from SRCP to DSTP by virtual address manipulation,
as much as possible. */
PAGE_COPY_FWD_MAYBE (dstp, srcp, len, len);
/* Copy from SRCP to DSTP taking advantage of the known alignment of
DSTP. Number of bytes remaining is put in the third argument,
i.e. in LEN. This number may vary from machine to machine. */
WORD_COPY_FWD (dstp, srcp, len, len);
/* Fall out and copy the tail. */
}
/* There are just a few bytes to copy. Use byte memory operations. */
BYTE_COPY_FWD (dstp, srcp, len);
In sysdeps/i386/i586/memcopy.h, WORD_COPY_FWD uses inline assembly to copy 32 bytes per loop iteration, but using %eax and %edx, not using SIMD instructions. It explains: /* Written like this, the Pentium pipeline can execute the loop at a
sustained rate of 2 instructions/clock, or asymptotically 480
Mbytes/second at 60Mhz. */
This is presumably what Jeff's code was a replacement for. Too bad he didn't contribute it to glibc, but he presumably wrote it at a time in Google's lifetime where the Google paranoia was at its absolute peak.PAGE_COPY_FWD sounds awesome but it's only defined on Mach. Elsewhere PAGE_COPY_FWD_MAYBE just invokes WORD_COPY_FWD.
My tentative conclusion is that Ulrich scared everyone else away from wanting to work on memcpy so effectively that it's been unmaintained since sometime in the previous millennium.
A CPU for games would have very fast cores, larger cache, faster (less latency) branch prediction, fast apu and double floating point.
Few games care about multicore, many "rules" are completely serial, and more cores doesn't help.
Also, gigantic simd is nice, but most games never use it, unless it is ancient, because compatibility with old machines is important to have wide market.
And again, many cpu demanding games are running serial algorithms with serial data, matrix are usually only essential to stuff that the gpu is doing anyway.
To me, cpus are instead are optimized for intel biggest clients (server and office machines)
Games are usually optimized for 4 cores max.
For example, i7-6700K (4 cores) perform better than i7-6950X (10 cores) in almost every modern game.
It's also quite likely that spreading out over more cores would be counterproductive. Many problems are not parallizeable indefinitely.
Large games will use all cores, our game uses 32 cores if you have them, we solved a bug because 1 guy had such a machine and reported an issue. 1 guy, from the public, and we fixed it the next update!
But no, developers don't care, we don't even like games! In fact, we hate games! Please don't enjoy our games!
[0] http://www.pcworld.com/article/3039552/hardware/tested-how-m...
[1] https://imgtec.com/blog/vulkan-scaling-to-multiple-threads/
The CPU industry stayed on this path for as long as it were physically possible, even long after the time when it hit diminishing returns on single thread performance divided by (area*power). Pentium 4 was the last CPU of this single-core era.
If you look closely at the microarchitecture of the modern desktop CPU, the out-of-order execution, caches and branch prediction are already maximized (to the point that >2/3ds of die area is cache). Multi-core has become mainstream only after all other paths became exhausted.
In my experience SIMD actually becomes more important when you want compatibility with old machines, because your ability to use the compute power of the GPU becomes more limited the further back you go in graphics library versions. For example, OpenGL has no compute shader before 4.6, no tessellation shader before 4.0, and no transform feedback before 3.0. When you can't make the GPU do what you want, SIMD becomes your best bet…
But as an engine programmer, I agree with the linked author. I'll take your points one at a time.
Most engines are multi-core, but we do different things on each core (and this is where Intel's hyper-threading, where portions are shared between the virtual cores, for cheaper than entire new cores, is a solid win). Typically a game will have at least a game logic thread (what you are used to programming on) and a "system" thread which is responsible for getting input out of the OS and pushing the rendering commands to the card along with some other things. Then we typically have a pool of threads (n - 1; n is the number logical core of the machine; -2 for the two main threads, +1 to saturate) which pull work off of an asynchronous task list: load files from disk, wait for servers to get back to us, render UI, path-finding, AI decisions, physics and rendering optimization/pre-processing, etc.
AAA game studios will use up to 4 core threads by carefully orchestrating data between physics, networking, game logic, systems, and rendering tasks (e.g. thread A may do some networking (33%), and then do rendering (66%), thread B might do scene traversal (66%), and then input (33%), see the 33% overlap?), they also do this to better optimize for consoles. But then they have better control of their game devs and can break game logic into different sections to be better parallelized, where as consumer game engines have to maintain the single thread perception.
SIMD is used everywhere, physics uses it, rendering uses it, UI drawing can use it, AI algorithms can use it. Many engines (your physics or rendering library included) will compile the same function 3 or 4 different ways so that we can use the latest available on load. It's not great for game logic because it's expensive to load into and out of, but for some key stuff it's amazing for performance.
That stuff the GPU is doing eats up a whole core or more of CPU time. So what if we are generally running serial algorithms, we need to run 6 different serial algorithms at once, that's what the general purpose CPUs were built for.
This is all the stuff you don't often have to deal with coddled by your game engine. The same way that webdevs don't have to worry about how the web browser is optimizing their web pages.
Pretty much all console games care about multicore.
> Do CPU designers spend area on niche operations such as _binary-field_ multiplication? Sometimes, yes, but not much area. Given how CPUs are actually used, CPU designers see vastly more benefit to spending area on, e.g., vectorized floating-point multipliers.
So, CPUs are not "optimized for video games", they are optimized for "vectorized floating-point multipliers". Something video game (and many others) benefits from.
This is also an optimization that compilers can readily take advantage of on a small scale (similar to pipelining) so the combination of benefit + ability to use + simplicity/low resource use makes it an inevitability.
Anything more specialized exists in the form of DSPs GPUs, etc..
edit: I checked. Microsoft at least used floating point in Excel until much later, until at least 2013. With the expected limitations. https://support.microsoft.com/en-ca/kb/78113
Here's what I remember from back then: early PCs (8-bit era: Apple // etc) had no socket for a FPU (i.e. not even an expansion option), every IBM-compatible PC had a co-processor socket starting with the 8086 (fun fact: the 8087 wasn't even shipping when the PC was designed, and boy were they expensive for what they did when they shipped) but almost no one bought one until the late-286/early-386 era for general-purpose computing, by the 386 era the FPU was pretty much standard equipment on any 'real' business PC. So the software evolution was: no FPU support, optional FPU support, FPU required. (i.e. by the last generation of DOS applications, several major apps had dropped their floating point emulation libraries and would crap out with an error along the lines of 'coprocessor not detected/installed')
The evolution of the FPU was very similar to how the GPU has played out in terms of becoming a standard, expected component. As with the FPU, CAD and all sorts of other scientific/business applications were the initial drivers of the, but gaming is what caused unit volumes to explode and the reason it is now standard equipment.
For very loose definitions of "quickly" and/or "around the time". Spreadsheets became the "killer app" with VisiCalc in 1979 -- before the IBM PC was even a thing; but Intel processors through the 386 didn't have an integrated FPU, and the 486 (around 1990) came in both integrated-FPU (486DX) and no-integrated-FPU (486SX) models. It wasn't until the Pentium that integrated-FPU was universal in the Intel line.
(AFAIK, most spreadsheets of the early PC era didn't even support using the FPU, and the main widely-used application category that leveraged FPUs in the optional-FPU era was CAD.)
Not true at all on modern X86. You just get half the FLOPS by using double precision instead of single precision -- 16 double precision FLOPS per core per cycle. Which is very good considering there's also twice as much data to process.
You could rather say X86 CPUs are highly optimized for double precision performance.
I still wouldn't say that they're highly optimized, though.
"In most cases, double precision calculations take no more time than single precision."
AES encryption support: https://en.wikipedia.org/wiki/AES_instruction_set
Hardware video encoding/decoding support (I presume for phones): https://en.wikipedia.org/wiki/Intel_Quick_Sync_Video
It's more that it's relatively easy to make some instruction useful to a variety of video game problems, but difficult to do the same for encryption or compression. You tend to end up with hardware support for specific standards.
The fact that Intel put an encryption feature in their chip, which does indeed make that algorithm faster, would tend to indicate they wanted faster encryption wouldn't it? That some other algorithm could be faster still isn't really contradicting that.
Gamers generally don't care about power consumption. When you've spent $1000 on the hardware an extra dollar or two on your electricity bill is no big deal.
Citation needed. Where did you get that idea? Please show how djb's vector code spends more power vs the built-in AES "dedicated hardware" instruction when, as he measures:
"* Both ciphers are ~1.7 cycles/byte on Westmere (introduced 2010).
* Both ciphers are ~1.5 cycles/byte on Ivy Bridge (introduced 2012).
* Both ciphers are ~0.8 cycles/byte on Skylake (introduced 2015)."
"even though AES-192 has "hardware support", a smaller key, a smaller block size, and smaller data limits" (his code is 256 bits and 12 rounds).
He's also comparing 12 rounds of ChaCha with 12-round AES, which may be fair, but realistically no one uses ChaCha12, they use ChaCha20.
http://www.intel.com/content/dam/www/public/us/en/documents/...
Intel has done a lot of things to try to balance this. One of those things is they don't even bother turning half the vector unit on unless you use it a lot. If you seldom issue an op with 512-bit operands, the CPU will actually dispatch them as multiple 256-bit operations, in which case you won't incur the drop in clock, but you also don't get the supposed benefit of double throughput. Furthermore the performance may be much worse if the CPU decides to turn up the remaining vector bits, because the clock drops dramatically while those units are charging up.
So you can see that for someone trying to wring out every last bit of performance on a recent Intel CPU using all the advertised vector capabilities, optimization can become quite complicated.
The big CPU hog and prime candidates for these vector operations nowadays seems to be video encoding.
http://blogs.unity3d.com/2014/07/08/high-performance-physics...
It's syntax is also pretty close to C and C++ which means developers with game dev background will feel at home as most game development is done in C++.
Unreal Engine uses Unreal Script which is now pretty much C++ but it is also not compiled directly (although with Unreal Engine 4 and onwards it's much closer to direct compile than any other scripting language).
Unity engine has it's own interpreter which then builds highly optimized C++ code and compiles it when you build the game.
Unity Engine is a pretty decent engine with kickass performance when optimized, without fine optimization any general purpose engine including Unreal 4 acts like utter crap. I'm alpha/beta testing a few UE4 games atm and you can see just how bad performance can get even on a solid defacto industry standard like UE4 like when dynamic shadows tank a GTX Titan X (Maxwell) SLI setup to below 20 fps any time there are light sources that are not properly fenced and culled - e.g. explosions.
Do they really need a bazillion shaders and dynamic shadows on everything?
There are more unreal engine titles for any given version than any other engine on the market on PC's and consoles.
On mobile unity is probably bigger atm.
CPU performance is important when there is a lot going on (most of the frame latency is CPU side), and when the CPU's are really busy dealing with other stuff like heavy AI or tons of NPC's you get poor frame rates and scalability with GPU processing power.
Good examples of real CPU hogs would be the latest Civ games as well as games like Assassin's Creed Unity. ACU is especially a CPU killer, if your CPU OC is stable playing ACU it's really stable, I've seen that game cause CPU's that are not technically overclocked crash on their normal boost clock when the memory XPS profile was loaded.
The two games your parent named basically are physics simulators.
Minecraft is also CPU bottlenecked pretty much but that's because well the game was built with Java :)
All those voxels take up a good chunk (heh) of processing time. They are not all static, you know. Not sure how much of that could be shifted to the GPU.
• Only works on Windows
• Only works with Nvidia graphics cards (and only them, cheap nvidia card for physx + AMD card for graphics will disable PhysX)
• Not guaranteed to be faster for all cases, needs to be evaluated on a case-by-case basis
So even if the physics engine can, not all game engines use it. The Unity devs e.g. stated that they won't bother with it, as the limitations make it unattractive to pour effort into.
At the end of the day, most scientific problems based on continuum mechanics need really fast level 1, 2, and 3 BLAS operations. That's mostly enough for simulating physics based on explicit time integrators. For elliptic problems or problems that require implicit time integrators, we need factorizations. Most of the time, we can get away with LU, Choleski, QR, and SVD. Both dense and sparse are required. For optimization with equality constraints, factorizations are also required.
By the way, if anyone wants to figure out how to do faster factorizations, dense and sparse, on whatever new hardware is coming out, that'd have an enormous impact on the scientific community. There are people working on it. There's not a lot of them.
And you will need some of those results on the CPU.
Physics for particles is not uncommon to be done on the GPU though. There is no feed-back to the CPU required so latency becomes a non-issue.
Because of the single core performance unless you are playing a really CPU intensive game which has been optimized for multiple cores (more than 4) (or running an insane SLI setup 3-4 way and need the extra PCIE lanes) the 6700K with as high overclock as possible (4.6-4.7ghz) is pretty much the bare minimum for GTX 1080 or better GPU's and even that isn't enough.
Here's Bioshock infinite: http://images.anandtech.com/graphs/graph8426/67045.png
Here is F1 2013: http://images.anandtech.com/graphs/graph8426/67044.png
So sometimes it matters almost as much as regular productivity task (like compiling software): http://media.bestofmicro.com/ext/aHR0cDovL21lZGlhLmJlc3RvZm1...
There are several outliers to this - some games are VERY cpu dependant (supreme commander, MMO's like planetside 2), and some which are EXTREMELY inefficient games (Minecraft) which require an order of magnitude more CPU power than it really needs due to architecture problems.
Bottlenecking just means that improving the bottlenecked thing will result in the biggest improvement. The second-biggest performance limiter can still be a significant performance limiter.
Many data centers go through a process of constant expansion and renewal, as they are competing with others and there is a larger financial incentive for them to do so.
Similarly, 4k displays require a lot of GPU power, though I think it's viewed as even more frivolous / excessive than building for VR.
As for datacenters and compute farms, the large ones are upgrading portions of their systems yearly, if not expanding on top of that.
In fact, games have always driven the modern computer industry. Even Unix started because of a game (http://www.unix.org/what_is_unix/history_timeline.html).
Honestly I would treat this the same as eg Ethernet - high end cards have hardware offload capabilities that the software stack can utilise to get better performance.
DJB's interest here is specifically in creating algorithms that work well on general-purpose popular CPUs.
It's quite possibly badly optimised too, but even if it weren't, Arma would eat CPU like crazy.
Some stories from back around 2000 when designing CPUs at Intel. Some people did bemoan the fact the few software actually needed the performance in the processors we were building. One of the benchmarks where the performance is actually needed was ripping DVDs. That lead to the unofficial saying "The future of CPU performance is in copyright infringement." (Not seriously, mind you)
However, here is a case where the CPUs were actually modified to improve one certain program.
From: https://www.cs.rice.edu/~vardi/comp607/bentley.pdf (section 2.3)
"We ran these simulation models on either interactive workstations or compute servers – initially, these were legacy IBM RS6Ks running AIX, but over the course of the project we transitioned to using mostly Pentium® III based systems running Linux. The full-chip model ran at speeds ranging from 05-0.6 Hz on the oldest RS6K machines to 3-5 Hz on the Pentium® III based systems (we have recently started to deploy Pentium® 4 based systems into our computing pool and are seeing full-chip SRTL model simulation speeds of around 15 Hz on these machines)"
You can see that the P6-based processors (PIII) were a lot faster than the RS6K's and the Wmt version (P4) was faster still? That program is csim and it is a program that does a really dumb translation of the SRTL model of the chip (think verilog) to C code that then gets compiled with GCC. (the Intel compiler choked) That code was huge and it had loops with 2M basic blocks. It totally didn't fit in any instruction cache for processors. Most processors assume they are running from the instruction cache and stall when reading from memory. Since running csim is one of the testcases we used when evaluating performance the frontend was designed to execute directly from memory. The frontend would pipeline cacheline fetches from memory which the decoders would unpack in parallel. It could execute at the memory read bandwidth. This was improved more on Wmt. This behavior probably helps some other read programs now, but at the time this was the only case we saw where it really mattered.
The end of the section is unrelated but fun:
"By tapeout we were averaging 5-6 billion cycles per week and had accumulated over 200 billion (to be precise, 2.384 * 1011) SRTL simulation cycles of all types. This may sound like a lot, but to put it into perspective, it is roughly equivalent to 2 minutes on a single 1 GHz CPU!"
Games were important but at the time most of the performance came from the graphics card. In recent years Intel has improved the on-chip graphics and offloaded some of the 3d work to the processor using these vector extensions. That is to reclaim the money going to the graphic card companies.