A few years later my brother was starting a web server to host a forum for the Digital TV switch-over focused on the Madison, WI market, shout-out and RIP madcityhd.com
Watching him setup Fedora on the box on the floor of his bedroom and using Compiz wobbly-windows was enough to hook me. I stole his install CD and nuked my drive (much to the irritation of my dad, who knew that I would be stealing one of his nights to reinstall Windows when I eventually realized could no longer play counter-strike), It was fun to see it come full-circle a year or two after I was ticked-off about my Anthlon64 not running a 64 bit operating system, when I started trying other distos and realized, “Hey, I can actually use the 64 bit one now!”
Fun times. I wonder how many kids got hooked on Linux by wobbly windows. I know that’s what brought me in, haha.
https://en.wikipedia.org/wiki/Windows_XP_Professional_x64_Ed...
It is entirely possible to make a modern system that use far less than they do today. If you're gaming it's a different beast, but for normal OS and programs it's certainly bloat.
If I've learned anything in my computer career, it is that this is an evergreen comment. I remember reading Q-Link posts from Commodore 64 users with this complaint when the wasteful C128 came out.
I wouldn't call it consensus, just a persistent claim by a large fraction of developers. They finally got to prove their claims with x32 ABI, which was ILP32 in amd64 mode, including the additional registers. Few if any people used it, nor did it show better performance in practice.
Regarding 64-bit support: by the time amd64 chips shipped people had been writing software for the 64-bit Alpha for 10 years, and 7 years for sparc64. IME most open source software was already 64-bit clean and worked well out-of-the-box. Back then the open source community, and especially GNU projects, heavily emphasized platform and hardware portability.
Perhaps the situation on Windows was different. It also didn't help that Windows kept long 32 bits, which had the effect of breaking code that cast between pointers and long (intptr_t didn't come until C99). The 64-bit ABI for all Unix platforms (AFAIU) carried forward the relationship between long and pointer. I don't think I ever recall seeing Unix or open source code casting a pointer type to int, only long; it was Windows software that presumed pointer->int conversions worked.
We use it for our specialized analysis framework. The real world performance gain is just shy of 30% which is impressive since all it took was a bunch of compiler flags.
Personally I think x32 is a vastly underused ABI. Every application that is unlikely to use more than 2GB should use it. It is literally free performance.
It only caused problems because some kernels used to expect that all physical memory is mapped at all times (and some hardware could only DMA to 32-bit physical addresses, but that's a problem with 64-bit CPUs as well).
Althlon64 was a swift kick in the nuts for Intel, but AMD really didn't followup until Ryzen. I wonder if they're going to stick around and fight this time.
This time, AMD may have stepped up with a steel pipe. We'll see how well they can use it...
:-))))
NetBurst was a bust. IA-64 flopped too. Both were attempts to progress from the P6 architecture - which they felt had reached its limits.
In the end, the market spoke, they want the P6. No one was willing to rewrite software (or even recompile). So back to the P6 in the form of the Pentium M followed by the Core /Core dual series and finally the i3/5/7/9 that we use today - which you would notice hasn’t improved much per core; as mentioned the P6 style design is pretty much tapped out and all they can do is “squeeze blood from stone” for a few percentage improvements here and there.
AMD recently caught up but frankly, they aren’t doing much better. Performance per core is just on par with Intel. Their main selling point is that they made it cheaper by splitting the L3 cache in 2 - the Ryzen is basically 2 quad-cores glued together for better yields.
Not sure which data you refer to? I just made a comparison on a single-threaded integer-heavy code between my old Core2 Duo and a Skylake, and just got x16 normalized perf improvement (from 2.5-cycles/byte to 0.15). Same code, same compiler.
So sure, progress have slowed down and they are adding more and more specialized stuff (AVX-512, AES-NI...) but still.
Edited for clarity.
That way I was compatible with games and most other things without having to deal with increased memory usage and compatibility hacks.
The idea is to expose the 64-bit instruction set and registers, but keep memory and pointers per-process 32-bit as most apps use less than 4 GB of memory.
And yes, before people flare me for this, don't pretend it isn't possible, recall that once upon a time 640K ought to be enough for anybody.
Do you know how much memory we can address with 64 bit?
(64 bit for addresses was arguably a mistake. 48 bit word size would probably have been better, but doesn't sound as cool.)
Having said that, 128 bit can make sense for certain calculations. So floating point calculations and GPUs support long registers for some of what they are doing.
(And having said that, Google figured out that they don't actually need all that precision when using GPUs for machine learning, and made TPUs with much smaller words.)
How else are you going to have more than 16 exabytes of RAM?
And I don't mean that in an absolute sense. People might very well get up to those orders of magnitude of RAM; but you are unlikely to have that amount of RAM available to a single processor.
The extra margin of the 64 bit might help with address space randomization, though.
The concept of having more physical address bits than virtual bits is reasonable, although it falls apart a bit with virtualization. The idea of having magic registers that fill themselves in for you and can’t be read at all (such as the PAE PDPTR registers) is a bad idea that unfortunately repeats itself in x86 design. Architects: don’t do this.
It is well and proper that Intel rounded up to 64 bits. It should serve us well for the next 40ish years, which is good enough for my professional lifetime at least.
It's interesting that you mention hitting that with memory mapping in practice. That is a valid concern.
I am gonna hold onto your quote, for enjoyment and giggles in 2032 :^)
Ie more mobile, and more parallelism on the server side.
But yeah, no 16 EiB on a single processor.
(Though we might see people memory map crazy amounts, without ever actually accessing all of them, of course.)
No more packet-switched serial storage I/O. You now have first-class ability to ask for any byte anywhere, really really really fast.
Because I/O request speed is now only limited by the the memory controller (which already goes at TB/sec in x86 hardware), a fast storage controller now has the opportunity to optimize and batch requests downstream to storage devices much much more efficiently. Because if the storage devices go at a certain speed but suddenly the addressing infrastructure is A LOT LOT faster, your optimization window just went through the roof and you can coordinate much more effectively.
I forget the exact architecture, but one of IBM's 128-bit boxen already does this. Various random bits of the hardware use MMIO as a first-class addressing strategy. The OS does the rough equivalent of `disk = mmap2(/dev/sda)` at bootup. Maybe this is a z series thing.
Predicting the future is difficult. 30 years ago (I am old enough to remember) it was hard to imagine that every household would have 10s of devices connected to the internet. My university VAX for 16 concurrent student users had less memory than required to show the splash screen when today's phone boots.
So if in 30 years the computing paradigm has changed and we directly address memory over the internet? I must admit that in my imagination we have reached a point where growth will slow down. But I have learned that my imagination is not always good enough.
That said, considering the current amounts of data Google holds, I could see the theoretical point about unique addressing every single byte they have. 64bit addressing only allow for a single order of magnitude of growth in that scenario.
And yeah, getting real world software, such as Postgres or Nginx, to work requires some fixes, but it’s really not that bad.