Revisiting 64-bit-ness in Visual Studio and elsewhere (2015)
blogs.msdn.microsoft.com
blogs.msdn.microsoft.com
The author acknowledges that this can happen:
> So, you’re now out of address space. There are two ways you could try to address this.
> a) Think carefully about your data representation and encode it in a more compact fashion
> b) Allow the program to just use more memory
> I’m the performance guy so of course I’m going to recommend that first option.
As a "performance guy", he should know that using a bit more RAM is essentially free, and much, much easier to code than doing the kind of bit-fiddling wizardry he suggests ("encode it in a more compact fashion"). In practice, that kind of low-level code will not get written anyway, at least not by modern Microsoft, and much less so by most extension authors. Which means that there's an arbitrary and very hard limit on how much you can customize/extend Visual Studio before you hit that 3 GB wall. And again, this is on modern workstation machines with gigabytes of unused RAM lying around.
In short, there may be very valid technical reasons why VS can't go 64-bit, but to claim that this doesn't hurt the product is in my opinion not justified.
It's still large relative to VS Code and other editors of course. But the "gigabytes of bloat" problem with Visual Studio isn't so much the case now.
Given VS Code has a fraction of the features VS has, it should be noticeably faster. It will eventually be GB of bloat and slower than VS.
What I'd like to see instead is features being broken out into separate components that can be used by any other editor/tool. Of course, if every component is going to required it's own dot file in my home directory to (maybe) turn off telemetry, I won't touch it, so I'd rather see those things developed by someone other than MS.
I've only seen VS Code run slow if you're running on 2gb of ram, or have hundreds of competing plugins installed, and even then. VS Code is way better than VS proper from my own subjective experience.
Personally, attaching a managed code port to the 64-bit migration sounds misguided. The managed code port sounds much more involved and disruptive than migrating the existing native code would be. It's probably some manager's pet project and it makes no technical sense.
Now with Roslyn and .NET Native, they are planning to increasingly move more runtime stuff into C#.
.NET Native backend is of course shared with Visual C++, aka C2, a kind of Microsoft's own LLVM-like approach.
I think it's good to have additional implementations, even if they are rather insular about it.
It was called Project Phoenix, developed at MSR.
I guess Phoenixey/LLVM type stuff is probably the way the winds were blowing in the compiler community at large. It's interesting how a lot of things that used to seem impossible are now in multiple software packages.
You figure that once MS is rolling out a compiler technology in a big way, it can't get much more mainstream than that.
I still don't totally understand why it was important for MS to push into JIT-compiled bytecode on Windows. From the end-user's standpoint, it just makes software slower.
It doesn’t always makes software slower.
Not all software is CPU bound, a lot is IO bound, and .NET framework implements event-based asynchronous I/O since version 1.0.
Probably the same one that caused the issues with Longhorn.
Now we have UWP with .NET Native, which is creating an OO ABI similar to Longhorn ideas.
Like most IDE's, it is a mediocre text editor with a bunch of compelling (for certain workflows) tooling for project and build management, debugger, etc. Plus usable introspection. As a text editor though, a solid C+ at best.
Is VS Code someone doing something radically different?
Completely open source, rich plugin/extension community. Pretty much has become the editor of choice if you're working on node or go... more than decent with other platforms, and as a general editor. There are extensions to improve workflows, but all optional extensions.
I have git history, git lense, eslint, docker, c# (works for .net core), go, jest, npm, npm intellisense, yandex-translate (makes vscode worthwhile by itself), and a couple others.
I don't think I can go back... having the treeview, git integration, integrated console, it's all just way awesome to say the least.
Sure, Visual Studio is a much much bigger challenge to make compatible. But come on, they've had over a decade during which it was obvious it should be 64bit compatible! A decade!
During the same decade-long timeframe they should have been working on lifting the 260 char path limit, another bug they "wontfix".
My guess is that they've had really strong product management - they really prioritize those snazzy features and releases, but aren't allowed to spend any time on the technical debt. (And, seriously, they've had a decade ...)
[0] https://blogs.msdn.microsoft.com/dotnet/2016/08/02/announcin...
As I recall, this was marked "wontfix" in 2013: https://visualstudio.uservoice.com/forums/121579-visual-stud...
It's been _possible_ using the NT Unicode APIs since Windows 2000: https://msdn.microsoft.com/en-us/library/930f87yf.aspx
But other common first-party windows development tools can still have the problem: https://github.com/Microsoft/msbuild/issues/53
Was really happy to see it was finally fixed in .Net and the Anniversary update, and long, long overdue.
I wanted the 64-bit transition only so I could properly use my tools at work (used to work in a .Net shop). We had resharper and several other plugins that ran in the same Visual Studio process as everything else and it was fairly common to hit the memory limit of a 32-bit process and Visual Studio would essentially die until it was killed and restarted.
Sure, you can say "stop using those tools" or "they should have written them better". But at the time I was required to use most of them (it wasn't only Resharper though Resharper was pretty nice).
Today I don't have this problem anymore. But I think because of that type of issue it would still likely be worth it. Honestly I feel like they could optimize Visual Studio at the same time; it has a ton of capability but it's also incredibly large and, from my understanding, carries a TON of legacy code and resources throughout.
- You really only get 3GB of RAM
- Actually, you get less than that, because of DLLs that get loaded into that space.
- Actually actually, fragmentation becomes a problem, and making very good use of the remaining memory gets awkward pretty quickly.
3GB is pretty crowded, especially when you're talking about (say) a game with tens of GB of footprint. You need to page stuff into that footprint, and be clever about the memory pressure not affecting the user experience.
On the server side of things, we regularly run processes with > 8GB footprints, including things like solr (at 32 to 160GB). Breaking this stuff up would involve a lot more disk and network chatter, as well as bugs involving OOM conditions, reducing global performance and reliability.
So while VS may be fine with 32-bit code+data (I am not convinced), real-world applications definitely need more. I'm guessing that making a 64-bit VS is hard for legacy reasons, and that the 32-bit space is actually holding the VS team back (and possibly plugin makers as well).
Many people in this thread and the corresponding one on reddit report VS regularly crashing on OOM or eating all the CPU garbage-collecting nonstop as it gets closer and closer to the limit.
I agree that it would probably be nice if those could be turned off, and the UI is sluggish, but VC6 was (apart from being fast) not much to write home about (compared to what we have now).
Anyway, this is is just 1 contrived example of something modern OSs do, and that OSs from the 90's didn't do.
Sure, there's some bloat, but a lot of it is the "Mozilla kind": "Mozilla is big not because it is full of useless crap. It is big because your needs are big".
Just booted my laptop (FreeBSD -current amd64, 1366x768 display) and started X: only 204M "Active" memory. Of the 204M: 60M is syncthing, 39M compton, 31M Xorg, 11M i3bar, 11M polkitd, 10M i3, 9M dunst, 5M wpa-supplicant… Not counting syncthing that's 144 megabytes. I think that looks reasonable. You can go lower and optimize for low memory (no polkit, no compton — 94M) or go higher and optimize for usability and fancy UI features (install gnome :D)
Let alone USB. During my training, there were lots of desktops running Windows NT 4.0, which did not know USB. By the time I had already become so used to copying data back and forth with a USB thumb drive (even though back then their capacity was still measured in megabytes), that it became fairly annoying to walk up to some computer, plug in the thumb drive, then realize this machine is running NT 4.0, not Windows 2000, cursing, then looking for another way to get some piece of data on that computer: Usually put it on some other machine copy it over the network.
Ah, those were the days... ;-)
This doesn't add anything, because contrived or not, it doesn't answer the question that was posed: "What does modern Linux do that Solaris from the 90's did not, that it requires 50x more memory?"
It hits the first part of the question, but not the second part. We could add wifi to a 90s OS without severely inflating the memory requirements. My Nintendo DS (with 4MB RAM) "seamlessly connects and disconnects to wireless networks", after all.
You could easily answer the question with a sensible, non-contrived answer, and I'm not sure why you didn't.
I recommend everyone to try and build an LFS (Linux From Scratch) at least once or at least set up an ArchLinux box to get a feel for how Linux actually works and where the bloat is.
As an example, the music player daemon (MPD) has 110 dependencies among which is wayland or x-org, I don't know why.
Also, using udev to only load needed kernel modules helps with a lot of memory usage.
Also, if you'd like to compare performance degradation over time, NixOS is a good choice to run because it can simply rollback to old configurations and you can see what changed.
I have small Linux systems today that work comfortable with 128MB of RAM used.
What my modern linux system does that Solaris didn't do in the 90s: runs a browser than can render absurdly detailed scenes using OpenGL where the object model itself exceeds 1GB (it fits nicely on the graphics card, but the computer is also loading that from the net at gigabit speeds and storing another copy in RAM).
That said, I think it's true that modern systems just waste oodles of resources unnecessarily.
It's salutary to consider that the memory for a reasonably spec'd machine a little over 20 years ago is now a rounding error
I remember at about that time having 80MiB of RAM: a) my machine flew, and b) plenty of people would ask why I needed so much
When we have less, we make due with less. When we have more, we consume more. It's sort of like money.
In short, due to VS having an already large working set it hurts more to let it grow even larger than having more registers helps.
I also think that many people took his stance as »32 bit should enough for anybody and no one should move to 64 bit«, whereas it was more a »performance characteristics for different types of applications are very different and for Visual Studio it's a choice that will make things slower«. Admittedly, that doesn't appease the people who experience crashes, although extensions like ReSharper are often more wasteful with memory than the IDE itself and those could just as easy be put in a separate process. Although for JetBrains it's probably a marketing argument to push ReSharper users to Rider (which actually uses an out-of-process model and thus not have the problem).
I do, however, buy people's arguments that they are feeling real pain by the memory limitations. Even if 32-bit is faster at the moment, I think it's a short-sighted decision to not port to 64-bit. First, the memory issue that everyone has brought up. Two, I find it likely that over time, using the now-standard bitedness (I know that's not a word, forgive me, I have a cold and mind is foggy) will yield benefits because it's the common path in the processor.
(ygra, I know you're not necessarily in agreement with everything Rico said, so I'm not really arguing with you.)
It's not "intel", it's x86. x86_64 has twice the number of registers.
The "typical ALU instructions" as add, cmp etc. can now encode 16 registers instead of 8. But the FPU stack still has 8 entries to encode and there are also still only has 8 MMX registers in 64 bit mode (they overlay the FPU stack as you surely know).
On the other hand with AVX-512 there will be 32 xmm/ymm/zmm registers available instead of 8 in 32 bit mode (4 times).
UPDATE: On the other hand in 64 bit mode there are only 2 segment registers available instead of 6 in 32 bit mode (OK, the 4 common ones were (IMHO unjustifiably, since there are some cool things that you can use them for if you know what you do) not used in modern 32 bit OSes (Windows NT, Linux), so they were left out).
TLDR: The multitude depends on the type of register that you consider.
I know that if you implement an algorithm via SSE/AVX it typically is much faster than if you use the FPU. But I still believe that the FPU has its uses: For example it also supports 80 bit precision floating point numbers, while SSE/AVX only support 32 or 64 bit ones. There are applications where this capability can be quite useful. This is one reason why the FPU is still supported (and sometimes used) in 64 bit mode.
If you look at generic floating point code generated for x86 by any modern C/C++ compiler, you won't see any FPU use, it's all SSE scalar, and occasionally SIMD for clever compilers.
Here are links about this topic:
> http://stackoverflow.com/a/35619528/497193
Hell, why not stick with 8 bit? We can just optimize everything to work on that, right?
Look at GUID partition tables. With MBR, we had to hobble along with only a byte to identify partition types. Now we have 128 bits. We can finally support more filesystem kinds now than there are atoms in the Sun, all in one system installation, and the bootloader just has to look at the GUID. A 32 bit "fourcc" partition label clearly wouldn't have been enough.
A 32 bit "fourCC" would be more than adequate. It could even be constrained just to readable characters, like LNXF (Linux Filesystem) and LNXS (Linux Swap).
If another OS happens to use LNXF for something, and you have that OS in the same darn system, it doesn't matter. You just don't have that "foreign" LNXF in your /etc/fstab, and likewise it doesn't have the Linux ones in its equivalent of /etc/fstab.
The only thing that needs a clash-free label is the EFI boot partition, so the boot firmware can unambiguously identify all these partitions on all attached devices and offer them as boot options.
If the argument you want to make is that they are of no consequence, then you need to answer why your proposal of still doing something is warranted.
Why not a nibble and just use the other 4 bits for flags?
The point of this line of questioning is you're quibling over bytes which definitely don't matter in any modern context, at the expense of masses of extra management complexity to try and avoid day-to-day problems when people want to stand up new systems.
With GPT, if I want to make a new filesystem type for some application, I just generate a GUID and it will not collide without me needing to coordinate with anyone.
16 bytes is an obvious example of the "second system effect" described by Fred Brooks in Mythical Man Month.
The fdisk utility now reduces the GUIDs to one byte codes that refer to the GUIDs. For instance, I remember that 29 is Linux RAID (previously FD). Will 29 always be Linux RAID everywhere? Probably not.
> With GPT, if I want to make a new filesystem type for some application ...
Four bytes could have an ample reserved range for local use by hobbyists.
Broad recognition of the code only matters if the application is very widely deployed.
Wrong (for x86-16 vs. x86-32). Just use an operand-size size override prefix (0x66) with your 16 bit real mode ALU (in this case 'add') instruction to make it a 32 bit ALU instruction. Works from 80386 on, where the 32 bit registers were introduced.
The one thing I found absurd with RISC-V is the 128bit variant. Most 64bit processors today don't even support a full 64bit virtual address space do they?
Except when they don't. Everyone already forgot tweet number 2147483648? :) https://techcrunch.com/2009/06/12/all-hell-may-break-loose-o...
My interpretation of the shift from < 32 bits into 32 bits is: before we had do do crazy things to algorithms we used to fit in those address spaces. When we transitioned to 32 bits, we didn't have to do that anymore.
So the question might be are there any surprising workarounds in the code because you're only dealing with 32 bit code where if you had 64 bits you could write some more elegant solution.
The only other example I can think of is the general "problem" of large databases. There's just a lot more paging and churn that has to happen in a 32-bit address space. Many NoSQL databases in particular have a memory model of mmap'ing an entire database, which runs into a hard limit on 32-bit address space.
In a 16-bit address space (64K), you hit the 16-bit limit _all the time_. Even a moderately sized text document will be bigger than 64K... and that's before considering images, videos, large data sets, etc.
32-bit takes you out to 4GB, which is much more likely to hold a typical working set, so the argument to go to 64-bit is much less pressing.
"why not stick with 8 bit?"
This conversation is about the size of the address space, not the machine word size. To my knowledge, there were no serious machines of any sort that were limited to an 8-bit address space. (Maybe something homebrew or embedded.) The closest I can think of is the 6502's preference for putting values in the zero page (which was 256 bytes).
Depends on the way 16 bit is implemented. For example x86-16 uses segmented memory - enabling adressing of a little bit more (including High Memory Area) than 1 MiB of memory. The Z180 uses as far as I know a MMU (but not completely sure). Another approach that is/was in common use is to use bank switching.
Depending on the kind of algorithm that you use this can make the coding much more complicated or can also be no problem, because the scheme that is used to address more memory than 2^16 bytes fits the algorithm quite natural.
One interesting hack for example when coding in real mode (x86-16) that I read about is rather to use some clever sharing of bits between the segment register value and segment index:
- One scheme is to consider the value of the used segment register as a pointer to a 16 byte block of memory and use the segment index to adress the specific byte in this block (with an option to increase the index "a little bit" if you want to go further)
- Another scheme is to (mostly) use only the 4 highest bits of the segment register (and zero all the other ones).
This only works in real mode, where the segment is shifted and directly added to the offset. In protected mode, it goes through a selector table. This can be made to work too, but it requires tiled allocations of segments with known delta between each segment. This is what __ahincr was about, if you remember it from the Win16 days.
Of course this is true. You wrote further above:
> In a 16-bit address space (64K), you hit the 16-bit limit _all the time_.
I wanted to outline that whether this is a problem or not depends a lot on the concrete 16 bit architecture.
I'd be interested to hears... I really can't. Probably the most capable architecture I'm familiar with that has a native 16-bit pointer type is the 80286, which provides 24-bit physical addressing, virtualizatoin, protections, etc. Even then, at least on Windows, key local heaps within the OS were confined to 64K, the default text editor was confined to a segment... writing image processing tools (which I did) required special handling for all but the smallest scale workloads. These all added to developer workload and reduced system capacity in unfortunate ways.
I get what you're saying that there are exceptions and hacks that make it possible to work within these limitations, but my point is that you have to care about them on 16-bit, and most of the time you really don't on 32-bit. My thesis for why that is goes back to what I was saying initially about the size of the core data types people tend to manipulate.
(And this goes back to the reason I posted in the first place, which was to explain the relative difference in motivation between the 16->32 switch and the 32->64 switch.)
Your question implies that you want one piece > 64 KiB of flat memory. With this you already stated a very strong implicit assumption about the data layout and the kind of algorithm that you want to use. My point rather is: Consider the capabilities that the 16b machine has and try to fit the data representation and algorithms that you use around it instead of wining about lack of machine capabilities. One will often find a solution using clever tricks that one would not have considered otherwise, which will often turn out to be surprisingly elegant and much better than the "naive" solution.
This way to program is of course nothing for the kind of programmer that want to write an unelegant and just working program in a short amount of time, I know. :-)
I do hear you, but one of the things I love about virtually all modern hardware is just how much capability it puts within the reach of totally naive and quick development strategies. If I can inelegantly solve two or three problems in the same time I can elegantly solve one, then that strikes me as a net win (at least for the people that need problems solved more than code written.)
I don't mind pushing the hardware and searching for elegance, but I'd rather be forced into it by the necessities of the problem I'm trying to solve.
Can someone explain why exactly 64-bit is generally slower than 32-bit?
I understand that more RAM will be used and I/O to it slowed down due to double the bits pushed around since "chunks" have double the length, which ends up being a lot of empty padding (is that correct?).
But everything inside the CPU, like registers or ALUs, are 64 bits wide anyway (right?), so computing in 64-bit mode would just make use of resources that were unoccupied in 32-bit mode. Or am I missing something?
Note this can also impact alignment of fields within structs--if you're on 64-bit, you want your pointers to appear before 32-bit fields; if you alternate, you waste a lot of extra space.
The same is true for L1/L2/L3 caches too.
The same is true to a much lesser extent for taking up more space in ram, and to an even lesser extent for reading things from storage.
It's an incorrect yet popular assumption on the internet.
64 bits pointers take more space, which affects caching => Slower
64 bits architectures have more registers and instructions, and it removes the need to have an intermediate 32to64bit abstraction layer from the OS => Faster.
Overall, the side effects are extremely complex.
Moving an application from 32 to 64 bits is as likely to be 1% faster as to be 1% slower.
About the only real benefit is that the system can handle more RAM (which of course is a good thing). For some applications that are very, very memory-hungry, this can improve performance dramatically, because the program no longer needs to juggle all that data and can just keep it in RAM.
This only applies to running a 64-bit kernel vs a 32-bit kernel; running a 32-bit process on a 64-bit kernel will incur the same page table cost as a 64-bit process.
https://tomssl.com/2015/03/31/always-run-visual-studio-as-ad...
I mean realistically don't know how much of a concern this is, but it's conceivable you might be running Visual Studio with elevated permissions on a shared machine and in turn want the ASLR protections.
https://blogs.msdn.microsoft.com/ricom/2016/01/04/64-bit-vis...
...In the VS space there are huge offenders. My favorite to complain about are the language services, which notoriously load huge amounts of data about my whole solution so as to provide Intellisense about a tiny fraction of it. That doesn’t seem to have changed since 2010."
But I'm sure that by remaining 32-bit, in another 5 years maybe everyone will magically optimize their stuff... Any second now...
Is this something VS could do?
I have the greatest of respect for the author, but everyone needs to be exposed to 'trust, but verify' at least occasionally, just in case they happen to be wrong this once.
And no, as any C++ developer will tell you, the answer is not to move to a souped-up text editor (Visual Studio Code).
Makes sense if you're running Java with -Xmx1024m
If I've got a box with 512GB of ram, it seems I'm supposed to spin up multiple instances to satisfy this, all because the JVM has a hissy fit if you go over ~30GB. This then means worrying about replication and ensuring we don't have both the primary and replicas sitting on the same box.
It seems insane that this is an actual issue in 2017.
What is the actual technical reason why the JVM cannot (easily?) address more than 32 GiB of RAM?
Except once you breach the 32GB limit, your 32 bit pointers grow to 64, and depending on your application you might need to grow your maximum heap into the high 40s to get room for new objects: https://blog.codecentric.de/en/2014/02/35gb-heap-less-32gb-j...
https://www.elastic.co/guide/en/elasticsearch/guide/current/...
There is also a limit for PDB files of 1GB, which can be workarounded to 2GB, but from the error above I think you are not hitting this limit.
?
> http://mash.wikia.com/wiki/Sherman_T._Potter
horse huckey:
> http://www.urbandictionary.com/define.php?term=horse%20hocke...
Colonel Potter curses (“Horse hucky!” is at the beginning):
Wait... did you just tell me to go fuck myself? ;)
The problem is probably that memory is crowded and plugins must do a lot in-process via COM.
If Microsoft really wanted to make VS fly for a lot of people they could just add the 10% most popular features from the 10most popular extensions. It would want nothing more than to throw out R#