Adventures of putting 16 GB of RAM in a motherboard that doesn’t support it
downtowndougbrown.com
downtowndougbrown.com
I hate stuff like this. What's the point of hiding the parameters, instead of tucking them away in the corner of the screen or something? It just makes things needlessly difficult.
I know the answer is going to be "but it makes the UI look beautiful and most users don't know what the parameters mean anyway!" Error UIs don't need to be beautiful, they need to be functional. Definitely have a nice, clear user-friendly error message, but keep the overall screen technically informative. I can imagine that the only thing more infuriating than having a computer that won't boot, is having one that won't tell you the problem unless you change registry settings (which you can't, because it can't boot). Leaving the extra information there does not harm, and it give people something to Google if they do run into a problem.
When app crashes Android simply tells you that "Unfortunately, App has stopped". Xiaomi UI additionally presents you the stack trace.
BTW: Android used to (?) have "Report" button, but it seems it's gone on new versions (Since 5.0?). Anybody knows where these reports were send? Probably to Android OS team, just like MS sends some crash infos of 3rd party apps to themselves.
The Report button is still there - maybe it is not shown if there is no one to report to? For apps on Google Play, the user reports can be seen in Google Play Developer Console.
Yeah, I was a bit surprised by that also. Even when turned on the extra debugging info isn't exactly front and centre, why not keep it there.
Arguably linux is the same - you can't edit files from grub - you need a bootable system to use vim.
When raising exceptions in code, I make sure the error message contains as much relevant detail as I can stuff in there. Users will just mail a screenshot to support regardless (if they want help before the full error report gets home), so might as well give as much info as possible to make troubleshooting easier.
https://en.wikipedia.org/wiki/File:Windows_NT_BSOD_at_GVA_ba...
Love to smack the goof who decided a giant sad face is more helpful than dumping even some minimal diagnostics fine print.
They need to be both. If it can’t be both, there’s as many arguments for it to be either human friendly or actionable.
I think it comes down to it being a human/machine interface, and what the human side does with it. If for more than half of the users the next step from the error UI is to rage phone their IT department, and a sympathic UI might actually stop them from doing it (or at least be considerate), I’d argue a friendly UI should be prioritzed over an overly informational one.
If it can be both, that’s better, but I wouldn’t be surprised if after user testing it appears it’s just damn hard.
Also having an error displayed on screen that can’t be retrieved any other way (logs, diagnostic dump on a console somewhere else, whatever) is a design choice I’d loathe way more than not having the error number on the BSOD.
In Windows, it would just spin a little circle forever until I rebooted. After installing Linux and getting actual error messages, I learned that my SATA AHCI controller wasn't always working at boot and was fixed by plugging my HDD into a different port.
The problem could have been fixed so much earlier had Windows just told me what was wrong.
It's been so long that I can't tell an e820 map from an MTRR, but this was still a fun read.
I'd love to see this code; doing memory management entirely in the OS and ignoring the BIOS sounds like fun (modulo working around BIOS-specific memory reservations so the OS doesn't get its memory stomped on by the BIOS).
I spent some time last year running various scripts to get an NVIDIA GPU working over thunderbolt in windows on a macbook pro. The problem is the DSDT table in many macbooks doesn't allocate enough space for pci-e devices. The NVIDIA driver in windows tries to allocate memory via that table to talk to the eGPU and it fails.
For some reason it works fine under MacOS - either macos ignores the DSDT table completely, or it allocates memory a bit differently than windows. In any case, the answer is to use obscure tools to download and patch the DSDT tables to allocate more RAM toward PCI-e. Doing this through UEFI feels very magic.
I don't know why so much of this stuff is up to the hardware vendor. Maybe there's a good reason, but I would expect the windows memory manager could do a much better job if it didn't have a bunch of memory range sizes hardcoded by apple.
Kids today...
In the case of RAM, "physically impossible" would mean something like the CPU not having that many address bits or the PCB traces not routed, but there are a lot of different not-well-advertised configurations of DIMMs available with the same size[1], so it could be that the manufacturer specified a lower maximum just to avoid having to answer subtle questions. For example, 2x8GB may be OK but 1x16GB not.
It could also be a "X if Y except Z else B" situation, and they just couldn't be bothered to document all the possible combinations and/or explain the details of the PC system architecture that result in such limitations.
To test if all the memory is present (and if it all works correctly), running something like MemTest86 might be sufficient.
More like:
from "physically impossible" to "we haven't tested that"
> they just couldn't be bothered to document all the possible combinations
Each combination adds exponential testing and documentation requirements.
Not tested != doesn't work. But that doesn't make not supporting it malicious. There are practical and financial cost to testing every combination, ultimately born by the consumer.
Can be:
- Something that will be rather obsolete soon, so it won't be tested
- Something that usually works, but has poorly understood corner cases or other implications
- Something that is so infrequently used that it's not worth it to build up proper testing infrastructure to prevent regressions
In this case, who thinks Foxcon has an engineer on call, who's familiar with the motherboard, who can actually answer this question? I personally doubt it.
But he wanted more. So, he purchased a pair of thick memory expansion board that brought it up to an unheard-of 640KB of RAM.
Of course, these slabs generated an impressive amount of heat, so we had to set up a series of cooling fans, or else computations would go awry and weird bugs would appear in programs as it overheated.
It had all manner of strange behaviors, and IBM engineers on the support contract would dutifully visit and give us stacks of floppies containing custom builds of MS-DOS to try and help us out when he called about problems with his mods.
Today, incidentally, I’m a sysadmin.
I would guess not? But then IBM were supporting it???
OT. I now have visions of young engineers joining IBM with dreams of working on big iron, but ending up in the burbs fixing PC jrs.
> All I know is it passes my RAM tests. Since Linux has been working fine with the 16 gigs of RAM installed, I am not too worried. It’s possible I would have problems if I had more PCI/PCIe cards installed or something, but in my use case, it seems to behave fine. Obviously, your mileage may vary
I run into this so often. "Look, it works for me. Don't try it at home. Who knows if it keeps working? What can I say I'm going to ride it until the wheels fall off".
Said on the last day to the new hire.
https://www.itprotoday.com/compute-engines/windows-nt-and-vm...
https://www.amazon.com/Showstopper-Breakneck-Windows-Generat...
EDIT: fix typo
:)
Toward the end of the Bell system, AT&T used Unix to run the control plane of phone switches and also to do administrative tasks for the phone network. I remember a paper in the Bell Systems Technical Journal where they did auditing of billing records for the whole U.S. with a set of tools kinda like grep, awk, etc. but with binary-format data records.
There are situations where operating systems get behind the 8 ball and there isn't a perfect thing to do. Giving up to prevent data corruption is one choice, but trying to soldier on and do the best you can is another.
Lists of recently switched processes, or the general category of errors it falls in.
Already this excludes most people who will see the error. Meanwhile, experts already have more sophisticated tools available to them like the event viewer. Obviously that wouldn't be useful in a situation like this where boot-up is blocked, but like I mentioned previously, there's only so much diagnostics you can reliably provide on a system that is in the middle of crashing.
So both sides have to be able to see the details.
Nobody is decoding QR codes by hand, so it isn't that big a deal to go from "https :// www.windows.com / stopcode ? code=ACPI_BIOS_ERROR" to "https :// www.windows.com / stopcode ? code=ACPI_BIOS_ERROR & p1=0000000000000002 & p2=FFFF9A0..."
I rather have a detailed error report in one language and then have to find someone who speaks it to translate it (if Google doesn't do a good enough job), than fifteen poorly written versions. Not to mention, those error screens should be as minimal as possible in terms of features so less goes wrong.
I don't want animations, fancy colors, dealing with the horror show that is localization, or anything of the sort. Because the system when it hits a bsod is in an undefined state, so it's best to exercise as little as possible of the system, just enough to get the error message out.
If it is for the user then being able to google the exact error is going to return more results than 47 translations.
Most users will in fact do nothing but report said error to their vendor whom will refer to their English language documentation.
In other news most programming languages aren't localized for example.
And they already provide just enough information to be able to google the error. I am just opposed to adding more detailed information, the kind of information that's only relevant for experts and is not guaranteed to be available in the middle of a crash.
> In other news most programming languages aren't localized for example.
Programming languages aren't consumer products like operating systems are
It should certainly be possible for someone who can diagnose an error to see the details, but making it the default can be problematic when the message has a chance of being shown to real end users.
I think normal users have a pretty good idea what to do with it -- show it to the nearest tech person or paste it into Google.
Just giving it to them up front makes things easier, not harder. They need help so they take a screenshot. If it contains the relevant information, someone can help them. If it just says "an error occurred" then nobody can help with that because it contains no information. Now they have to go back and have the tech walk them through trying to extract the information from some log file, which is seventeen clicks through an interface they've never used and now there are 5000 entries and they don't remember exactly what time it happened etc.
It also deprives the user of any opportunity to learn. The user has a problem, they get a message, they ask the tech what to do and the tech tells them. Now the next time they get the same message they know what to do. But if the message is always the same regardless of the cause then they lose the ability to match known problems with known solutions and give up.
Not all users care to learn that sort of thing, but you take it when you can get it, not purposely inhibit the user from doing it.
I suspect the modern trend of making problems opaque comes from companies that do it on purpose. If the user can solve their own problems, what do they need with your expensive support contract? Why have the user fix problems with their existing device when you can sell them a whole new one the first time anything happens?
And then other developers who don't even use those business models still cargo cult the same UX.
I was not writing code at the time and had nothing to contribute.
So apparently they fixed it later.... then came the security issues, unexpected behavior, shit going sideways.
Sometimes a straight crash is not such a bad thing.
One of the things I like about Kubernetes is that the ecosystem (generally) tries to adhere to the "Not healthy? Then crash and keep crashing" mantra when something doesn't work. If I see something is in a CrashLoopBackoff I at least know its b0rked. Stuff that reports its up and running when it's actually hosed is really annoying.
The worst sort of thing to track down is working ... but not. God knows when it started or what it has impacted.
Worse if it impacts real data or data goes where it shouldn't.
Linux needs to come in after the fact and run on whatever garbage happened to ship.
(FWIW: in this case the root cause was a host bridge in the tables which had been granted a truly outrageous memory space despite having no devices in it. It was likely a typo, or some test stuff that got left in.)
You're assuming this was an issue. Manufacturers religiously create artificial tiers in software to upsell users. This is true of hardware just as much as it is software.
(It also took some Perkele by Linus to make it work correctly)
Perkele means evil spirit in Finnish and is a popular Finnish profanity.
For example, I have a devboard that only support 1 GiB of DDR2 RAM, but it's a 64-bit system and the memory controller on the CPU was supposed to support at least 2 GiB of RAM. Meanwhile, another board that uses the identical chip runs 2 GiB RAM without problems.
The engineers of the devboard briefly explained that the problem was electrical. The memory controller itself has inadequate drive strength, adding more RAMs would increase the load on the DDR bus and destabilize the system. On the other hand, the other board had better PCB layout so the problem did not occur.
So it seems memory is a general problem among Mini-ITX boards? Perhaps the reason is that these boards have less available space for routing, fewer layers, and targets a lower price, so they trend to have worse electrical characteristics?
The BIOS does a (relatively) quick memory check in the POST to detect how much memory is actually available, basically by writing a series of patterns to all addresses and then reading them back to confirm; some desktops have a "fast boot" option which mostly skips it (I believe it's something like testing one byte per 4KB instead of every byte), and servers usually have a much more thorough test that can take many minutes.
The best way to check whether the memory is functional when 16GB is detected is to run a memory tester like MemTest86.
The Synology web UI only shows 8 GB and it breaks the graph display in the memory usage monitor (although when I first installed it, it used to show 16 GB but after some update, or random reboot, it stopped. Haven't tried rebooting to see if it's just random)
https://hardforum.com/threads/hp-proliant-microserver-owners...
> It seems to be confirmed, that these wonderfull little boxen can actually support 16Gb Memory, not the 8Gb HP Says they are limited to.
So it would appear that HP doesn't officially support 16Gb, which puts us firmly back in the territory of this post, YMMV, etc.
It was not uncommon to have a small boot partition at the beginning of the disk for that purpose.
[1] Bit of a simplification. In reality, the BIOS being 16bit real mode code meant that you had to jump through very elaborate hoops if you wanted to use it in your protected mode OS in any way, and then for questionable gain.
- I'm using CompactFlash as many do, since it's fast, cheap, and reliable. Some of my CF cards are smaller than 8 GB, and if I hard code it I can't boot with those.
- I actually do run OSes on that machine that thunk out to BIOS (Win98 for example.) If I really wanted a computer for useful things, no doubt I'd run Linux on it, but all the same, I've got plenty of smaller computers with unilaterally more power (even a RPi is much faster.)
(FWIW: autodetect properly detected the right parameters, it just locked up on boot. My guess is it was some kind of simple integer overflow bug or something.)
still had to that a few weeks ago because on a dell r720 grub did not want to boot my zfsonlinux rootfs on a 6tb pool - it's either the dell hba controller, grub oder some other limitation but once you go beyond 2tb or 4tb disk I always run into strange behavoir.
Maybe UEFI solves that, I don't know.
Western Digital called their program EZBIOS, see if you can scrape a copy of that up somewhere.
The Lenovo T480 supports 64 GiB. A 13" hackintosh-friendly laptop with:
- dual m.2 slots
- WQHD 2560 x 1440
- Thunderbolt 3
- water-resistant keyboard with drains
- 9 hours of battery life with the second, extended battery
- officially user-serviceable parts/guides
- & 64 GB!
The iPhone 6S, even with the headphone jack, is IP67 in all but name but didn't sell it as a feature... it has all the gaskets of the 7 but supposedly the headphone jack was an issue... mine's been in the shower a few times for YouTube morning news and still works.
You can make any data & charging connector USB-C magnetic & right-angle for $30. https://www.aliexpress.com/item/USB-C-L-Tip-Magnetic-Adapter...
macOS and most apps have nice UX and mostly work well together. Docker usually works on top of the built-in macOS intel-origin hypervisor. (VMware Fusion is another $olid option for $$; had problems with VirtualBox.)
What's neat about the T480 is there's an utility on the Windows partition to flash the BIOS startup boot logo from red "Lenovo" to whatever you want.
Intel, This is the reference, most operating systems use this (linux, freebsd, macos)
Microsoft
Openbsd, yes the openbsd project is crazy enough to build their own acpi implementation, keep fighting the good fight openbsd.
Because have you seen the spec? good lord what a train wreak.
A somewhat disappointing trend in user interface design…
I was expecting this to be "Linux roolz, Microsoft droolz," but I really enjoyed this. It's been a while since such simple, fun hardware exploration came across my desk.
It's not a RAM issue though. AFAIK there's some bug in the UEFI/BIOS even on the latest version which Windows handles just fine but Linux randomly doesn't. I can boot Linux every time if I disable ACPI at boot but then I lose power management and battery percentage. Latest kernels (can't remember the version exactly, 4.somethingoranother) should handle the issue according to some forum posts I found but it doesn't for me.
Running memtest it was clear that there were issues with accessing ram in various segments, but I didn't have the knowledge to muck around with the acpi tables to see if that would fix it.
I'd love to give it another shot if possible.
My every day workstation is a Dell Studio 540 from 2009 that I've upgraded to the max in every possible way but without changing the motherboard and case.
The only upgrade I haven't pursued (because I thought it was impossible) is bumping up the RAM beyond 8GB. If anyone has experience bumping 2009 era Dell motherboards (in particular the M017G motherboard) beyond 8GB RAM and has any tips to offer, please let me know!
If I could bump up the RAM to 16GB, I might be able to keep using my beloved desktop for another 10 years :)
In theory (!) for older generation chips like yours and the one in the article, the maximum listed there (if present) will be correct.
For example, for the i5-750 listed in the article it shows 16GB maximum, which the article author was then able to make work:
https://ark.intel.com/content/www/us/en/ark/products/42915/i...
If yours shows 16GB as well, you may be in luck. :)
It works!
I ordered 4 x 4GB 2Rx8 PC2-6400U DDR2-800Mhz 240PIN UDIMM from a business on eBay called Memory Masters and it worked perfectly. No hackery required.
My desktop from 2009 now runs with 16GB RAM and will be hopefully be in use for another 10 years :)
> All I know is it passes my RAM tests.
I guess you could allocate a humongous matrix with random cell values and do a CRC or something (once while you're filling it in, and again from square 1) to make sure the values are indeed what you think they are.
I remember custom flashing a motherboard to support a pair of SSD's in RAID-0, and allow it as a boot device.
bug report?
If there is problem with reported values from a BIOS, it's a good thing Windows does not continue. BIOS issues (such as a RAM address out of the reported range) can lead to unpredictable behaviour. Troubleshooting such issues is next to impossible.
For a second I thought you might say the contrast was too low, and which CSS style you had to adjust.