Building a personal, private AI computer on a budget
ewintr.nl
ewintr.nl
The workstation was refurbished for just over 600 bucks, and another 120 bucks for the GPUs and another ~60 for the fans.
Edit: and before someone asks - no I have not uploaded the STL's anywhere cause I haven't had the time but also since this is a very niche use case, though I might: the back(exhaust) bracket came out brilliant the first try - it was a sub-millimeter fit. Then I got cocky and thought that I'd also nail it first try on the intake and ended up re-printing it 4 times.
Well, for a dedicated LLM box it might be feasible to suffer with drivers a bit, no? What was your experience like with the software side?
I thought the problem was that those cards have loads of RAM but lack really important compute capabilities such that they're kind of useless for actually running AI workloads on. Is that not the case?
it is - they're laughably slow and not even supported by latest CUDA
> NVIDIA Driver support for Kepler is removed beginning with R495. CUDA Toolkit development support for Kepler continues through CUDA 11.x.
friend you shouldn't make comments like this unless you understand the definitions of the words. Deepseek wrote some parts of their kernels using PTX. newsflash: PTX support for features is lockstep with CUDA support for the same features ie the fact that CUDA doesn't support it means you couldn't write the PTX to use those features either.
Thank you for providing the information to clear up ignorance though.
> is deepseak's use of PTX instead of CUDA relevant here?
this is a conclusion/assumption thinly veiled as a question
> Deepseek R1 doesn't use CUDA, so ... it isn't a big deal?
note, genuine questions don't already presuppose an answer.
You don't get great out-of-the-box performance but it only took me three work days or so with no experience writing these to adapt, test, and validate a kernel using the acceleration hardware that was available (no prior experience writing these kernels).
They're not as powerful as others but still significantly better than running on a CPU alone and I'd bet my kernel is missing more advanced optimizations.
My issue with these was the power cable and fans. The author touches on the fans and I did try a 3D printed shroud and some of the higher pressure fans but I could only run the cards in short stints. I ended up making an enclosure that went straight out of the case using two high pressure SAN array fans I harvested from the IT graveyard per card and making a hole with an angle grinder.
The power cable is NOT STANDARD on these. I had to find a weird specific cable to adapt the standard 8-pin GPU connector and each card takes two of these bad boys.
I've heard arguments both for and against this, but they always lack concrete numbers.
I'd love something like "Here is Qwen2.5 at Q4 quantization running via Ollama + these settings, and M4 24GB RAM gets X tokens/s while RTX 3090ti gets Y tokens/s", otherwise we're just propagating mostly anecdotes without any reality-checks.
Still, it’s way to early and there are simply way to many hardware and software combinations that change almost weekly to establish “the best practice hardware configuration for training / inferencing large language models locally”.
Some day there will be established guides with solid. In fact someday there will be be PC’s that specifically target LLMs and will feature all kinds of stats aimed at getting you to bust out your wallet. And I even predict they’ll come up with metrics that all the players will chase well beyond when those metrics make sense (megapixels, clock frequency, etc)… but we aren’t there yet!
Saying "Apple seems to be somewhat equal to this other setup" doesn't really contribute to someone getting an accurate picture if it is equal or not, unless we start including raw numbers, even if they aren't directly comparable.
I don't think it's too early to say "I get X tokens/second with this setup + these settings" because then we can at least start comparing, instead of just guessing which seems to be the current SOTA.
But you can likely find similar threads for the llama.cpp benchmark here: https://github.com/ggerganov/llama.cpp/tree/master/examples/...
These are good examples because the llama.cpp and whisper.cpp benchmarks take full advantage of the Apple hardware but also take full advantage of non-Apple hardware with GPU support, AVX support etc.
It’s been true for a while now that the memory bandwidth of modern Apple systems in tandem with the neural cores and gpu has made them very competitive Nvidia for local inference and even basic training.
Still, thanks for the links :)
* hardware spec
* inference engine
* specific model - differences to tokenizer will make models faster/slower with equivalent parameter count
* quantization used - and you need to be aware of hardware specific optimizations for particular quants
* kv cache settings
* input context size
* output token count
This is probably not a complete list either.
What's hard about it? You get the hardware, you run the software, you take measurements.
But how are we supposed to get enough people doing those things if everyone say "There isn't enough data right now for it to be useful"? We have to start somewhere
total duration: 24.919887458s
load duration: 39.315083ms
prompt eval count: 37 token(s)
prompt eval duration: 963.071ms
prompt eval rate: 38.42 tokens/s
eval count: 441 token(s)
eval duration: 23.916616s
eval rate: 18.44 tokens/s
I have a gaming PC with a 4090 I could try, but I don't think this model would fitWhat quantization are you using? What's the runtime+version you run this with? And the rest of the settings?
Edit: Turns out parent is using Q4 for their test. Doing the same test with LM Studio and a 3090ti + Ryzen 5950X (with 44 layers on GPU, 2 on CPU) I get ~15 tokens/second.
Only settings I did were the ones shown in the blog post
OLLAMA_FLASH_ATTENTION=1
OLLAMA_KV_CACHE_TYPE=q8_0
Ran the model like ollama run gemma2:27b --verbose
With the same prompt, "Can you write me a story about a tortoise and a hare, but one that involves a race to get the most tokens per second?"Example: the default model weights for Llama 3.3 70b, after hitting the “view all” have this hash and size listed next to it - a6eb4748fd29 • 43GB
Now scroll down through the list and you will find the one that matches that hash and size is “70b-instruct-q4_K_M”. That tells you that the default weights for Llama 3.3 70B from Ollama are 4-bit quantized (q4) while the “K_M” tells you a bit about what techniques were used during quantization to balance size and performance.
total_duration: 10530451000
load_duration: 54350253
prompt_eval_count: 36
prompt_eval_duration: 29000000
prompt_token/s: 1241.38
eval_count: 460
eval_duration: 10445000000
response_token/s: 44.04
Fast prompt eval is important when feeding larger contexts into these models, which is required for almost anything useful. GPUs have other advantages for traditional ML, whisper models, vision, and image generation. There's a lot of flexibility that doesn't really get discussed when folks trot out the 'just buy a mac' line.Anecdotally I can share my revealed preference. I have both an M3 (36gb) as well as a GPU machine, and I went through the trouble of putting my GPU box online because it was so much faster than the mac. And doubling up the GPUs allows me to run models like the deepseek-tuned llama 3.3, with which I have completely replaced my use of chatgpt 4o.
total duration: 10.5922028s
load duration: 21.1739ms
prompt eval count: 36 token(s)
prompt eval duration: 546ms
prompt eval rate: 65.93 tokens/s
eval count: 467 token(s)
eval duration: 10.023s
eval rate: 46.59 tokens/sThe same on Nvidia (various models) https://github.com/ggerganov/llama.cpp/issues/11474
[1] this is a the model: https://huggingface.co/unsloth/DeepSeek-R1-GGUF/tree/main/De...
I'm not sure I'm reading the results wrong or missing some vital context, but that sounds unlikely to me.
Where the a100 and other similar chips dominate is in training &c, which is mostly a question of flops.
I don't think they do.
From Wikipedia:
> the M2 Pro, M2 Max, and M2 Ultra have approximately 200 GB/s, 400 GB/s, and 800 GB/s respectively
From techpowerup:
> NVIDIA A100 SXM4 80 GB - Memory bandwidth - 2.04 TB/s
Seems to be a magnitude of difference, and that's just the bandwidth.
I've heard that Macs are pretty slow with XL and borderline unusable for flux requiring minutes at a time to generate a single image - whereas an RTX4090 can generate a 1024x1024 image with the higher quality Flux Dev model (not schnell) in 14 seconds.
OP is probably correct that if you want to branch out of just strictly LLM's, cuda is the way to go. I've never heard of anyone getting LTX or hunyuan running on a Mac for example.
Around half that price tag was attributed to the blogger reusing an old workstation he had lying around. Beyond this point, OP slapped two graphics cards into an old rig. A better description would be something like "what buying two graphics cards gets you in terms of AI".
Meaning what? This is largely what you do on a budget since RAM is such a difference maker in token generation. This is what's recommended. OP could buy an a100, but that wouldn't be a budget build.
Why do you say this? I thought the p40 only had a memory bandwidth of 346 Gbytes/sec. The m4 is 546 GB/s. So the macbook should kick the crap out of the p40.
I have OpenWebUI and LibreChat running on my local “app server” and I’m quite enjoying that but every time I price out a beefier box I feel like the ROI just isn’t there, especially for an industry that is moving so fast.
Privacy is not something to ignore at all but the cost of inference online is very hard to beat, especially when I’m still learning how best to use LLMs.
But to get commercially competitive models you need 5 figures of hardware, and then need to actually run it securely and reliably. Pay as you go with multiple vendors as fallback is a better option right now if you don't need harder privacy.
I think sooner or later I'll break down and buy a server for local inference even if the ROI is upside down because it would be a fun project. I also find that these thing fall in the "You don't know what you will do with it until you have it and it starts unlocking things in your mind"-category. I'm sure there are things I would have it grind on overnight just to test/play with an idea which is something I'd be less likely to do on a paid API.
Exactly. Once the price and performance get to the level where buying stuff for local training and inferencing… that is when we will start to see the LLM break out of its current “corporate lawyer safe” stage and really begin to shake things up.
Same, esp. if you factor in the cost of renting. Even if you run 24/7 it's hard to see it paying off in half the time it will take to be obsolete
Core Ultra Arc iGPU boxes are pretty neat too for being standalone and can be loaded up with DDR5 shared memory, efficient and usable in terms of speed, though that's definitely low end performance, plus SYCL and IPEX are a bit eh.
But you should still play with proxmox, just not for this purpose. My recommendation would be to get an i7 HP Elitedesk. I have multiple racks in my basement, hundreds of gigs of ram, multiple 2U 2x processor enterprise servers etc.... but at this point all of it is turned off and a single HP Elitedesk with a 2nd NIC added and 64GB of ram is doing everything I ever needed and more.
> This runs the 671B model in Q4 quantization at 3.5-4.25 TPS for $2K on a single socket Epyc server motherboard using 512GB of RAM.
If there are things you cannot send to a random party, you might want to look at hosted versions with agreements (if it's a code issue, if you're fine with github then azure is probably fine too).
Outside of that, if you really need to then sure, but these are the kinds of things that really benefit from being able to get high usage on GPUs for short periods of time.
The thing is, however, that at 2k one is not paying good money, one is paying near the least amount possible. TFA specifically is about building a machine on a budget, and as such cuts corners to save costs, e.g. by buying older cards.
Just because 2k is not a negligible amount in itself, that doesn't also automatically make it adequate for the purpose. Look for example at the 15k, 25k, and 40k price range tinyboxes:
It's like buying a 2k-worth used car, and expecting it to perform as well as a 40k one.
This is the problem.
If your use case is getting a small handful of non-urgent responses per day then it's not a problem. That's not how most people use LLMs, though.
Let me know if you find some config that really leverages more cores!
In the future, I expect this to not be the case, because models will be far more efficient. At this pace, maybe even 6 months can make a difference.
1 https://www.techpolicy.press/shining-a-light-on-shadow-promp...
What makes Apple attractive is (as the author mentions) that RAM is shared between main and video RAM whereas NVidia is quite intentionally segmenting the market and charging huge premiums for high VRAM cards. Here are some options:
1. Base $599 Mac Mini: 16GB of RAM. Stocked in store.
2. $999 Mac Mini: 24GB of RAM. Stocked in store.
3. Add RAM to either of the above up to 32GB. It's not cheap at $200/8GB but you can buy a Mac Mini with 32GB of shared RAM for $999, substantially cheaper than the author's PC build but less storage (although you can upgrade that too).
4. M4 Pro: $1399 w/ 24GB of RAM. Stocked in store. You can customize this all the way to 64GB of RAM for +$600 so $1999 in total. That is amazing value for this kind of workload.
5. The Mac Studio is really the ultimate option. Way more cores and you can go all the way to 192GB of unified memory (for a $6000 machine). The problem here is that the Mac Studio is old, still on the M2 architecture. An M4 Ultra update is expected sometime this year, possibly late this year.
6. You can get into clustering these (eg [1]).
7. There are various Macbook Pro options, the highest of which is a 16" Mackbook Pro with 128GB of unified memory for $4999.
But the main takeaway is the M4 Mac Mini is fantastic value.
Some more random thoughts:
- Some Mac Minis have Thunderbolt 5 ("TB5"), which is up to either 80Gbps or 120Gbps bidirectional (I've seen it quoted as both);
- Mac Minis have the option of 10GbE (+$200);
- The Mac Mini has 2 USB3 ports and either 3 TB4 or 3 TB5 ports.
An M4 Pro still has only 273GB/s, while even the 2 generations old RTX 3090 has 935GB/s.
Oh and the top end Macbook Pro 16 (the only current Mac with an M4 Max) has 410GB/s memory bandwidth.
Obviously the Mac Studio is at a much higher price point.
Still, you need to spend $1500+ to get an NVidia GPU with >12GB of RAM. Multiple of those starts adding up quick. Put multiple in the same box and you're talking more expensive case, PSU, mainboard, etc and cooling too.
Apple has a really interesting opportunity here with their unified memory architecture and power efficiency.
So lets say we'd run a model on a Mac Mini M4 with 24GB RAM, how many tokens/s are you getting? Then if we run the exact same model but with a RTX 3090ti for example, how many tokens/s are you getting?
Do these comparisons exist somewhere online already? I understand it's possible to run the model on Apple hardware today, with the unified memory, but how fast is that really?
Needless to say, ram isn't everything.
That said, the base storage option is only 512GB, and if this machine is also a daily driver, you’re going to want to bump that up a bit. Still, it’s an amazing machine for under $3K.
Samsung 990PRO 2TB is $170 and Acasis T5 80Gbps is €300. So it makes sense to buy external for ≥ 2TB, more flexible as well :-)
For 1TB it makes more sense to buy built-in as you note above.
If that is enough for your use case, it may make sense to wait 2 months and get a Ryzen AI Max+ 395 APU, which will have the same memory bandwith, but allows for up to 128GB RAM. For probably ~half the Mac's price.
Usual AMD driver disclaimer applies, but then again inference is most often way easier to get running than training.
On the other hand I am wondering about what is the state of the art in CPU + GPU inference. Prompt processing is both compute and memory constrained, but I think token generation afterwards is mostly memory bound. Are there any tools that support loading a few layers at a time into a GPU for initial prompt processing and then switches to CPU inference for token generation? Last time I experimented it was possible to run some layers on the GPU and some on the CPU, but to me it seems more efficient to run everything on the GPU initially (but a few layers at a time so they fit in VRAM) and then switch to the CPU when doing the memory bound token generation.
Look into RPC. Llama.cpp supports it.
* https://www.reddit.com/r/LocalLLaMA/comments/1cyzi9e/llamacp...
> Last time I experimented it was possible to run some layers on the GPU and some on the CPU, but to me it seems more efficient to run everything on the GPU initially (but a few layers at a time so they fit in VRAM) and then switch to the CPU when doing the memory bound token generation.
Moving layers over the PCIe bus to do this is going to be slow, which seems to be the issue with that strategy. I think it the key is to use MoE and be smart about which layers go where. This project seems to be doing that with great results:
* https://github.com/kvcache-ai/ktransformers/blob/main/doc/en...
If the goal however is not to tinker but to really build and learn AI, it is going to be financially better to rent those GPUs/TPUs as needs arise.
Yes, it will not rival OpenAI, but it's 100% local with no monthly fees and depending on the model no censoring or limits on what you can do with it.
Not necessarily. For non-professional purposes, I've spent zero dollars (no additional memory or GPU) and I'm running a local language model that's good enough to help with many kinds of tasks including writing, coding, and translation.
It's a personal, private, budget AI that requires no network connection or third-party servers.
People can play with "small" or "medium" models less powerfull and cheaper cards. A Nvidia Geforce RTX 3060 card with "only" 12Gb VRAM can be found around €200-250 on second hand market (and they are around 300~350 new).
In my opinion, 48Gb of VRAM is overkill to call it "on a budget", for me this setup is nice but it's for semi-professional or professional usage.
There is of course a trade off to use medium or small models, but being "on a budget" is also to do trade off.
1080Ti might even be a better option, it also has a 12gb model and some reports say it even outperforms the 3060, in non-rtx I presume.
Note that CUDA version numbers are confusing, the compute number is a different thing than the runtime/driver version.
less than $500 total feels more fitting as a ‘budget’ build - €1700 is more along the lines of ‘enthusiast’ or less charitably “I am rich enough to afford expensive hobbies”
If it’s your business and you expect to recoup the cost and write off the cost on your taxes, that’s one thing - but if you’re just looking to run a personal local LLM for funnies, that’s not an accessible price tag.
I suppose “or you could just buy a Mac” should have tipped me off though.
A beefed up home pod with a local LLM-based assistant would be a more typical Apple product. But they'd probably need LLMs to become much, much more reliable to not ruin their reputation over this.
With a big glaring exception: developer laptops are overwhelmingly Apple's game right now. It seems like they should be able to piggyback off of that, given that the decision makers are going to be in the same branch of the customer company.
Using cloud infrastructure should help with this issue. It may cost much more per run but money can be saved if usage is intermittent.
How are HN users handling this?
Combine the best of both worlds. I have a local assistant (communicate via Telegram) that handles tool-calling and basic calendar/todo management (running on a RTX 3090ti), but for more complicated stuff, it can call out to more advanced models (currently using OpenAI APIs for this) granted the request itself doesn't involve personal data, then it flat out refuses, for better or worse.
I’m not saying that dumping $10k into rapidly depreciating local hardware is the more economical choice, just that people often discount the likelihood and cost of making mistakes in the cloud during their evaluations and the time investment required to ensure you have the correct safeguards in-place.
There's a healthy secondary market for GPUs.
I'm seeing 24GB M40 cards for $200, 24GB K80 cards for $40 on eBay.
* Cheap
* Fast
* Decent amount of RAM
Pick two.
These old GPUs are as cheap as they are because they don’t perform well.
So Cheap and Decent amount of RAM work for me.
I'm about to plunge in as others have to get my own homelab running the current crop of models. I think there's no time like the present.
> How are HN users handling this? I’m working on a startup for end-to-end confidential AI using secure enclaves in the cloud (think of it like extending a local+private setup to the cloud with verifiable security guarantees). Live demo with DeepSeek 70B: chat.tinfoil.sh
To your point about cloud models, these are really quite cheap these days, especially for inference. If you're just doing conversation or tool use, you're unlikely to spend more than the cost of a local server, and the price per token is a race to the bottom.
If you're doing training or processing a ton of documents for RAG setups, you can run these in batches locally overnight and let them take as long as they need, only paying for power. Then you can use cloud services on the resulting model or RAG for quick and cheap inference.
Plenty of people in tech earn enough to support a family and drive a fancy car, but choose not to. A used RTX 3090 isn't cheap, but you can afford a lot of $1000 GPUs if you don't buy that $40k car.
Other options include only running the smaller LLMs; buying dated cards and praying you can get the drivers to work; or just using hosted LLMs like normal people.
In this setup the model is sharded between cards so data must be shuffled through a PCIe 3.0 x16 link which is limited to ~16 GB/s max. For reference that’s an order of magnitude lower than the ~350 GB/s memory bandwidth of the Tesla P40 cards being used.
Author didn’t mention NVLink so I’m presuming it wasn’t used, but I believe these cards would support it.
Building on a budget is really hard. In my experience 5-15 tok/s is a bit too slow for use cases like coding, but I admit once you’ve had a taste of 150 tok/s it’s hard to go back (I’ve been spoiled by RTX 4090 with vLLM).
How would you setup NVLink, if the cards support it?
Mode-collapse. One reason that the tuned (or tuning-contaminated models) are bad for creative writing: every protagonist and place seems to be named the same thing.
We probably have different definitions for "budget", but I just ordered a super janky eGPU setup for my very dated 8th gen Intel NUC, with a m2->pcie adapter, a PSU, and a refurb Intel A770 for about 350 all-in, not bad considering that's about the cost of a proper Thunderbolt eGPU enclosure alone.
The overall idea: A770 seems like a really good budget LLM GPU since it has more memory (16GB) and more memory bandwidth (512GB/s) than a 4070, but costs a tiny fraction. The m2-pcie adapter should give it a bit more bandwidth to the rest of the system than Thunderbolt as well, so hopefully it'll make for a decent gaming experience too.
If the eGPU part of the setup doesn't work out for some reason, I'll probably just bite the bullet and order the rest of the PC for a couple hundred more, and return the m2-pcie adapter (I got it off of Amazon instead of Aliexpress specifically so I could do this), looking to end up somewhere around 600 bux total. I think that's probably a more reasonable price of entry for something like this for most people.
Curious if anyone else has experience with the A770 for LLM? Been looking at Intel's https://github.com/intel/ipex-llm project and it looked pretty promising, that's what made me pull the trigger in the end. Am I making a huge mistake?
I'm seeing A770s for about $500 - $550. Where did you find a refurb one for $350 (or less since you're also including other parts of the system)
It's out of stock now unfortunately, but it does seem to pop up again from time to time according to Slickdeals: https://slickdeals.net/newsearch.php?q=a770&pp=20&sort=newes...
I would probably just watch the listing and/or set up a deal alert on Slickdeals and wait. If you're in a hurry though, you can probably find a used one on Ebay for not too much more.
In my budget AI setup I use 7840 Ryzen based miniPC with USB4 port and connect 3090 to it via the eGPU adapter (ADT-link UT3G). It costed me about $1000 total and I can easily achieve 35 t/s with qwen2.5-coder-32b using ollama.
How unfortunate that people are discounting the likelihood that American AI agents will avoid saying things their master think should not be said. Anyone want to take bets on when the big 3 (Open AI, Meta, and Google) will quietly remove anything to do with DEI, trans people, or global warming? They'll start out changing all mentions of "Gulf of Mexico" to "Gulf of America", but then what?
1 https://www.techpolicy.press/shining-a-light-on-shadow-promp...
And q8_0 already halves the memory usage compared to fp16.
One of the ollama Devs called the quality impact negligible at q8_0: https://smcleod.net/2024/12/bringing-k/v-context-quantisatio...
But perhaps quantifying the KV cache does not scale as gracefully as the model itself?
1. https://arxiv.org/abs/2412.19437v1
2. Quibble over the exact figure. Far less than Open AI, doing more with less.
But it should be said that basing model quality on its size in GB is like qualifying a video based on its size in GB. You can have the same video be small or huge with anywhere from negligible to huge differences in quality between the two.
You will be running quantizied model weights, which can range in precision from 1 to 16 bits per parameter (the B for billion in the model name). Model weights at Q8 are generally their parameter size without the B in GB (Llama 3 8B at Q8 would be ~8GB). There are many different strategies for quantizing as well, so this is just a rough guide.
So basically if you can't fit the 48GB model into your 48GB of VRAM, just download a lower precision quant.
For example, if you want to run "CodeLlama 70B" from https://huggingface.co/TheBloke/CodeLlama-70B-Python-GGUF where's a table saying the "Q4_K_M" quantised version is a 41.42 GB download and runs in 43.92 GB of memory.
Of course I do want my own local GPU compute setup, but the juice just isn't worth the squeeze.
My personal DL machine has a 24 core CPU, 128GB RAM and 2 x 3060 GPUs and 2 x 2TB NVMe drives in a RAID 1 array. I <3 it.
https://www.hp.com/us-en/shop/mdp/business-solutions/z440-wo...
No integrated graphics.
Author's explanation of the problem:
The Teslas are intended to crunch numbers, not to play video games with. Consequently, they don't have any ports to connect a monitor to. The BIOS of the HP Z440 does not like this. It refuses to boot if there is no way to output a video signal.
Its funny to see people independently "discover" these builds that are a year plus old.
Everyone is sleeping on these guides, but I guess the stink of 4chan scares people away?
If that isn't "ancient" in terms of AI workstation build guides, then I don't know what is.
https://stratechery.com/2025/deep-research-and-knowledge-val...
> Unless, of course, the information that matters is not on the Internet. This is why I am not sharing the Deep Research report that provoked this insight: I happen to know some things about the industry in question — which is not related to tech, to be clear — because I have a friend who works in it, and it is suddenly clear to me how much future economic value is wrapped up in information not being public. In this case the entity in question is privately held, so there aren’t stock market filings, public reports, barely even a webpage! And so AI is blind.
(edited for clarity)
Why trust the good will of a company, over a box that you built yourself, and have complete control over?
https://deepgains.substack.com/p/running-deepseek-locally-fo...
Again, though, my experience is limited. I imagine others know something I do not and would absolutely love to hear more from people who are running tiny models on low-end hardware for things like code assistance, since that's where the use-case would lie for me.
At the moment, I subscribe to "cloud" models that I use for various tasks and that seems to be working well enough, but it would be nice to have a personal model that I could train on very specific data. I'm sure I am missing something, since it's also hard to keep up with all the developments in the Generative AI world.
"AI computer" sounds pretentious and misleading to outsiders.