There’s plenty of room at the Top: computer performance after Moore’s law
techrepublic.com
techrepublic.com
The paper's URL is https://science.sciencemag.org/content/368/6495/eaam9744 but the text seems to be paywalled. We've put its title above.
https://www.google.com/search?q=there's+plenty+of+room+at+th...
[1] https://thesquareplanet.com/
Talks
[2] https://youtu.be/s19G6n0UjsM
[3] https://youtu.be/DnT-LUQgc7s
Rust Details(my favorites)
There's plenty of room at the bottom. Another million times the transistors is hardly implausible; a thousand times is practically a given.
> Although other manufacturers continued to miniaturize—for example, with the Samsung Exynos 9825 (8) and the Apple A13 Bionic (9)—they also failed to meet the Moore cadence.
This just factually isn't fair; transistor density has been on an unwavering straight line since 1970.
https://docs.google.com/spreadsheets/d/1NNOqbJfcISFyMd0EsSrh...
Power may soon become the bottleneck as it's not clear that energy per op is going down as fast as transistors are going up
The only way we can really use all that transistor density today is to artificially limit clock speeds to conserve power.
Productive uses for a billion transistors that can be mostly turned off most of the time are hard to find.
All of it is parallel and benefits from a sheer increase in transistors and density. Maybe in a few generations we'll see chips that are 99% ML compute and 1% traditional CPU.
However, these lower power chips are still designed like classical computers and are deterministic. What if you had something like a few thousand analog multipliers in parallel computing matrix products? An analog multiply can use far less power than a digital one at the expense of accuracy and determinism, which might not matter for some ML tasks.
TSMC 5nm, 3nm, 2nm, none of them will hit 2x density improvement compared to its previous node within 24 / 30 months period.
TSMC 3nm does sound less ambitious, but it's also supposedly coming early, and sounds too way early to call for sure, at least for anyone not under NDA.
I started learning to code as a kid in the 80s on 8-bit machines. Seemed like there was rapid progress in the availability of CPU power, RAM, etc., up until the early 2000s and then things seemed to... slow down.
Which is about the time everybody started to focus hard on horizontal scalability.
Google Docs isn't the product. It's free. You're the product. And it only needs to be good enough to keep competitors from creating something so great that it removes their first-mover advantage. It's clear that Google has given up developing new features for it.
[ADDED: And as other have mentioned while the collaborative editing of Google Docs doesn't scale all that well, it's a great remote collaboration tool. Never want to go back to sending docs around.]
I actually never much cared for WordPerfect with all its formatting codes and so forth. I was actually a bigger fan of early DOS-based Microsoft Word at the time.
We never went past 64 bits because it's enough bits to address memory for the foreseeable future, and applications needing more are rare. It's mostly cryptography.
My first paid job was in the late 80's. I had to do something about a system that was basically a lot of menus written in Oracle Forms running in terminals on SCO Unix or maybe Xenix.
I was unbearable slow. Users had to wait 5+ seconds every time they pressed a key. And it was a nightmare to try to optimize it. But then I scrapped it and rewrote the whole thing using plain C and curses and system() calls to do the actual work. Hard-coding the menu hierarchy structure which nobody was changing anyway. That made the system very fast. The company was very surprised that navigating menus could be so fast.
Constraint based creativity needed!
Oh and for some reason their language and APIs are OO despite the fairly tight memory limitations, and despite not having any real need to do that (nor enough memory to really take advantage of any OO features even if you wanted to). So you've really got to watch (haha) yourself.
I tell you what, though: having a declarative UI where "declarative" is "now draw text at coords [some coords]. Then draw a square of dimensions [dimensions] at coords [other coords]" is a fucking refreshing break from webdev. Even more so than regular mobile dev. There's simply no room on the devices for much screwing about. You've got your lifecycle methods and some calls that draw things. You can listen for button presses. That's about it as far as the UI goes.
P.S. Please don't shoot the messenger, I'm just linking it.
Also, you still worry about small amounts at Google scale:
> “It’s pretty slow,” Jeff said. He leaned forward, still relaxed. “So that one was a hundred twenty kilobytes,” Sanjay said, “and it was, like, eight seconds.” “A hundred twenty thousand stack calls,” Jeff said, “not kilobytes.” [...] In a sense, they had been occupied by minutiae. Their code, however, is executed at Google’s scale. The kilobits and microseconds they worry over are multiplied as much as a billionfold in data centers around the world. – https://www.newyorker.com/magazine/2018/12/10/the-friendship...
If you are in a startup where you rather worry about product-market fit, of course such minutiae is not relevant (yet).
It won't happen with this generation of the industry.
That being said, there are areas where users care about speed. For example, FPS is very important to gamers, so games are often highly optimized.
So that was a factor, but so was the fact that the entire subsystem was pretty much written in COBOL, though I would occasionally write small C programs to work around limitations. Even newer versions of the ERP had a lot of the old COBOL still in it, just with a shinier new interface, but they were gradually replacing that code with (I think) Java. Probably because of the ability to keep finding & retaining COBOL talent, along with the need to make the product more readily interfaced via the web rather than a custom presentation layer from the vendor. (which was a green-screen terminal based interface, and significantly faster than any interface you'd see today. And with modern terminal emulators you weren't limited to green, you could choose any color! )
but slow internal tools are a direct tax on developer productivity
this is an area where faster / smaller software does make a difference
We need to start thinking about performance as a necessary and basic feature, not some nice to have that can be worked on later. A new program needs to designed from the ground up to be fast, from the language choice to the data structures.
And I think this shift is being made, except in web development circles. Most new and upcoming languages list performance as a basic feature: Rust, Julia, Nim, Zig, etc. Also, native compilation is often listed as a feature, as there's been a huge backlash against the vm cargo cult on the late 90s and 2000s.
As for browsers shrug, turn off JS and add a blocklist and it'll run an order of magnitude faster.
Mostly, company-mode lags when given a large autocomplete list. Also, large CSV files take a while to load.
Large CSVs, Perhaps use LFV (large file viewer) if you want a read-only look.
Otherwise I just loaded a 26MB binary file in less than a second. I created a 2.46 Gigabyte (not MB) CSV, opened it in a new emacs, took less that 15 seconds (fundamental mode). I think your problem lies elsewhere - it may be an unnecessary mode being invoked.
How did the browser run when you turned off JS, a whole lot faster I imagine?
Lloyd, S. Ultimate physical limits to computation. Nature volume 406, 1047–1054 (2000)
Bremermann, H. J. Minimum energy requirements of information transfer and computing. Int. J. Theor. Phys. 21, 203–217 (1982)
Bekenstein, J. D. Energy cost of information transfer. Phys. Rev. Lett. 46, 623–626 (1981)
"As miniaturization wanes, the silicon-fabrication improvements at the Bottom will no longer provide the predictable, broad-based gains in computer performance that society has enjoyed for more than 50 years. Performance-engineering of software, development of algorithms, and hardware streamlining at the Top can continue to make computer applications faster in the post-Moore era, rivaling the gains accrued over many years by Moore’s law. Unlike the historical gains at the Bottom, however, the gains at the Top will be opportunistic, uneven, sporadic, and subject to diminishing returns as problems become better explored. But even where opportunities exist, it may be hard to exploit them if the necessary modifications to a component require compatibility with other components. Big components can allow their owners to capture the economic advantages from performance gains at the Top while minimizing external disruptions."
(My emphasis on "economic advantages from performance gains.")
Databases are still getting faster. You don't get a faster DB by buying a new server and then running a 10 year old version of SQL Server or PostgreSQL on it.
Language runtimes are still getting faster. You don't get better Java or JavaScript performance by running a 10 year old release of the HotSpot JVM or V8.
3D renderers, video encoders, maximum flow solvers (one example examined in depth in the article): they're all getting faster over time at producing the same outputs from the same inputs.
The key is incentives. There are probably people in your organization who care a lot if database operations slow down 10x. But practically nobody in your organization cares if Slack responds to key presses 10x slower or uses 100x as much memory as your favorite lightweight IRC client. (I'm trying to make a neutral observation here. Looking at it from 10,000 feet, I get annoyed when I see how much memory Slack takes on my own machine, but that annoyance wouldn't crack the top 20 priorities for things that would improve the productivity of the business I'm in.)
The incentives problem is also why the Web seems slow. Browsers too are still getting faster. I used to run multiple browsers for testing and old Firefox and IE releases were actually much slower at rendering identical pages than current stable releases. But pages are getting heavier over time. Mostly it's not even a problem of people trying to make "too fancy" sites that are applications-in-a-browser. It's mostly analytics and advertising that makes everything painfully slow and battery-draining. I run a web site that has had the same ad-free, analytics-free, JS-light design since the early 2000s. It renders faster than ever on modern browsers. It's not modern browsers that are the problem -- it's the economic incentive to stuff scripts in a page until it is just this side of unbearable.
For certain kinds of human-computer interaction, people will pay a lot to reduce latency. Competitive gamers will pay, for example. Sometimes people will invest a lot of effort to reduce memory footprint -- either because they're shaving a penny off of a million embedded devices or because they're bumping up against the memory you can fit in a 10 million dollar cluster. But the annoyances that dominate Wirth's Law discussions on HN -- why do we have to use Electron apps?! -- are unlikely to get fixed because few people are willing to pay for better.
Now we have apps that are like 40+ MB, eat RAM and CPU cycles to do fairly simple things.
So yeah it's disheartening to see programs with the same functionality of an old 16 bit 0x86 program using hundreds of MB's of RAM.
Games like Destiny weighs in at >100GB with it's latest updates due to map and image needs, and the slack client runs smoothly across 5 Operating systems.
Turns out, you can! There's a big domain that needs your help, today: deep learning.
A lot of the software is currently crude. For example, to train a StyleGAN model, the official way to do it is to encode each photo as uncompressed(!) RGB, resulting in a 10-20x size explosion.
There's plenty of room at the top, and never moreso in AI software. Consider it! Every one of you can pivot to deep learning, if you want to. There's really nothing special or magical in terms of knowledge that you need to study. A lot of it is just "Get these bits from point A to point B efficiently."
There's also room for beautiful tools. It reminds me a lot of Javascript in 2008. I'm sure that will sound repugnant for a majority of devs. But for a certain type of dev, you'll hear that ringing noise of opportunity knocking.
Enterprise software!
But we’re more likely to see true General Artificial Intelligence before we see a 2x increased in enterprise software performance.
https://www.uipath.com/blog/whats-the-difference-between-rob...
Keep in mind, though, that RPAs often link together different pieces of software, and if the GUI changes for one of them you'd only have to retrain on that piece. I wouldn't be surprised if enterprise software vendors start optimizing their programs for RPAs so that they don't have to rely on hacky GUI monitoring as much.
The end game is that large, frequently run RPAs will serve as a flag for the organization to develop, find, or outsource an end-to-end program that executes the same process without an RPA. RPAs, then, will always be the scout at the frontier of automatable office work, finding who can be freed from rote drudgery next and serving as a bridge to the best programmatic solution.
What you need is determination. There's no substitute for this.
If you have that, there are all kinds of resources. Here are a few...
Resource 1: a community. We've set up a discord server for AI dev. It has 360 users. At any given time, around ~50 people are online, of which ~10 are skilled devs. Come join! https://discordapp.com/invite/x52Xz3y
Resource 2: Find something fun for you, and pursue that. I like generative AI, so for me that's been GPT-2 and StyleGAN. Gwern has some lovely tutorial-type articles on both.
GPT-2: https://www.gwern.net/GPT-2
StyleGAN: https://www.gwern.net/Faces
Peter Baylies' StyleGAN tutorial notebook is a hands-on resource. This was actually how I started, nearly a year ago.
github: https://github.com/pbaylies/stylegan-encoder
notebook: https://colab.research.google.com/drive/179SPYbBC8pKDxVRjZep...
(The original notebook was broken; this is a copy I've updated with some fixes.)
Lastly, follow a bunch of AI people on twitter. Here are a few to get you started: https://twitter.com/i/lists/1160386581850730496/members
The reason to follow them is, whenever you see something that seems interesting or fun, tweet at them and say so! Ask questions. Ask how to get started. Everyone is shockingly nice and helpful. My theory is, the software is so crude and often hard to use, that we all like to celebrate together whenever one of us gets it working, and we're happy to share that knowledge however we can. (Twitter is a bit chaotic right now due to world events, but I imagine it might return to normal within a couple weeks.)
And yes, you're right about fast.ai and other courses. You can go that route if you like it. I found it more exciting to dive into the deep end, though, and try to tinker with stuff.
But, have you achieved something singificant in the context of deep learning, as per your earlier comment? There's no information about that in your profile and a cursory glance at ddg and google results for "shawn presser" doesn't turn up anything very relevant.
So, I have to ask: without having studied CS, what contributions have you made in deep learning that are widely recognised?
I hope you agree this is a reasonable question to ask, and that you are not offended by it. Otherwise, I apologise because it's not my intention to offend you.
- A Newsweek article https://www.newsweek.com/openai-text-generator-gpt-2-video-g...
- GPT-2 chess https://www.theregister.com/2020/01/10/gpt2_chess/
- ... which DeepMind referenced: https://twitter.com/theshawwn/status/1226916484938530819
- GPT-2 music https://soundcloud.com/theshawwn/sets/ai-generated-videogame...
- Swarm training (WIP) https://www.docdroid.net/faDq8Bu/swarm-training-v01a.pdf
To clarify, what I was hoping to see is, at best, an article published at a reputable venue for AI research, a conference or a journal, or at a minimum an arxiv article that at least looks like it was meant to be submitted to a conference or journal. And at worst, a software tool that can be used in deep learning research. But it seems to me that your achievements are mainly having fun with and in one case finding an interesting use for tools that are already available.
Again, I'm not trying to be harsh, neither do I want to say that all this is not worth the trouble. But it should not be held up as an example of what people can achieve without studying CS. Because, I think you'll agree, they are kind of underwhelming when compared to what people routinely achieve who have studied CS.
You really need to get out more!
With other ML resources I find they get bogged down trying to explain the concepts/details, and I would usually lose interest before getting to the implementation of things.
For instance, I'm sure that self-studying some fundamental subjects in CS, like formal languages and automata and complexity theory, will give you important tools to tackle any CS-related problem, be it programming a more efficient deep learning implementation or really, anything else you like.
Spending the same effort to get into deep learning instead, will most likely leave you with only superficial knowledge that will not transfer to other subjects.
I have no idea if I would or would not be better served by a traditional CS background in this case, so part of this is to find out if that is or is not the case.
In that case it's very difficult to go wrong with acquiring a traditional CS background. There is really nothing you can do with computers that will not benefit from a CS background, including anything that may have to do with deep learning etc. On the other hand, learing about deep learning will only help you when it comes to working with deep learning.
So, regardless of whether getting into deep learning may help "put bread on the table", getting a well-rounded CS education, certainly will.
And really this is only for software that runs inside their servers. The costs for software running on client machines is just distributed to the user.
If google makes a change that makes Chrome run less efficiently that means I'm paying for it with my time and electricity. If I want it to run better I'm shelling out money from my own pocket for new hardware.
This is why all classes for scientists I've seen teach numpy.
(Incidentally: total time to "optimise" and TEST is less than 5 minutes.
Total time to rewrite in Rust: ongoing after 2 hours, with issues filed on the best available crate for the task, whose documentation examples fail to compile. Not even going to try in bare Rust due to the issues with > 32 length arrays.)