10GHz at under 1V by 2005 - The future of Intel’s manufacturing processes [2000]
anandtech.com
anandtech.com
Intel introduced a 65 nm (0.065 micron) process in 2006. The "Cedar Mill" Pentium 4 processor ran at 3.6 GHz at a whopping 1.3V although a small double-pumped part of the processor ran at 7.2 GHz. It could be overclocked to 4.5/9.0 GHz at 1.4V.
The discrepancy between 0.85V and 1.3V was caused by the end of Dennard scaling. Basically transistors require much more voltage than predicted and thus consume far more power than predicted. Although the transistors can technically run at 9 GHz, the resulting power density is very difficult to cool.
But nowadays we have processors with multiple cores, where sometimes you need only 1 core (and it needs to be fast). So would it be an idea to increase clock frequency for those cores, but multiplex them quickly to allow them to cool?
Some BIOSes have settings to completely disable a bunch of cores to enable more turbo boosting.
This isn't as sophisticated in thermal as what you outlined, but it saves on cache coherence.
https://en.wikipedia.org/wiki/Northstar_engine_series
When the engine did NOT have coolant, it would run only one bank of cylinders at a time and alternate between the two to let them cool.
The system was at least somewhat effective. There's a story from the time about a journalist that was testing the feature in the desert. He stopped at a truck stop after conducting the test and amazed the folks at the stop by opening the hood, filling the engine with coolant, and driving off like nothing was amiss.
introduced in ~2007 in Core 2 mobile CPUs (Penryns I think). You can load Middleton modded bios into your oldshool Thinkpad R61/T61 and it will enable dual IDA[1] in top CPUs.
http://forum.notebookreview.com/threads/t61-x61-sata-ii-1-5-...
You can even load modded(by chinese) modded middleton bios, solder one wire and plop $8 T9550 CPU for 2.66GHz boosting to 2.8GHz. You could even use quads, but its beyond uneconomical with cpus being more than whole used thinkpad T420.
http://forum.thinkpads.com/viewtopic.php?f=29&t=110620
This was Intels first steps and not a lot of OEMs enabled IDA. Turbo Boost was next, introduced in Nahalems.
https://en.wikipedia.org/wiki/Intel_Turbo_Boost
[1] dual IDA is a trick forcing IDA on both cores. Throttlestop utility is also able to turn it on.
Obviously there's some cost there in sharing registers, cache, etc, but it's an interesting notion.
If performed at the hardware level, you may also get some improvement by having the core push its register and low level cache state directly to the next processor, rather than have them pull the data through shared cache or RAM. Unlike a typical context switch, the process is not being resumed from idle, but has an active cache state.
Its not the rapid growth of the 90's and early 2000's, but it is still growth.
That 10GHz talk was a lie on the part of Intel to intimidate people away from AMD, not only would a 10GHZ P4 melt down, but it would be stalled all the time from memory latency. So many things did not work that it was not an honest mistake.
Today there is talk of a big clock rate bump (to 200 GHZ or so) if they go to a different semiconductor, but at that point you probably need a fiber optic or terahertz wave link to memory to keep the pipeline full.
You talk as if there couldn't possibly be a benefit to an increase in speed without a corresponding increase in memory bandwidth. Whilst it wouldn't be an optimally efficient system, if we /could/ bump to 9GHz (or 200GHz), wouldn't it be worth doing so for at least some kinds of calculations, even if the memory can't keep up?
edit: Both responses were super-interesting. Don't wanna reply to both, but thanks all :)
Most of the market is for things that are generalizable; maybe you could make some kind of hyper-DSP for millimeter wave base stations or something like that, but you have to spread out the development cost across a low number of units.
Are there are apps that have high computational intensity? Sure, matrix multiply is one of them. That's one of the reasons why dense linear algebra serves as the standard benchmark to determine the top 500 supercomputers in the world.
But even in HPC (high performance computing), many if not most apps actually have relatively low computational intensity (i.e. in the range of one or so compute operations per word of memory loaded). In this regime, it really doesn't make sense to grow compute out of proportion with memory bandwidth because you'll just be idling the processors.
And while I have no proof, I'd expect HPC applications to generally be more computationally intense than general consumer computing tasks. So I'd expect that computational intensity goes mostly down from here.
[1] https://www.amazon.com/Intel-Xeon-2620-Processor-BX80621E526...
It reminds me of Malthus thinking the economy/world was doomed because we were about to run out of trees, and then we started making significant use of coal (coal had been used for a long time, but only on a small scale).
--
[1] Imagine being able to speak normally with your computer as you would a secretary sitting next to you and have your computer accurately and quickly take notes from your speech.
[2] Imagine logging onto your computer not via a user name and a password but by sitting in front of your display and having it scan your face to figure out if you are allowed access to the computer.
Do you remember any specifics of where & how exactly these features were available?
For voice there were things like dragon that could do reasonably accurate full voice dictation. IME it was more accurate the google voice is today because it was trained.
Moving these to the cloud is more about controlling data than it is about lack of local computing power.
https://en.wikipedia.org/wiki/Dragon_NaturallySpeaking
Windows had an API for basic command recognition, in late 90's giving the ability to navigate through telephony menus, for example:
https://en.wikipedia.org/wiki/Microsoft_Speech_API#SAPI_1-4_...
Example 2 requires the use of depth-sensing and heat-sensitive cameras to avoid trivial "show a photo of an authorized user" attacks - that's not really CPU-dependent either.
The impressive amount of processing power available on many smartphones today certainly contributes to this being a practical unlock method.
You can build basic functionality into a speech bot in Powershell ("PowerShiri, what time is it?" "PowerShiri, what is the weather?") as a weekend project: https://news.ycombinator.com/item?id=11663029
https://myactivity.google.com/item?product=29
(If the Product 29 parameter doesn't work, click on "Item View" in the left and filter by "Voice & Audio".)
Modern speech models are quite big, not so big that you couldn't load it on your desktop, but big enough that you would notice. Couple that wih the fact that search/or a service is going to happen on a server anyways, client side processing doesn't make sense beyond a few functions.
I don't know of any smartphone or similar device that would interpret voice commands locally.
I assume Google encodes the data in something like iSAC [1], requiring 32 kbit/s for good quality speech, so an hour of dictating is 3600 x 32/(8 x 1024) = 14 MB.
[1] https://en.m.wikipedia.org/wiki/Internet_Speech_Audio_Codec
As an example, if I ask Alexa to play a certain kind of music and she plays the wrong one, I'll have to specify it further to get the correct music. Next time I'd expect her to get the clue, but she'll make the same mistake again.
Even though the most annoying thing is that chain commands don't really work with any assistant. I'd expect to give 4-5 commands at once and get them executed. Activating it for each command is very annoying.
Between Apple, Google, and Microsoft. To me only Cortana is worth using. For example. If I say "text prentice" neither Siri or Google Now can figure it out. But 100% of the time Cortana picks it up.
I do miss my windows phone :(
A coworker once told his smartphone, "Tell me a dirty joke". It kind of worked, it opened a web search for "dirty joke". Cortana really impressed me by actually telling a joke; it was about dirty laundry (but I suspect that was intentional to avoid offending people). ;-)
It works wonderfully well. I use it obsessively to the point that people now joke that if you ask me a question I'll simply ask Google about it.
E.g. I randomly tried: What do I do when I burn my finger? and it answered. Same for: How do I get ink stains out of my shirt?
Also, there might be something wrong with the mic if you're using an external bluetooth mic?
This is just my experience however on my Nexus 5 with very good wifi. Hopefully others are having a better experience, or there is a fix for the lag, because despite these problems, I use OK Google (Google Now?) quite often.
Well, about those accurately and _your_ computer parts… :)
It's refreshing that -- despite not having 10GHz CPU's available -- my 6yo CPU/mobo still does not feel in any way slow or in need of upgrade.
Yes, other components have offset the relatively slow pace of CPU improvements (GPU's for gamers, SSD's for everyone, etc) but it feels like we're enjoying a lengthy era of 'good enough grunt' on the CPU front (for most of us).
The limiting factors have changed- or rather, the limiting computers have changed.
Throwing more client horsepower or internet speed at the problem can't solve the fact that the server (and to a limited extent, an internet connection) has to deliver both a huge amount of JavaScript and the (typically image/video/audio-heavy) content itself.
The client is always underutilized- the fact that smartphone apps (or rather, cached local copies of what would normally be a website) have similar performance to a modern desktop PC should speak volumes about that. It's probably why Microsoft smartphone-ized Windows; though their execution of that was awful and it didn't help that WinRT wasn't mature on release.
And raw server performance (like desktop performance) has been at a plateau for a while now as well. The lack of competition for Intel doesn't help that (maybe AMD's new processor line will start driving improvements again, but there are no hard numbers on performance; and higher-TDP ARM designs aren't currently competitive with Intel's in this space either).
Until this changes, and there's nothing to indicate it will on Intel's roadmaps, clients will continue to be good enough- it might be the first time in history where computers (since 2008 or so) are replaced because the hardware failed and not because they were insufficiently fast.
Aside -- I grew up in the days before network reliance (or even network presence) consequently I don't think of my 32GB / 8 core / 8TB / 32" monitored computer as an enhanced VT52. :) Though I certainly respect the fact that for many people, once you exceed the grunt required to render HTTP(S) at an acceptable speed, there's no great interest in a faster computer.
Your observations on Intel -- is it possible they were a bit more prescient than we typically give them credit for, insofar as not pursuing ever faster CPU's (which would now have been considered relatively unnecessary for the majority of installations)?
You're forgetting two things:
1 - Games
2 - VR
Games especially have a very long history of having a bottomless appetite for more and more powerful hardware every year as they push the limits of what's possible.
VR is dramatically upping hardware requirements, and those requirements are going to exponentially increase as consumers start to demand and expect 8k per eye, full motion, high framerate, 360 degree, 3D, wide field of view, interactive VR -- all on smaller and lighter headsets, ideally wirelessly.
Sure, though the CPU has less to do with that. For reference, most modern games still perform just fine on first-gen i7s and second-gen i5s; the most recent of those was released 6 years ago. GPUs are a replaceable part where CPUs are not.
At least GPU technology is still advancing in big ways, though I think that's more a property of how they're built, what their functions are, and (especially for higher-end cards) what kind of power budget users are willing to accept. It's unusual for the newest mid-range card not to match the previous high-end card.
Intel, on the other hand, has never released a CPU with TDP over 150W (AMD had a couple at 220W), even though most in the overclocking community know that 5GHz is regularly attainable on modern CPUs. But that that's been mostly true since 2013.
It'd be interesting to know the sort of IPC delta between a Netburst architecture and Kaby Lake. Also, it'd be interesting to know how theoretically fast you'd have to push Netburst to see the same performance as Kaby Lake.
tl;dr: Core i7-2600k @ 3GHz, is about 2,67x faster than a Pentium 4 HT 660 (Prescott ua) @ 3Ghz. In the benchmark they tested single core, and equal clock, to keep everything as level as posible.
As for single core performance, Kaby Lake would be between 30 to 40% better than Sandy Bridge (which could be called the peak improvement architecture in the Core i brand), so it'd be about ~3,5x faster than the Prescott based chips.
Just how close are you sitting to your monitor?
http://www.tomshardware.com/news/amd-fx-8150-overclock-9ghz-...
At 10GHz, light travels approximately 3cm between each period of the clock.
Note that transistors which operate above 10GHz are not rare and used in microwave applications; as I understand it, the difficulty is in creating logic circuits with them and at a scale suitable for a CPU.
Light doesn't change speed, so this statement confuses me.
Practically speaking, it does change speed. The speed of light in a vacuum is fixed, but it varies based on what it's traveling through.
The propagation velocity of an EM signal in RG coax cable is 80% of c*, PCB traces can be as low as 50% and somewhat surprisingly both fiber optic cable and cat6 cable are about the same, ~60-70%. I don't know if there is any good public information about the velocity factor with modern CPU transmission lines.
- 10 GHz
- long pipelines, I think we're currently at around 14 stages, which is more than P6 but less than Netburst (Pentium 4)
- EUV lithography is still not a thing at 10 nm
Die I miss any important failed predictions?http://valid.canardpc.com/records.php
https://www.youtube.com/watch?v=UKN4VMOenNM (the world record at 8.429GHz)
was a statement made by anandtech. Did Intel actually strive for those numbers?
Edit: I probably exaggerate. I just mean that higher clock speed still have benefits.
I think that you would cease to see any acceleration from increased clock speed unless you also sped up other parts of the processor / computer.
I think your sentiment is correct. If we can take advantage of higher clock speeds, then why not do it?
(Caveat lector: Quoting from memory here)
But the consumer market is more than just this. Several popular "consumer" applications, such as gaming and photo/video editing, continue to see benefit from increases in CPU performance, and single-core performance (what most people actually want when they ask for higher clock speeds), continues to be very relevant to this day. Not all workloads are easily parallelizable.
Lots of (most?) software development typically charges ahead without much regard to resource usage, until performance becomes a problem; then things are optimised until performance is no longer a problem, and the charge resumes. This results in software with performance which is just about acceptable, regardless of what resources are available. It was the case in 2000, it is the case now, and it would be the case if we had 50GHz machines.
This is the case for tasks where the main bottleneck is 'has anybody bothered to implement this yet?'; I'd say your examples of gaming and video editing are tasks where performance is a major part of the bottleneck. Arguably, Trello didn't exist in the 90s because nobody had bothered to make it yet; Skyrim didn't exist in the 90s because the machines weren't up to it.
They have released a process that is worth upgrading to: https://ark.intel.com/products/97129/Intel-Core-i7-7700K-Pro...
https://ark.intel.com/products/52214/Intel-Core-i7-2600K-Pro...
There are a massive host of improvements between those two processors. Each release since Sandy Bridge has continued to increase performance between 5 to 10% in most tasks and more in some specifics. Over 4 releases that is a noticeable effect.
That processor in the US is 350 Dollars.
Although your light turns on very quickly when you flip the switch, and you find it impossible to flip off the light and get in bed before the room goes dark, the actual drift velocity of electrons through copper wires is very slow. It is the change or "signal" which propagates along wires at essentially the speed of light. refer https://en.m.wikipedia.org/wiki/Drift_velocity
Yes. Notice how after it was clear that P4 was a dead end they went and dusted off their P3 mobile line that they had been making more power efficient the whole time?
Surprisingly, power consumption also made huge impact. As tablets and laptops got more popular than desktop battery life became a major concern and thus TDP played major role in research.
Try this fun experiment: Underclock your CPU by half a GHz and see if you notice the difference in your day to day work.
Dude, Intel spends something like $80B/yr on R&D. This is closer to hitting fundamental laws of physics barriers.
They killed off their P4 line and developed their mobile line for a reason.
https://www.fool.com/investing/2017/02/05/intel-corporation-...
No amount of R&D spending can bend the laws of physics to overcome the inherent limitations of silicon. I'm sure Intel also looked into alternative semiconductors (e.g., III-V) before giving up on the 10 GHz dream.
That a secretary typing a document or someone who only spends time on facebook doesn't notice the difference is irrelevant- consider, for example, the massive capital outlay by the financial industry to have servers located as closely to the world's trading hubs as possible. If they are willing to pay whatever it takes to shave milliseconds off a round trip, faster CPUs are a part of that equation.
I think the GP did not debate that but pointed out the for CPU speed/throughput, clock speed is only part of it. Adding functional units and allowing the CPU to process more instructions in parallel can have a big impact, so can e.g. larger cache, better branch prediction and so forth.
If you give people faster CPUs, they will cheer and find something to keep them busy. ;-) And for some people, there is no such thing as "fast enough". But for a fairly large share of desktop/mobile users, the is not the limiting factor as much as memory bandwidth and I/O.
As other commenters mentioned, you can counteract this with better cooling, but eventually the marginal $/watt of cooling isn't worth the tradeoff of lower speeds & multicore processors.
So, because each switch consumes the same amount of power, the power consumed is directly proportional to the switching frequency. Double the frequency = double the power.
As other posters have alluded, Intel tried to work on the power issue by decreasing the voltage. Unfortunately, they were not able to figure out how to decrease the voltage as they had in the past. With the high power required to run high frequency circuits, cooling would have been more of a problem than could be overcome economically.
We're also at the point where transistor leakage current matters. All transistors leak a tiny amount of current even without switching. Alone this might only be a few µA, but pack enough of them together on a die and that might add up to hundreds of mA. Combine that with higher voltages so the chip can run at high clock speeds, and the fact that leakage increases with temperature (which will be driven by higher voltages and frequencies), and all of a sudden you have a feedback loop which can cause you to burn a few extra Watts before you've even done anything.
Because of the special relativity theory nothing including electrical current can transmit information faster than the speed of light (c = 3*10^10 cm/s). So if maximum size dimension of your CPU is say 2cm (I'm ignoring that they actually can have multiple smaller cores), then the maximum physically achievable frequency is upper bounded by size/c = 15 GHz.
Although it has not been reached yet, it has the same order of magnitude as the record frequencies achieved by overclockers using liquid nitrogen cooling (something a little below 9 GHz). Also it shows that reducing transistor size and as a consequence overall CPU size increases maximum physically allowed frequency.
One of my first contributions to the Linux kernel was a bugfix to the bogomips routine. It stored the result in a 32-bit variable and our 8+ GHz chilled test machines would cause that to wrap.
It was then determined that the processor would be marketed using the slow clock frequency. This was the right answer, but it didn't feel like it at the time.
That article predates the marketing change. It might have made them realized they needed that change.
For General Purpose IPC Single threaded performance, we basically had little to no improvement since Sandy Bridge apart from Clock Speed. SSD / IO Speed helped to fill the performance improvement gap in the past 4 - 5 years.
Now we are waiting for the next big improvement to come. If there are any.
how many cycles does the average method/function/procedure need anyway?
1. https://en.wikipedia.org/wiki/Parallel_computing#Amdahl.27s_...
If 90% of your runtime can be done in parallel, you still have to wait for that last 10%. You can hit this limit with 10 cores (1 processing the 10% that's serial, the other 9 processing the 90% in parallel). If you throw 1000 cores at the problem, you'll have 999 cores processing the 90% in parallel, each performing 0.09% of the workload, but you'll still have 1 core doing the 10% that's serial. Those 999 cores will be idle for 99% of the time, waiting for that last core to finish.
The counter-argument is Gustafson's Law https://en.wikipedia.org/wiki/Gustafson%27s_law
This says that people don't choose a particular task, then wait for a computer to do it. Instead, the choice of which task to perform depends on what the computer can manage. Hence the user of a 1000 core machine will choose to do different tasks than the user of a 10 core machine, or a 1 core machine.
Whilst Gustafson's Law is clear from experience (a PlayStation 4 isn't used to run Pacman really fast), Amdahl's Law is the one that's relevant for compilers: a "sufficiently smart compiler" can alter your code in all sorts of ways, but the resulting executable must still perform the same task (otherwise it's a bug!).
There might be an approach based on e.g. writing an abstract specification and deriving a program which is suitable for the given hardware, but that's a long way off (for non-trivial tasks, at least).
Are consumers ever going to see a general purpose computer based on any of those technologies on their desktops?
Still gradually in development, presumably coming someday. It's very difficult to get quantum systems of any real complexity to not decohere before they can do useful computation.
> DNA computing
DNA-based storage apparently exists in the lab and might exist someday: http://arstechnica.com/information-technology/2016/04/micros.... Current costs are something like $40 million/GB, so that'll have to come down "a bit". I don't know of anyone doing computation with DNA, biology seems too slow for that.
> Optical computing
It's very difficult to make light interact with other light (which is sine qua non for computation) except inside a material with electrons, where you get significant losses and there's no real advantage over just plain old transistors. Not likely to become a useful technology IMO.