Neural network spotted deep inside Samsung's Galaxy S7 silicon brain
theregister.co.uk
theregister.co.uk
"Dynamic Branch Prediction with Perceptrons" (PDF)
http://www.atarihq.com/danb/files/64doc.txt:
"If an instruction does not store data in memory on its last cycle, the processor can fetch the opcode of the next instruction while executing the last cycle."
The 6502 also sometimes does excess reads and writes. For example, "Read-modify-write instructions (like INC) read the original data, write it back, and then write the modified data." (http://forum.6502.org/viewtopic.php?t=2).
One could call that speculative execution, where the processor writes back the value, just in case adding one to it doesn't increase it, and then, having discovered that adding one changes the value, writes the right value :-)
Either it's doing a read or it's doing a write, there is no option to use the bus unused. So the 6502 spams dummy reads and occasionally dummy writes whenever it's doing something internally.
Edit: And that's not the only interesting thing I've learned recently about the 6502, There is a very good reason why the stack is on page 1 and not some other constant page. There are only 3 constant values which can be pushed on the SB/ADH buses, 0xff, 0x00 and 0x01.
Ox00 is loaded into upper address register for zero page instructions. For Push operations, it loads 0x01 into the upper address register and puts the stack pointer + 0xff into the adder, which subtracts one from the stack pointer. (The 0xff is just the default bus value, if nothing else is pulling it low, so you get it for free)
For Pop operations, it pushes the same 0x01 into both the upper address register and the adder, at the same time.
So if they wanted to put the stack page at another constant page, or at a variable page they would also need to provide more gates to provide another constant one somewhere.
Edit: my layman knowledge about branch prediction and neural networks is showing its gross inadequacy :/
OTOH I see this as the optimization of hardware to fix the cap that software, paradigms and devs mean.
This is maybe not the best approach for certain games or benchmarks, but for normal use this has proven to be great.
Search google for "hashed perceptron branch predictor"
Edit: also, Intel branch predictors up to Core2 were well characterized and were not believed to be using NN.
[1] https://hal.inria.fr/hal-01100647/document ,discussed last year at HN.
Even a simple dot product can be called a "neural network," albeit a small uninteresting one. Your features could be (say) the state of cache, the number of jumps, the number of recent stalls / stack size / returns, and so on. Put them into a 10-dimensional vector or whatever, take the dot product between that and a set of 10 learned weights, and preload the branch if the result is above some threshold. Even a simple model could perhaps work really well here.
Of course it's not like they have some deep learning convolutional net running on the program input. I also bet the weights are set at the factory so it's not "learning" anything.
Learning on the fly would consume much more power as it back-propagates, so doesn't really make much sense as the potential performance gain would be marginal and would probably cause a performance hit when switching to a different process or thread.
> The perceptrons are trained by an algorithm that increments a weight when the branch outcome agrees with the weight’s correlation and decrements the weight otherwise.
A "perceptron" seems to be a single linear function, akin to a single neuron in a neural network. So complex learning algorithms like back-propagation don't apply here.
The complete process, from the paper, is:
> 1. The branch address is hashed to produce an index "i" into the table of perceptrons.
> 2. The "i"th perceptron is loaded from the table into a vector register "P" of weights.
> 3. The value of "y" is computed as the dot product of "P" and the global history register. [This is a shift register containing values of -1 or 1 for the last N branches seen. -1 means not taken, and 1 means taken.]
> 4. The branch is predicted "not taken" when "y" is negative, or "taken" otherwise.
> 5. Once the actual outcome of the branch becomes known, the training algorithm uses this outcome and the value of "y" to update the weights in "P".
> 6. "P" is written back to the "i"th entry in the table.
The article says that AMD's microarchitecture uses a hashed perceptron system in its branch prediction. So, from the above, it seems that "hashed" means that you hash the branch address to pick which perceptron (model) to use, and "perceptron" is just a dot product. (Actually the wikipedia article[2] suggests that the name "perceptron" refers to the training algorithm, i.e. how you update the weights.)
[1]: Daniel A. Jiménez and Calvin Lin. 2002. Neural Methods for Dynamic Branch Prediction. https://www.cs.utexas.edu/~lin/papers/tocs02.pdf
> AMD's Zen architect Mike Clark confirmed to us his microarchitecture uses a hashed perceptron system in its branch prediction. "Maybe I should have called it a neural net," he added.
Neural networks can be very complicated, but they can also be very simple. What little I know about perceptrons is that they are very simple. If you google "hashed perceptron", many of the papers that come up mention branch prediction in their titles.
I'm not as familiar with mobile phone architecture as with that of PCs, but the number and types of operations influenced by the device's various sensors (light sensor, GPS, accelerometer, gyroscope) could conceivably be more than normal, naïve branch prediction can handle effectively.
Also I do not see how the number or quality of sensors could in any way affect the prediction rate.
http://www.redmondpie.com/galaxy-note-7-vs-iphone-6s-real-wo...
Video (3:29 sec): https://www.youtube.com/watch?v=3-61FFoJFy0
PS: I highly recommend you check out the video... iPhone completely obliterated Note 7.
Unless they do something abusive like keep the speaker or mic on, which then permits them to be more abusive like keep location services active, which then rapidly drains battery.
http://solutionowl.com/the-ultimate-guide-to-solving-iphone-...
http://9gag.com/gag/a5PmrLq/technically-correct-man-the-man-...
Still, it's interesting to see the single A9 with coprocessor outperform dual M1s - like was said about the reduced core count in the Snapdragon 820:
“The most important thing to have is peak single-threaded performance, as most of the time only one or two cores are active." [0]
This is apparently true, with the iPhone outperforming despite its A9 running up to 400mhz slower than the Android cohort's Snapdragon.
[0] http://trustedreviews.com/opinions/snapdragon-820-vs-snapdra...
But for a hardware test, I'd expect a single running app, executing shared native codebase, in no disruptions airplane mode, measuring specific part of the hardware. And that's not even close to measuring before/after improved branch prediction, which could be useful in both phones.
The test is a valid observation, but is irrelevant to this article, or to the technology.
Why measure something that doesn't show up in a standard user's test? Here are some examples: lower power usage, lower latency of small operations (rather than throughput of large actions), smaller design (saving chip space), etc. And finally - if this is a good technology, Apple can use it too and get even faster.
I would argue that would be a completely useless test unless you're explicitly testing something for development purposes only. Users use phones. Nothing else matters but the user experience. If you can make your phone beat another phone in your limited, no disruptions, network traffic off example that doesn't mean anything.
I feel this is a fair test. You test the phones at what they're supposed to be able to do. If one does it better it's legit to point that out regardless of what the underlying structure looks like.
Bringing up user experience when talking about branch predictors is meaningless. Just as bringing up branch predictors when talking about perceived performance. (unless you're proving that this exact feature being available/not available, with other things being controlled for, makes a major difference in the test)
If working on branch predictors does not impact the user experience, then arguably that work is meaningless in the context of building real products -- it may still be interesting for academic purposes.
But, for the final user, it doesn't matter if each technology alone is better than another. If you have a background 64-core processor, driven by another 8-core... What matters is how fast something loads, performs, and how long batteries last.
Of course, this might be outplaced in a thread discussing NN branch prediction.
Err? It seems to me that the only tests that make any sense are those that count for end users. Everything else is pointless. It doesn't matter if it's the same system or code base or whatever other engineering fetish.
Users who compare their phones to see which goes faster look at how the same app opens and runs on two separate phones. Which incidentally is what this test is about. It's the only sensible test you can do.
If you are specifically comparing two specific phones against each other then yes this is the only sort of test that matters, and if you are comparing manufacture's flagship models then that is what you are doing.
From every other point of view, either comparing performance of individual components or comparing like-for-like performance of devices with very similar spec it is less valid though: the screen resolution is a huge variable for a game so it isn't a "fair" test to compare performance of two devices like that. You could equally state that the higher resolution produces a better result because of the resolution difference and you would similarly be called to task by for not comparing like for like. Maybe the a device with the same hardware other than the better screen would outperform the iDevice on the test, or at least not appear to underperform by as much, this test can't tell us that.
This is difficult of course because few phone models are practically identical and few vary by just one factor, and can make a developers life more difficult when trying to support a varied market like Android based devices.
Better understanding individual components lets engineers build better overall systems. Millions of choices and decisions go into something like a phone, most of which can't have their result easily measured by just looking at the end result. But if enough of them are good, you end up with the iPhone.
If enough of them are bad, you end up with mass recalls, and a class action lawsuit because your phones catch fire when charged overnight because your engineers didn't think to test that specific use case, didn't test their capacitors, didn't test their QA process, didn't unit test their battery management code, or otherwise failed to indulge their "fetish" (read: job.)
Proper hardware tests might be pointless to end users, but that doesn't make them pointless.
Anything besides the concrete end-user experience of a system doesn't matter at all.
(In theory it could matter for cpu/compiler designers etc -- but even what matters to them is inconsequential if in the end of the line it doesn't matter to the end-user).
On the other hand, without a neural net, it might be even slower. Android supports tons of different architectures and platforms, of course you can't optimize it like Apple can with a single platform.
Not really. Sure Android can run on multiple architectures but 99% of the time it's mostly the same. Why couldn't they provide optimizations for multiple architectures anyway? Whenever something can run on more than 1 architecture I usually see the argument about how it can't be as optimized on any specific architecture but I don't see why not. Optimize for the 80% use case and you're probably covered. Go beyond that to get more performance out of other platforms but it's unlikely necessary.
Because Android is usually developed by at least two companies: Google does generics and OEM which is responsible that Android runs smoothly on their hardware. Sure, given enough time and resources, you could do optimization for every hardware combination possible, but that is not how things work in real world, especially where release cycles are such insane as in smartphones market.
You make it sound as if every version of CPU, GPU and chipset have to be optimized in every combination possible but typically you do optimizations based on various CPU targets, optimizations based on GPU targets, etc; the vast majority of smart phones use the same architecture and many end up targeting the same CPUs and GPUs (I mean how many phones ended up / still use the Snapdragon 820?).
Yes they can't optimize for every combination but optimizing for the typical CPUs and GPUs seems like something they're likely already doing. If the stack was squished and Google made it from hardware to software I'm not sure that they would be necessarily doing optimizations any different on the hardware / driver level.
My recently purchase S7 is the first ever cellphone I return to the store because: 1) It's way too hard to flash their exynos based hardware and 2) the default install comes with 14Gb of garbage included. That includes the full microsoft office and a slew of garbage apps I'll never use (I'd include the launcher in that, but I understand some people like it, so I'll it a pass)
Annecdata, but I have now become this guy who will tirelessly slap Samsung's garbage software "strategy" every time I have the opportunity.
edit: The article points out that the resolution of the Samsung display is a lot higher, so it's not quite comparing apples and oranges here. I'm guessing that a much higher resolution is pure win when drawing mostly vector graphics i.e. apps but really kills you when loading games as presumably they're paging in much larger assets from disk.
Yep, it is. D: Well here is a tip that should be somewhat obvious: unscientific measurements of which phone loads apps faster when you click from the home screen isn't any basis to claim a phone "beats" the other, especially when we were talking about processor performance.
By the way, the iPhone beats the S7 on single core and the S7 wins on multi core performance. Which makes sense seeing as you are competing a dual core with an octa core.
Disclaimer: as I said, this is an analogy. Brains are not literally executing instructions in a pipeline.
The iOS works well with the iPhone simply because it is fine tuned for the product. On the androids side the OS is tuned to add bloatware of the respective HW manufactures. To add the so called speed the HW manufactures put in more RAM.
It seems like they go for redundancy, use two suppliers etc for common parts and chips. Which means these are built based on the designs of Apple. Samsung just manufactures them.
I don't follow closely though so I might be wrong.
Now the original iphone was probably 80% Samsung design and manufacturing, but these days the design part is closer to 0%.