Online Labs – ARM servers in the cloud
labs.online.net
labs.online.net
The servers have four logical processors like this
processor : 0
model name : ARMv7 Processor rev 2 (v7l)
Features : half thumb fastmult vfp edsp thumbee fpv3 tls idiva idivt vfpd32 lpae
CPU implementer : 0x56
CPU architecture: 7
CPU variant : 0x2
CPU part : 0x584
CPU revision : 2
[1] https://community.cloud.online.net/ Hardware : Marvell Armada 370/XP (Device Tree)
Revision : 0000
Serial : 0000000000000000 free -m
total used free shared buffers cached
Mem: 2020 97 1923 0 6 40
-/+ buffers/cache: 49 1970
Swap: 0 0 0
Server seems to be located in France {
"as": "AS12876 ONLINE S.A.S.",
"city": "",
"country": "France",
"countryCode": "FR",
"isp": "Tiscali France",
"lat": 48.86,
"lon": 2.35,
"org": "Tiscali France",
"query": "212.47.232.90",
"region": "",
"regionName": "",
"status": "success",
"timezone": "Europe/Paris",
"zip": ""
}Someday comes a point where apps that actually are compute-bound might want to use more, slower cores for power/density/cost/etc.--I just don't think that cutover is tomorrow for the kind of apps (most of) you or I work on.
Further out: This is a Marvell-designed core that looks slower than the Cortex-A15-based Tegra K1 in a Chromebook (posted results elsewhere in the comments; it could be a clock-speed issue, not anything inherent to the core designs). Further out, there're some 64-bit ARM cores (Cortex-A57, X-Gene, Project Denver though that may not wind up in servers) and at process bumps (like TSMC 20nm). Related, check out http://www.anandtech.com/show/8580/hp-appliedmicro-and-ti-br... if you haven't. Of course, Intel isn't sleeping, and low-power x86 chips will improve, too; there will be 14nm versions of the Atom-based Xeons someday. As ever, fun times.
http://www.marvell.com/embedded-processors/armada-xp/
Along with low power, make it very interesting.
I can see someone put 64, 128 of them in 1U chassis. This might be interesting low cost system for someone who need simulcast live video streams to millions of users at really low cost.
"64 4 cores CPU each with integrated 16 GIGE ports for fan out live video stream that potentially fit inside 1U chassis"
If that hold true, 128 SOC in 1U can be 2-300 Watts. That can true ly go against x86 for "some" applications.
P.S. I believe in 64 in 2U, though.
ubuntu@c1-10-1-2-29:~$ dd if=/dev/zero of=test1 bs=1M count=512
512+0 records in
512+0 records out
536870912 bytes (537 MB) copied, 5.43834 s, 98.7 MB/s
ubuntu@c1-10-1-2-29:~$ dd if=test1 of=test2 bs=1M count=512
512+0 records in
512+0 records out
536870912 bytes (537 MB) copied, 5.869 s, 91.5 MB/s
ubuntu@c1-10-1-2-29:~$ dd if=test2 of=/dev/null bs=1M count=512
512+0 records in
512+0 records out
536870912 bytes (537 MB) copied, 0.60429 s, 888 MB/sI wonder what advantages this product have if any?
| Online Net | RK3188
Coremark (single) | 3288.391976 | 4745.333755
Coremark (dual cpu) | 6579.488445 | 8505.209441
Coremark (4-cpu) | 12985.958932 | 13930.001741
dhrystones | 3204101.2 | 5810575.0
linpack_dp | 96396.624 | 280673.07
linpack_sp | 141666.344 | 286485.812
nbench ASSIGNMENT | 5.4852 | 11.569
nbench BITFIELD | 1.8627e+08 | 3.1242e+08
nbench FOURIER | 4346.6 | 9248.2
nbench FP EMULATION | 94.734 | 143.8
nbench HUFFMAN | 1035 | 1444.5
nbench IDEA | 2035.4 | 1963.3
nbench LU DECOMPOSITION | 183.26 | 459.03
nbench NEURAL NET | 8.0897 | 13.392
nbench NUMERIC SORT | 573.47 | 733.99
nbench STRING SORT | 56.463 | 108.94
scimark Composite Score | 113.53 | 230.20
scimark FFT | 121.34 | 199.90
scimark LU | 92.63 | 279.92
scimark MonteCarlo | 64.20 | 81.53
scimark SOR | 191.28 | 420.49
scimark Sparse matmult | 98.23 | 169.17
stream Add | 1239.3447 | 1615.2325
stream Copy | 1168.6045 | 1147.0347
stream Scale | 926.4318 | 1599.6019
stream Triad | 1066.3372 | 1528.7064
stream_omp Add | 3159.7990 | 1271.3516
stream_omp Copy | 2603.3387 | 1193.8474
stream_omp Scale | 2372.8052 | 1653.1527
stream_omp Triad | 2595.0168 | 1245.1069Don't they have to follow YOU to be able to do that?
If the provider tells you what you're getting you can always ignore that information if you don't care. But some customers care about hardware specifics.
sudo rm / -rf --no-preserve-root
but it's not as fun as it sounds10$/mo droplet
$ sysbench --test=cpu --cpu-max-prime=2000 run
sysbench 0.4.12: multi-threaded system evaluation benchmark
Running the test with following options:
Number of threads: 1
Doing CPU performance benchmark
Threads started!
Done.
Maximum prime number checked in CPU test: 2000
Test execution summary:
total time: 1.5297s
total number of events: 10000
total time taken by event execution: 1.5219
per-request statistics:
min: 0.14ms
avg: 0.15ms
max: 4.68ms
approx. 95 percentile: 0.16ms
Threads fairness:
events (avg/stddev): 10000.0000/0.00
execution time (avg/stddev): 1.5219/0.00
C1 $ sysbench --test=cpu --cpu-max-prime=2000 run
sysbench 0.4.12: multi-threaded system evaluation benchmark
Running the test with following options:
Number of threads: 1
Doing CPU performance benchmark
Threads started!
Done.
Maximum prime number checked in CPU test: 2000
Test execution summary:
total time: 27.0053s
total number of events: 10000
total time taken by event execution: 26.9926
per-request statistics:
min: 2.69ms
avg: 2.70ms
max: 2.84ms
approx. 95 percentile: 2.72ms
Threads fairness:
events (avg/stddev): 10000.0000/0.00
execution time (avg/stddev): 26.9926/0.00If I'm not mistaken you only have 1 CPU for a droplet at this price.
ubuntu@c1-10-1-18-157:~$ sysbench --test=cpu --cpu-max-prime=2000 --num-threads=4 run
sysbench 0.4.12: multi-threaded system evaluation benchmark
Running the test with following options:
Number of threads: 4
Doing CPU performance benchmark
Threads started!
Done.
Maximum prime number checked in CPU test: 2000
Test execution summary:
total time: 6.7674s
total number of events: 10000
total time taken by event execution: 27.0485
per-request statistics:
min: 2.69ms
avg: 2.70ms
max: 7.00ms
approx. 95 percentile: 2.70ms
Threads fairness:
events (avg/stddev): 2500.0000/17.36
execution time (avg/stddev): 6.7621/0.00 sysbench 0.4.12: multi-threaded system evaluation benchmark
Running the test with following options:
Number of threads: 1
Doing CPU performance benchmark
Threads started!
Done.
Maximum prime number checked in CPU test: 2000
Test execution summary:
total time: 8.8170s
total number of events: 10000
total time taken by event execution: 8.8083
per-request statistics:
min: 0.83ms
avg: 0.88ms
max: 21.43ms
approx. 95 percentile: 0.95ms
Threads fairness:
events (avg/stddev): 10000.0000/0.00
execution time (avg/stddev): 8.8083/0.00
Total time is 2.4926s with --num-threads=4.I have 20 of those Xeon cores and 128Gb of RAM in a 2U.
Comparing the ratio of bogomips you'd have to get 598 of those ARM machines in a 2U to get the same bogomips.
Like I said this isn't even slightly scientific but is at least interesting trivia.
I can probably push 40 of those onto it without it bending too terribly. If I knock the RAM down to 2Gb an instance I could probably quite happily get 64-100 on it in theory. I think memory bandwidth might kill it before CPU does.
We have two almost full (18 each) 42U racks of those machines (bar switches) so across the 720 E5 cores with 4.6TiB of RAM there is about 4.3 million bogomips.
Fun :)
(most of this is corporate fileservers, exchange, AD, various crappy apps, network appliances, web servers, SQL servers and idles at around 20% in use). If it all went off you'd need earplugs and fireman's equipment.
However if you want one fast core, you're screwed :)
http://www8.hp.com/us/en/products/proliant-servers/product-d...
AMD will enter the market soon, too, but I think Applied Micro will hold its first mover advantage with its 3rd gen chips coming next year.