Command-line Tools can be 235x Faster than a Hadoop Cluster (2014)
adamdrake.com
adamdrake.com
My response: nope, MySQL! (plus an ORM, a very hated programming language and very few config optimisations on the server side).
This project DB has a couple of tables with thousands of records, not billions, and, for now, a few users (40-50).. a good schema and a few well done queries can do the trick.
I guess some people are so used to see sluggish application that as soon as they see something that goes faster than average they think it must use some cool latest big tech.
Lots of people seem to want some silver bullet magical software/technology to solve their problems instead of learning how to use the tools they have. That's not software development.
Well, I'm not sure "seniority" is the right word - the more tech stuff you know, in general, the _less_ seniority you're going to achieve in terms of org charts, decent seating, respect and actual pull within an organization. You can achieve job security and higher pay that way, though.
Large companies would benefit from dev not over-engineering simple apps and spending less time on their own tools and ticket systems, and spending more time instead on solving/automating more problems for the company.
[1] https://www.tamingdata.com/wp-content/uploads/2010/07/tree-s...
Because it fit in the CPU cache the speedup was around 10000x on a single machine (numerical simulations, amirite?).
Because it was so much faster all the code required to split it up between a bunch of servers in a map reduce job could be deleted, since it only needed a couple cores on a single machine for a ms or three.
Because it wasn't a map-reduce job, I could take it out of the worker queue and just handle it on the fly during the web request.
Sometimes it's worth it to just step back and experiment a bit.
Good understanding of data access patterns and the right algorithm go a long way in both spaces as well.
Computing history's not a circle, but it's damn sure a spiral.
I think that is my new favorite phrase about computing history. Everything old is new again. There's way too much stuff we can extract from past for current problems. It's kind of amazing.
https://platypus1917.org/2011/06/01/lenins-liberalism/
A similar sentiment is that history doesn't repeat but it does rhyme.
In that example the idea is that feudalism represents a form of "community property", which though owned by a feudal lord in the final instance is in practice shared. Capitalism then "negates" that by making the property purely private, before communism negates that again and reverts to shared property but with different specific characteristics.
The three key principles of Hegels dialectics comes from Heraclitus (the idea of inherent conflict within a system pulling it apart), Aristotle (the paradox of the heap; the idea that quantitative changes eventually lead to qualitative change), and Hegel himself (the "negation of the negation" that was popularised by Marx in Capital; the idea that a driven by inherent conflict, qualitative changes will first change a system substantially, before reversing much of the nature of the initial change, but with specific qualitative differences).
The idea is known to have been simultaneously arrived at by others in 19th century too, at least (e.g. at least one correspondent of Marx' came up with it independently), and it's quite possible variations of it significantly predates both Marx and Hegel as well.
If it was a circle we'd have better code and framework reuse. Sigh.
(Obviously this wouldn't work for FPSes, though, since hitscan weapons mess the independence-of-tiles logic all up.)
What's really ironic is the PS3 actually had the right architecture for this, the SPUs only had 256kb of directly addressable memory so you had to DMA everything in/out which forced you to think about memory locality, vectorization, cache contention, etc. However X360 hit first so everyone just wrote for its giant unified memory architecture and then got brutalized when they tried to port to PS3.
What's even funnier in hindsight is all that work you'd do to get the SPUs to hum along smoothly translated to better cache usage on the X360. Engines that went <prev gen> -> PS3 -> X360 ran somewhere between 2-5x faster than the same engine that went <prev gen> -> X360 -> PS3. We won't even talk about how bad the PC -> PS3 ports went :).
One awk command later my file was flipped. Run same exact sort command on this but without specifying field. Completed in 12 seconds.
Morals:
1. Small changes can have a 3+ orders of magnitude effect on performance
2. Use the Google, easier than understanding every tool on a deep enough level to figure this out yourself ;)
wget "https://data.cityofnewyork.us/api/views/kku6-nxdu/rows.csv" rows.csv
sqlite3
.mode csv
.import ./rows.csv newyorkdata
SELECT *
FROM newyorkdata
ORDER BY `COUNT PARTICIPANTS`; wget "https://data.cityofnewyork.us/api/views/kku6-nxdu/rows.csv"
csvsql --query "SELECT * FROM newyorkdata ORDER BY `COUNT PARTICIPANTS`" rows.csv > new.csvSpecific command:
LC_ALL=c sort -n filename -r > output
k2,2
have been faster, as it limits to just the second field?How large was the 10 million line file, and do you know whether sort was paging files to disk?
Used these commands:
cat /dev/urandom | tr -dc 'a-zA-Z0-9' | fold -w 32 | head -n 10000000 > text
cat /dev/urandom | od -DAn |tr -s ' ' '\n' |head -10000000 | sed -e 's/$/\/9.0/'> num
bc -l num > num2
(that's to create decimal numbers)
paste -d, text num2 > textfirst
paste -d, num2 text > numfirst
time LC_ALL=c sort -n numfirst -r > output
real 0m10.712s user 0m21.959s sys 0m1.644s
time LC_ALL=C sort --field-separator=',' -k2 -n -r textfirst > output
real 0m9.638s user 0m21.940s sys 0m1.293s
> The COST of a given platform for a given problem is the hardware configuration required before the platform outperforms a competent single-threaded implementation. COST weighs a system’s scalability against the overheads introduced by the system, and indicates the actual performance gains of the system, without rewarding systems that bring substantial but parallelizable overheads.
However for iterative calculations like pagerank the overhead for distributing the problems is often not worth it.
The target data went fro almost 200MB of SQLite to 7MB of binary that could just be mapped into memory. Oh, and lookup on the device also became 1000x faster.
There is a LOT of that sort of stuff out there, our “standard” approaches are often highly inappropriate for a wide variety of problems.
iOS used to kill processes above 50 megs of RAM usage. Good times!
Not so sure it's really obsolete, because connections are often spotty and local data can provide some initial information while waiting for the servers.
Github?
But I've sent it to you in email - check it out if you're interested.
(And some would say that it then went to "optimize everything for resume keywords," which is almost always inappropriate, but I don't want to be too cynical.)
Assuming the locality of data is not a big issue. It can be extremely fast. However, depending on the system architecture reading from the drives can be a bottle neck. If your system has enough drives for enough parallel reads you will turn through data pretty quickly. Moreover from my experience most systems or clusters with a few petabytes have enough drives that one can read quite a lot data in parallel.
However, the worst is when the data is referencing non-local pieces of data. So your processing thread will have to fetch data from either a different node or data not in main memory to finish processing. This can be pain since that means either the task is just nor really parallelize-able or the person who originally generated the data did not take into account certain groups of data may be referenced with each other. Sadly, what happens here is that commutation or read costs for each thread start to dominate the cost of the computation. If it's a common task and your data is fairly static it makes sense to start duplicating data to speed up things. Also restructuring data can also be quite helpful and pay off in the long run.
I like that better than bringing racks into it, because once you have multiple machines in a rack you've got distributed systems problems, and there's a significant overlap between "big data" and the problems that a distributed system introduces.
Single-server setups larger than 2U but (usually) smaller than 1 rack can give tremendous bang for the buck, no matter if your "bang" is peak throughput or total storage. (And, no, I don't mean spending inordinate amounts on brand-name "SAN" gear).
There's even another category of servers, arguably non-commodity, since one can pay a 2x price premium (but only for the server itself, not the storage), that can quadruple the CPU and RAM capacity, if not I/O throughput of the cheaper version.
I think the ignorance of what hardware capabilities are actually out there ended up driving well-intentioned (usually software) engineers to choose distributed systems solutions, with all their ensuing complexity.
Today, part of the driver is how few underlying hardware choices one has from "cloud" providers and how anemic the I/O performance is.
It's sad, really, since SSDs have so greatly reduced the penalty for data not fitting in RAM (while still being local). The penalty for being at the end of an ethernet, however, can be far greater than that of a spinning disk.
At a theoretical level, as a sysadmin, I learned the theoretical capabities by reading datasheets for CPUs, motherboards (historically, also chipsets, bridge boards, and the like, but those are much less irrelevant), and storage products (HBAs/RAID cards, SAS expander chips, HDDs, SSDs). Make sure you're always aware of the actual payload bandwidth (net of overhead), actual units (base 2 or base 10) and duplex considerations (e.g. SATA).
From a more practical level, I look at various vendors' actual products, since it doesn't matter (for example) if a CPU series can support 8 sockets if the only mobos out there are 2- and 4-socket.
I also look at whatever benchmarks were out there to determine if claimed performance numbers are credible. This is where sometimes even enthusiast-targeted benchmark sites can be helpful, since there's often a close-enough (if not identical) desktop version of a server CPU out there to extrapolate from. Even SAS/SATA RAID cards get some attention, not in a configuration worthy of even "medium" data, but enough for validating marketing specs.
As a software engineer who builds their own desktops (and has for the last 10 years) but mostly works with AWS instances at $dayjob, are there any resources you'd recommend for learning about what's available in the land of that higher-end rackmount equipment? Short of going full homelab, tripling my power bill, and heating my apartment up to 30C, I mean...
That's probably better, since it'll scale a bit better with technological improvements. The problem is, it doesn't have quite the clever sound to it, especially with the numbers and dollars.
Now, the other main problem is that, though the cost of a workstation is fairly well-bounded, the cost of that medium-data server can actually vary quite widely, depending on what you need to do with that data (or, I suppose, how long you might want to retain data you don't happen to be doing anything to right at that moment).
I suppose that's part of my point, that there's a mis-perception that, because a single server (including its attached storage) can be so expensive, to the tune of many tens of thousands of (US) dollars, that somehow makes it "big" and undesireable, despite its potentially close-to-linear price-to-performance curve compared to those small 1U/2U servers. Never mind doing any reasoned analysis of whether going farther up the single-server capacity/performance axis, where the price curve gets steeper is worth it compared to the cost and complexity of a distributed solution.
> are there any resources you'd recommend for learning about what's available in the land of that higher-end rackmount equipment?
Sadly, no great tutorials or blogs that I know of. However, I'd recommend taking a look at SuperMicro's complete-server products, primarily because, for most of them, you can find credible barebones pricing with a web search. I expect you already know how to account for other components (primarily of concern for the mobos that take only exotic CPUs).
As I alluded in another comment, you might also look into SAS expanders (conveniently also well integrated into some, but far from all, SuperMicro chassis backplanes) and RAID/HBA cards for the direct-attached (but still external) storage.
That's primarily because I'm aware of the variability that random access injects into spinning disk performance and that 10GE is now common enough that it takes more than just a single (sequentially accessed) spinning disk to saturate a server's NIC.
Plus, if you're talking about a (single) local spinning disk, I'd argue that's a trivial/degenerate case, especially if compared to a more expensive SSD. Does my assertion stand up better if I it had "of comparable cost" tacked on? Otherwise, the choice doesn't make much sense, since a local SSD is the obvious choice.
My overall point is that, though one particular workload may make a certain technology/configuration appear superior to another [1], in the general case, or, perhaps most importantly, in the high performance case, to have an eye on the bottlenecks, especially the ones that carry a high incremental cost of increasing their capacity.
It may be that people think the network, even 10GE now, is too cheap to be one of those bottlenecks, arguably a form of fallacy [2] number 7, but that ignores the question of aggregate (e.g. inter-switch) traffic. 40G and 100G ports can get pricey, and, at 4x and 10x of a single server port, they're far from solving fallacy number 3 at the network layer.
The other tendency I see is for people not to realize just how expensive a "server" is, by which I mean the minimum cost, before any CPUs or memory or storage. It's about $1k. The fancy, modern, distributed system designed on 40 "inexpensive" servers is already spending $40k just on chasses, motherboards, and PSUs. If the system didn't really need all 80 CPU sockets and all those DIMM sockets, it was money down the drain. What's worse, since the servers had to be "cheap", they were cargo-cult sized at 2U with low-end backplanes, severely limiting existing I/O performance. Then, to expand I/O performance, more of the same servers [3] are added, not because CPU or memory is needed, but because disk slots are needed and another $4k is spent to add capacity for 2-4 disks.
[1] This has been done on purpose for "competitive" benchmarks since forever
[2] https://en.wikipedia.org/wiki/Fallacies_of_distributed_compu...
[3] Consistency in hardware is generally something I like, for supportability, except it's essentially impossible anyway, given the speed of computer product changes/refreshes, which means I think it's also foolish not to re-evaluate when it's capacity-adding time after 6-9 month.
Is it safe to say that such situations are often found with embedded or otherwise specialized hardware?
Sure, a lot of tyros are working on distributed systems because it's cool or because it enhances their resumes, but there are also a lot of professionals working on distributed systems because they're the only way to meet requirements. Cherry-picking examples to favor your own limited skill set doesn't seem like engineering to me.
But the article is kind of wrong. It depends on your data size and problem - you can even use commandline-tools with Hadoop Map/Reduce and the Streaming API and Hadoop is still useful if you have a few terabytes of data that you can tackle with map and reduce algorithms and in that case multiple machines do help quite a lot.
anything that fits on your local ssd/hdd probably does not need hadoop... however you can run the same unix commands from the article just fine on a 20tb dataset with Hadoop.
Hadoop MapReduce/HDFS is a tool for a specific purpose not magic fairy dust. Google did build it for indexing and storing the web and probably not to calculate some big excel sheets...
a) resume enrichment by way of buzzword addition b) huge budget grants and allocation, purportedly for lofty goals while management is really unaware of real technology needs/options
Much has been talked about this already; sharing them again!
[1] https://news.ycombinator.com/item?id=14401399
[2] https://www.chrisstucchio.com/blog/2013/hadoop_hatred.html
[3] http://widgetsandshit.com/teddziuba/2010/10/taco-bell-progra...
[4] https://www.reddit.com/r/programming/comments/3z9rab/taco_be...
-- my actual employer
(the actual problem is our SAN is getting us a grand total of 10 iops. No, that's not wrong, 10. The "what the christ, just buy us a system that actually works" momentum has been building for 2 years now but hasn't managed to take over. Hardware comes out of a different budget than wasted time, yo.)
1) minimize hardware spend at all costs
2) maximize lost time (if spent under the manager's control)
Really this just goes to show how impressive RDMSs like Postgres are. There's nothing out there that's drastically better in the general case. So alternative database systems tend to just nibble around the edges.
My rule of thumb is always try to implement it using a relational model first, and only when that proves itself untenable, look into other more specialized tools.
Ever notice how with cloud it's actually in the interest of cloud providers to have the worst programmers, worst possible solutions for their clients ? Those will maximize spend on cloud, which is what these companies are going for.
Of course, they hold up the massive carrot of "you can get a job here if you ...".
I hope the NoSQL hype is over by now and people are back to choosing relational as the default choice. (The same people will probably chasing block-chain solutions to everything by now...)
Most nosql SaaS platforms are significantly easier to consume than the average service oriented RDBMS platform. If all the DBMS is doing is handling simple CRUD transactions, there's a good chance that relational databases are overkill for the workload and could even be harmful to the delivery process.
The key is to take the time to truly understand each workload before just assuming that one's preferred data storage solution is the right way to go.
You can have minimally relational data such as URL : website, but that's still improved by going URL : ID, ID : website because you can insert those ID's into the website data.
Now plenty of DB's have terrible designs, but there I have yet to year of actually non relational data.
But, there are plenty of other ways to slice and dice that data, for example a URL is really protocol, domain name, port, path, and parameters, etc. So, it's a question of how you want to use it.
PS: Using a flat table structure (ID, URL, Data) with indexes on URL and ID is really going to be 2 or 3 tables behind the scenes depending on type of indexes used.
I'm also approaching this from the perspective of someone who spends more time in the ops world than development. I won't argue that NoSQL would ever "outperform" in the realms of data science and theory, but I question whether a business is going to see more positive impact from finding the perfect normal form from their data or having more flexibility in the ability to deliver new features in their application.
* I'm fully aware that this can be as much of a curse as a blessing depending on the data and the architecture of the application, which reenforces understanding the data and the workload as a significant requirement.
Although that is true in principle, in reality that results in the messes I see around me where a small startup (but this often goes for larger corps too) has a plethora of tech running it fundamentally does not need. If your team’s expertise is Laravel with MySql then even if some project might be a slightly better fit for node/mongo (does that happen?), I would still go for what you know vs better fit as it will likely bite you later on. Unfortunately people go for more modern and (maybe) slightly better fit and it does bite them later on.
For most crud stuff you can just take an ORM and it will handle everything as easily as nosql anyway. If your delivery and deployment process have a rdbms, it will be natural anyway and likely easier than anything nosql unless it is something that is only a library and not a server.
Also, when in doubt, you should take a rdbms imho, not, like a lot of people do, a nosql. A modern rdbms is far more likely to fit whatever you will be doing, even if it appears to fit nosql better at first. All modern dbs have document, json/doc storage built in or added on (plugin or orm) : you probably do not have the workload that requires something scaleout like nosql promises. If you do, then maybe it is a good fit, however if you are conflicted it probably is not anyway.
No, there are not. In 99% of applications, the data is able to be modeled relationally and a standard RDBMS is the best option. Non-relational data is the rare exception, not the rule.
Apparently, in 2017, that was somewhere around 16 terabytes. https://www.theregister.co.uk/2017/05/16/aws_ram_cram/ Heck, you can trivially get a 4TB instance from Amazon nowadays: https://aws.amazon.com/ec2/instance-types/
The biggest DBs I've worked on have been a few tens of billions of rows, and several hundreds of gigabytes. That's like... nothing. A laughable start. You can make trivially inefficient mistakes and for the most part e.g. MySQL will still work fine. Absolutely nowhere near "big data". And odds are pretty good you could just toss existing more-than-4TB-data through grep and then into your DB and still ignore "big data" problems.
BTW in R data.table is better for larger datasets than the default dataframe... or use the Revolution Analytics stuff on a server.
It makes me want to cry, knowing we handled that with a single server and a relational database 10 years ago.
Lets also not forget that everyone today forgets that the majority of data actually has some sort of structure. There is no point in pretending that every piece of data is a BLOB or a JSON document.
I have given up on our industry ever becoming sane, I now fully expect each hype cycle to be pushed to the absolute maximum. Only to be replaced by the next buzzword cycle when the current starts failing to deliver on promises.
I sometimes wish I was unscrupulous enough to cash in on these trends, I’d be a millionaire now. Instead I’m just a sucker who tries to make solid engineering decisions on behalf of my employers and clients. It’s depressing to think what professionalism has cost me in cash terms. But you’ve got to be able to look at yourself in the mirror.
"Resume driven development" is a great term.
IIRC the term "big data" was coined to refer to data volumes that are too large for exiting applications to process without having to rethink how the data and/or the application was deployed and organized. Thus, although the laptop reference is on point, large enough data volumes that make Pandas choke in a desktop environment does fit the definition of big data.
What about too large for Excel, is that “big data” too?
(Neither are)
This is obviously an over-exaggeration, but I don't think a dataset that breaks Pandas in a desktop environment even comes close to big data.
At the time, I thought "wow - this is big!", and it was for me.
[0] we didn't call it microservices back then though that's clearly what it was. The term I used was "event-driven, loosely coupled, transactional programs processing units of work comprised of a day's worth of meter readings per meter". Doesn't roll off the tongue quite so easily.
For 16TB you'll definitely get benefits if you are doing something embarrassingly parallel and processor bound over a cluster of a handful of machines. It's just parallelism, and it's a good thing.
I totally agree about the hundreds of Gigs - that is unless you are in a setting where many teams need to access that database and join it with others, in which case a proper data lake implemented on something beefy is a good idea. Hadoop has the benefit of distribution and replication, but another data warehouse might work better - say Oracle or Teradata if you are a small shop.
100Gb of data may not be that much to you, but it's another story if you have to process all of it every minute. Also everything is fine if you have nicely structured ( and thus indexed ... ) data, but I do have a few 100s Gb of images as well...
cat *.pgn | grep "Result" | sort | uniq -c
This pipeline has a useless use of cat. Over time I've found cat to be kind of slow as compared to actually passing a filename to a command when I can. If you rewrite it to be: grep -h "Result" *.pgn | ...
It would be much faster. I found this when I was fiddling with my current log processor to analyze stats on my blog. cat /dev/zero | pv > /dev/null
pv > /dev/null < /dev/zero
The second one is much faster. Argument list too long! ulimit -s 65536
https://unix.stackexchange.com/a/45584/115135find -name \*.pgn -print0 | xargs -0 grep -h "Result"
Each grep invocation will consume a maximum number of arguments, and xargs will invoke the minimum number of greps to process everything, with no "args too long" errors.
Tried spark on my laptop: waste if time. After 4h I killed al processes because it didn't read 25% of the file yet.
Same for hadoop, python and pandas, and a shiny new tool from google whose name I forgot long time ago.
Finally I installed cygwin con My laptop and 20 minutes later 'cut' gave me the results file I needded.
cut is line oriented, like most Unix style filters. It needs to keep only one line in memory at most.
If you say:
pd.read_csv(f)[[x,y,z]]
It has to read and parse the entire 28GB into memory (because it is not lazily evaluated; cf Julia).If you actually need to operate on three columns in memory and discard the rest, you should:
pd.read_csv(f, usecols=[x,y,z])
Then you get exactly what you need, and avoid swapping.The lack of lazy evaluation does inhibit composition--just look at the myriad options in read_csv(), some of which are only there to enable eager evaluation to remain efficient.
Parsing isn't actually a tough problem – https://github.com/dw/csvmonkey is a project of mine, it manages almost 2GB/sec throughput _per thread_ on a decade old Xeon
Okay, sounds reasonable, that's larger than most machines' memory...
> Tried spark on my laptop
Yikes, how did we get here? Not to shame you or anything, but that's like two orders of magnitude smaller than the minimum you might consider reaching for cluster solutions...and on a laptop...I'm legitimately curious to hear the sequence of events that led you to pick up Spark for this.
In my opinion this isn't something you should be leaving the command line for. I'm partial to awk; if you needed to get, say, columns 3, 7, 6, 4 and 1 for every line (in that order):
awk '{print $3"\t"$7"\t"$6"\t"$4"\t"$1}' file.csv
...where you're using tab as the delimiter for the output.While I'm at it, this is my favorite awk command, because you can use it for fast deduplication without wasting cycles/memory on sorting:
awk '!x[$1]++' file.csv
I actually don't know of anything as concise or as fast as that last command, and it's short enough to memorize. It's definitely not intuitive, however...Extracting only the lines with "[Result]" in them into new file (using grep) takes about 3 seconds.
Importing that into a local Postgres database on my laptop takes about 1.5 seconds.
Then running a simple:
select result, count(*)
from results
group by result;
Takes about 0.5 seconds.So the total process took only 5 seconds (and now I can run much more aggregation queries on that)
The further my clients move to the cloud, the more shell scripts they write at the exclusion of other languages. and just like this, I have clients who have ripped out expensive enterprise data streaming tools and replaced them with bash.
The future of enterprise software is going to be a bloodbath.
We generate svg maps in psuedo-realtime from this data-set : 2MB maps render sub-second over the web, which feels 'responsive'.
I only mention this as many marketing people will call 50Million rows or 1TB "Big Data" and therefore suggest big / expensive / complex solutions. Recent SSD hosts have pushed up the "Big Data" watermark, and offer superb performance for many data applications.
[ yes, I know you can't beat magnetic disks for storing large videos, but thats a less common use-case ]
a) basically a preprocessed static inverted index on keywords [ using GIN / tsv_ etc ]
b) geo location index on GIS geometry field - relying on postGIS to be performant over geo queries
The data changes infrequently, so these are computed at time of import / regular update.
We do have a fair amount of ram set aside for postgres [ circa 20GB ]
[0]: https://www.chrisstucchio.com/blog/2013/hadoop_hatred.html
Many thanks for the feedback and comments. If you have any questions, I'm also happy to try to answer them.
I'm also working on a project, https://applybyapi.com which may be of interest to anyone here hiring developers and drowning in resumes.
The info available to non-signed-up people is simply not enough. It would be great if you had a demo account so one can get a feel for how this works.
E.g. I couldn't find info on why I have to add a job description. Wouldn't I post job board, with the posting liniing to the ApplyByAPI test?
I could also not find any info on the tests candidates have to do. Yes, APIs, I get it. But an example wouldn't hurt.
Interesting idea!
We will certainly consider those points, and try to make things both more clear from a communication perspective and also consider adding some demo account (great idea!).
To answer the question about the job description, you're right that most customers so far have a post on a jobs board which they use to get traffic over to their ApplyByAPI posting. Once the candidate is there, they get a page with the job description and some API information they can use to generate their application.
We'll work on improving the language and information presentation to potential customers, and thank you again for taking the time to give your perspective!
MonrtDB: columnar DB for a single machine with a psql like interface: https://www.monetdb.org/
Where MapReduce Hadoop etcera shine is when a single dataset one of many you need to process is bigger than the biggest single disk availible - this changes with time
Back when I did M/R for BT the data set sizes where smaller - still having all of the Uk's larges single PR!ME superminis running your code was dam cool - even though I used a 110 baud dial up print terminal to control it.
1. Cache line misses. 2. So called definition of BigData. (if data can be easily fit into memory, then its not Big period! )
Many times, I have seen simple awk / grep commands will outperform Hadoop jobs. I personally feel, its lot better to spin up larger instances, compute your jobs and shut it down than bearing the operational overhead of managing hadoop cluster.
... | mawk '/^\[Result' ...On the other hand, for more intensive work, I’ve found it doesn’t perform as well as JVM-based applications. A lot of this is either JRuby or data models in Java doing interesting things with data structures.
Finally, for sheer processing speed, we’ve started converting some of our components to Rust. This has a 3-4x speed advantage to Java and Go for similar implementations.
Again - all anecdotes and I’m not really in a position to deeply explain how it all works (proprietary, etc). Just my two cents.
For dev tools, oh hell yes. Stick to Java. Go comes with some nice ones out-of-the-box which is always appreciated, but the ecosystem of stuff for Java is among the very best, and even the standard "basic" stuff vastly out-does Go's builtins.
[1]: https://twitter.com/brianhatfield/status/634166123605331968?... and there's also a blog/video? about these same charts and how it got there. pretty neat transformations.
I’ll be in my corner doing asking “do we really need this?”/“have you tested this?”
On the other hand, processing data over several steps with a homegrown solution needs a lot of programming discipline and reasoning, otherwise your software turns into an unmaintainable and unreliable mess. In fact this is the case where I work right now...
the command-line solution probably took that person a couple or more hours to perfect
If you want to find out what tools are available on the CLI that operate on jpegs, you'd try 'apropos jpeg' and get a list of things that mention jpeg in their man page.
The real issue is why do we silo ourselves so? It's fucking stupid. I do so much more than Ruby but conveying that to the new breed of tech people is just impossible. All they see is your biggest resume badge, and all they care about is how wide you will spread your legs for them.
Find (take a look at -exec option)
Cut (or awk)
Sed (for simple text/string substitutions)
Xargs
Dd (that beast can do charset transtation from ASCII to EBCDIC to use in mainframes and can also wipe disks)
And bash of course :)
I've got some similar ones for work that take data csv like data and output sql. A little bit of vim-foo on the csv (yy10000p) and I've got all the test data I need.
"awk"
brb 15 years
Not a program, but another fun bash thing I've learned about recently is brace expansion: https://www.linuxjournal.com/content/bash-brace-expansion
-P max-procs
Run up to max-procs processes at a time; the default is 1.
If max-procs is 0, xargs will run as many processes as
possible at a time. Use the -n option with -P; otherwise
chances are that only one exec will be done.
-n max-args
Use at most max-args arguments per command line.
Fewer than max-args arguments will be used if the size
(see the -s option) is exceeded, unless the -x option is
given, in which case xargs will exit.
So here is a "map/reduce" job which takes the log files in a directory
and processes them in parallel on eight CPUs, then combines the results. find . -name "*.log" | sort | xargs -n 1 -P 8 agg_stats.py | sort | merge_periods.pygotta SSH into 100 machines and ask them all a simple question? xargs will trivially speed that up by at least 10x, if not better, just by parallelizing the SSH handshakes.
https://mywiki.wooledge.org/BashPitfalls#Non-atomic_writes_w...
e.g. testing in CI a JSON file is properly formatted:
jq . < $f | diff -u $f -I rewrote a Hadoop job to use Streaming (pipes, basic unix model) from a Java job to native. Just that switch was 10x faster. Mostly because of Hadoop overheads.
The problem tends to come along with the visibility that comes when one of these get released publicly. People use these things without doing the same level of due diligence to see if something less complicated fits their needs. Either deliberately (resume driven development) or simply because it doesn't occur to them to do so and "if it's good enough for this unicorn it's good enough for me". Or in some cases they do due diligence, but incorrectly evaluate their needs and rule out simpler solutions.
But the most common reason I've ran into is simply that it's more fun to play with these projects than it is to use crufty old stuff, no matter how tried and tested and potentially appropriate it is.
At a previous company, I had a python script running on a cron job every 5 minutes to do some data processing needs. Once or twice a month, the batch would be so large that it took 6-7 minutes to complete and the cron job would trigger the script again before it finished, causing the second instance of the script to see the lock file, log an error, and exit. It didn't cause any problem for these periodically skipped runs, because the business need only required data to be processed within 24 hours of coming in. The 5 minute cron job was just to even out resource usage throughout the day instead of doing a larger nightly batch job. A piece of data not getting processed for 8 minutes instead of 5 did not make any material impact.
Another team had noticed the errors popping up in the log and were in the process of testing out a whole bunch of real time data pipelines like Kafka, leveraging my error messages to justify the need for a "real time" system without ever even asking me about the errors. After I found out other people were noticing those superfluous errors, I moved the cron job to a 30 minute window to stop them from happening anymore. Turns out they didn't have any other justification for their data pipeline greenfield project and weren't happy to go back to their normal day to day work. I offered to let them maintain and expand my python scripts if they were interested in data processing work, but for some reason they never took me up on that offer. :(
Right, Yahoo made Hadoop because (probably) they needed it. 99% of companies... just don’t.
I think this is a dangerous, irritating mythology that programmers permit to their detriment. Skipping work entirely to go play pool or watch a movie is "fun". Evaluating a new technology stack to see if it fits business needs (present or future) might be _intellectually stimulating_, but deriding it as "fun" - and allowing management to write it off as time-wasting - hurts everybody. This isn't "fun", it's research, just the same as particle physics experiments are, and it's a big part of what we went to college to learn to do effectively.
"This is an absolute classic which although ancient in Computing years is an absolute gem full of relevance yet."
short-descriptions and tons of examples.. I had fun and learned a lot doing this
link: https://github.com/learnbyexample/Command-line-text-processi...
Did anyone try the unicage solution?