1,592 karma · joined November 12, 2010
[0] https://npmjs.org/doc/registry.html#Can-I-run-my-own-private...
It would be cool to see some stats on throughput and latency relative to # of producers / consumers and amount of data currently accumulated in the producers (since there is no middleman).
http://wurfl.sourceforge.net/help_doc.php
Mind their (new) license tho if you plan to use them specifically.
Seems like a great fit so the user does not have to perform their own cross-referencing.
What gives?
Testing is a good point. If you're running Pig through their APIs it is definitely easier to test than command-line running scripts. We've written test code that reads and runs pig scripts through the API using fixed sample data (stored in HDFS for easy access), read the results, and compare it to expected results (also stored in HDFS). Honestly, you don't need too much input data to prove the correctness of the query.
Remember that Pig also supports placeholders in your scripts so you that you can set them in run-time to define input/output paths, etc. This makes testing easier.
Dependencies can also be stored in HDFS which makes it simple to run your scripts w/o the need to distribute jars around.
www.wearelistening.org
What i always tho is the ability to run search queries that also involve dynamic grouping (like grouping by random combinations of facets) and providing those aggregated results.
Only thing i've seen that can do this "on the fly" is SenseiDB. CloudSearch/Solr/etc seem to need preprocessing to get this right.
Now if someone could put SenseiDB on the cloud, i'd pay for it...
Here are some that come to mind right now that are very useful:
- Be smart about your commit strategy if you're indexing a lot of documents (commitWithin is great). Use batches too.
- Many times, i've seen Solr index documents faster than the database could create them (considering joins, denormalizing, etc). Cache these somewhere so you don't have to recreate the ones that haven't changed.
- Set up and use the Solr caches properly. Think about what you want to warm and when. Take advantage of the Filter Queries and their cache! It will improve performance quite a bit.
- Don't store what you don't need for search. I personally only use Solr to return IDs of the data. I can usually pull that up easily in batch from the DB / KV store. Beats having to reindex data that was just for show anyway...
- Solr (Lucene really) is memory greedy and picky about the GC type. Make sure that you're sorted out in that respect and you'll enjoy good stability and consistent speed.
- Shards are useful for large datasets, but test first. Some query features aren't available in a sharded environment (YMMV).
- Solr is improving quickly and v4 should include some nice cloud functionality (zookeeper ftw).
Do I get any benefit from you guys regarding this?
1. Listen to bad reviews. 2. Ignore good reviews.
Also remember that the very nature of JavaScript might make the classic idea of how an ORM is implemented a little different.
From my experience, performance is a tough question to answer because you really have to think of the system as a whole. The weakest link (slowest piece) usually dictates the overall system perceived performance. Having a slow DB isn't going to make your choice of Java or Node.js matter that much...
So assuming that everything else is equal (like architecture, data structure choices, database driver quality, database design and queries, etc) it might be that Java is faster.. but it might be a moot point.
Writing asynchronous code in Node.js is much more convenient than Java and helps in doing more at the same time so that you aren't waiting for all those database queries to run one after the other, for example. Check out the Node.js async module for some cool patterns to use (https://github.com/caolan/async).
In all, I wouldn't make the choice to move to Node.js because it has better performance or not, but the overall holistic view.
First to get it out of the way: Consultant != Contractor. Some people think they mean the same thing. They don't. Contractors vs employees is just a hiring logistic. Consultants are usually contractors, but the scope of work is usually more specific and specialized.
I'd define a consultant as a subject matter expert who helps clients with a problem domain for a fixed interval.
"Being a consultant" to me means usually helping a company with a problem or role that they can't help themselves in. That means you have to be not only good at what you do, but confident in your decisions and suggestions. You are the authority in their point of view - or at least close to one so you'd probably want to find that area of expertise where you feel that level of confidence.
Having said that, reading books might be a good start, but that's definitely not enough to become a highly paid consultant who gets great referal gigs. "Been there, done that" is a good expression to describe where your experience would best come from. Obviously, you'll get better as you consult more, but it would be wise not to improve your skills on the behalf of others checkbooks in a consulting role.
So assuming you're ready to go that route, you need to fill your calender with enough work to keep yourself busy (paid) and your clients happy. I personally don't have a secret formula for that, i just try to do the best work i can. Leaving a good impression tends to make it easy for people to recommend you to others which brings you more opportunities.
I'd start simply by working for someone else (i.e. a consulting firm). They'll make some dollars off your work, but you'll start to get a feel of the process and build a client base.
I hope that is at least a bit helpful.
I did see Chrome installed as well so that might be their solution :)
The general consensus on this is to structure your data so that it encapsulates your business needs in one document structure (which is atomic on changes), but i find it hard to always conform to in the real world.
So now i have to use zookeeper (memcached also works) to setup global locks on those specific batch update actions. I guess it's a small price to pay right? Right?
It might not be exactly what you were looking for, but it could have the same outcome.
Tablets and smartphones will be dying slowly too in a few years to be replaced by some other idea.
The point is that you're keeping your original data around is the key part.