We also built https://techusearch.com which uses Sphinx indexes (think of a SaaS version of Elasticsearch backed by Sphinx).
505 karma · joined May 11, 2013
We also built https://techusearch.com which uses Sphinx indexes (think of a SaaS version of Elasticsearch backed by Sphinx).
I guess the title should be a little more mild - scp isn't going away, or rsync via SSH for that matter.
Have you also done any testing with ssh + gzip?
Also, as you note at the end, the security concerns are not trivial.
It is actually a transcompilation to Bash from a functional language, using Python for the intermediate processing.
Nice effort, Redis is perfectly fine but I believe that the storage layer should be somehow more separated in case someone wants another type of storage, e.g. in-memory SQLite is adequate and already installed in most systems.
> id is an incremental value, I choose 11 as a length
for this primary key but this value is defined
by the number of pages you’ll need to index
This is a bit confusing. You'd generally be better off with INT UNSIGNED as it doubles the range for auto-increment columns.Also the visited field would be better to be represented by a timestamp, choosing the right datatype does matter in large tables.
- You should certainly use Requests http://docs.python-requests.org/en/latest/
- The Story class seems somewhat redundant. You could possibly use collections.namedtuple as a container for properties or simply a dictionary. The print_story method could just be the __str__ special method.
- JSON output would be useful.