556 karma · joined July 20, 2010
In fact, there's even a callout to Result<> "..but the only thing you can't do, is ignore them" . :)
I ‘ve tried all kinds of ideas before settling for this setup, and they all felt forced or just too much trouble for what they were to me. Text files, Dropbox, and Spotlight have been a perfect combination for my needs for years now.
It's not as a big of a deal in Python because as others have mentioned, you usually don't end up writing large functions to begin with, and tab identation is supported. It still means that if you try to send a code snippet to someone over email or Slack, IM, etc it may not work because whitepsace may be trimmed etc. With YAML, it's way worse for reasons the author outlined.
Something may look good on paper (or in screenshots) but practical considerations need to factored in when designing a grammar, and there are myriads pretty-formatting utilities that could be used to that end if one cares for that sort of thing (see also: clang-format).
IR is for all intents and purposes a solved problem -- in fact it was solved a long time ago, and I highly recommend the seminal book “Managing Gigabytes”. I also recommend https://github.com/phaistos-networks/Trinity/wiki/IR-Search-... this page(disclaimer: I am maintaining it) for some interesting/important links to IRC technologies, developments, etc. While some novel ideas come out from time to time, the fundamentals haven’t changed -- progress there is incremental and mostly specific to different encoding schemes or ways to execute queries faster by using JIT or more cache-aware datastructures, etc.
Managing and queries documents based on keywords and boolean operators is one thing, and Lucene/Solr, and Trinity (https://github.com/phaistos-networks/Trinity) among other technologies can be used to take care of those challenges. But that’s the easy part (assuming you can do this fast enough, because you almost always can’t afford long-running queries):
- User Interfaces: Not just how results are presented, but also how users can construct or input queries. What options can be come available for filtering matches? - Ranking: Precision is key, and rather simple formulas (tf/idf, BM25, etc) generally don’t work well for many/most domains. Furthermore, ranking is almost always not just about relevancy. It factors in static context scores (e.g document “popularity”), personalisation biases(how likely is it for user to mean Soccer or American Football for [football]),and other signals, fused together somehow to determine the final ranking of matched documents. - Scale: Getting everything right is one thing, getting everything right at massive scale is whole different game. What may work on small scale(algorithms, technologies, services) may not work at all when you scale out. - Everything else not directly related to search but either important or fundamental to a good experience/business: from matching queries to ads, to analytics, to autosuggestions, to training ML models to power all that, etc.
Web search is not a zero sum game. Bing makes over 3nb / year and while it may not have a chance to catch up with Google anytime soon, that’s a great business right there. Ditto for DDG. There are also companies that offer a different or better experience and access to datasets google doesn’t yet.
So, all told, search may be solved only in terms of the basic IR technology that makes it all work, and arguably a lot better than it used to be in terms of user interfaces, ranking, etc, but it will take a lot longer until those other aspects of web search may be considered ‘solved’.
( I never drink coffee )
2. We didn't need transactions - if the producer would fail while it was executing the REPLACE statements, we 'd start it over and it wouldn't be a problem (idempotency)
3. We didn't care for crash recovery either -- if aything would go wrong, we 'd rebuild those tables (we only cared for 2 weeks or so worth of rows, rebuilding them wouldn't take long).
I think you ignore that, despite MyISAM's deficiencies, it's really fast if you don't care for the aforementioned properties/warranties provided by more modern engines. And it was -- for our dataset, it was almost twice as fast as InnoDB.
We have been running mySQL in production since release 3.x; we moved to it from mSQL. It may not mean much, but we know it mySQL well, at least some of our folks do.
As I said in another reply, we didn't use native table partitioning because we didn't get the expected benefits in a different use-case/dataset, but we certainly should have considered it.
Thank you for the suggestions though :)
We just figured out exactly what we wanted, stored the data in chunks, compressed(we used 2 var-int encoding schemes, and snappy compression), indices(skip-lists) for each file and each chunk and for queries, we parellize access to as files required across multiple OS threads (scatter-gather). In fact, I am sure we could have gotten better performance if we wanted to spend more time on that problem. It wasn't novel or particularly interesting or hard anyway. Just something that needed to be done to help us solve a problem.
Again, there are probably tunables and practices specific to MyRocks that we just didn't consider. We just didn't want to pursue this any further.
Doesn’t mean others will have a similar experience(in terms of performance and scalability) with us. Obviously. YMMV.
InnoDB wasn't working out for us (INSERTs and SELECTs were too slow). MyISAM fared a lot better; INSERTs were a lot faster(obviously?), and SELECTs some 50% faster. But it too wasn't as great as we hoped it'd be. It got to where it was too slow for production use (some operations would take over 4s, which was a deal breaker for us).
We then switched to MyRocks. It was great initially -- much smaller on-disk footprint, INSERTs were fast, SELECTs were fast enough. But two weeks later it also got to where it was too slow. Slower than InnoDB and MyISAM even; also, our mySQL server would often starve for memory, and restarting it was the only practical way to “fix” it.
We would DELETE from those tables every few hours because we were only interested in the last 2 weeks worth of rows, so the dataset size was constrained/bounded, so slow downs weren't a result of tables getting larger.
In the end, we just gave up on MySQL, wrote our own thing that stores and accesses data on-disk directly. Disk footprint is over 2 orders of magnitude smaller, and all operations take constant time(no more than 20ms at 99pc), whereas in the past we ‘d get around 3s at best, and 10 or maybe 20s on average(not even at 99pc).
This is not about us doing anything “better” or about rolling your own alternative to generic datastores. It’s about MyRocks performance’s initially being good, but deteriorating very quickly - and how it compares with InnoDB and MyISAM, at least how it did for us. Also, using an RDBMS wasn't likely the right choice for what we were doing anyway, but for various reasons that's what we used.
It was in beta at the time, and I am sure there are tunables we could change to maybe get a better performance, but we didn’t really bother with any of that.
Nothing about that or the JS engine is impressive really. That was my point, more or less. Of course, there's a difference between building something that works, and something that works exceptionally works etc etc(all the stuff on top of that), but all told, I still don't think building a browser justifies assembling such a huge org, even if there is no reliance on third-party technologies.
Browsers have evolved - because standards did, and expections along with them, but in the past they too were quite complex(see also Netscape's play: browsers, email, usenet client etc, suit) and again, it was harder then that it is nowadays to build such systems.
Whenever a value for such a unique (key, arguments) cell would be required, the application would collect all such that need updates -- or, you could click a button to update them all --, create a message that would contain all paths and arguments, and auth info (each user would require to authenticate with some centralized system, so that the auth key would be used for that RPC, and that the service could check for authentication and authorization ), call out to some service and get values for each of those cells (value could depend on the authenticated user, or could even be denied depending on privileges ).
Furthermore, the application would cache each such cell value (auth key, resource path, arguments digest, value) and it would periodically update them, or do so on demand (so that you can use it say on a plane, and refresh them when you land). That’s the gist of it, but there’s more to it and there are some optimization opportunities I am ignoring here, but it could work.
Then again, for all I know, there are already such services, or apps that kind of do this sort of thing.
What is odd is that is seems to be related to rendering; sometimes when I scroll on either direction on Tweetbot or on Safari, the iPhone will stall for maybe 500ms and the battery will drop by 20% or so instantly. Other times, I will app-switch to Slack and it will do the same. Reading content in black background/white text on the other hand, doesn't seem to be draining the battery. So I think its safe to assume this is an upgrade related problem. Also, I have disabled spotlight indexing and pretty much everything else I could, and it's also been quite some time since the upgrade for any upgrade tasks to be running to completion still.