HNHacker News
TopNewBestAskShowJobs

markpapadakis

556 karma · joined July 20, 2010

@markpapadakis | Bytes conjurer; Seeking Knowledge 24x7. CTO at Phaistos Networks. Simple is Beautiful
submissionscomments
markpapadakis··on Hashicorp Vault v1.0
Apologies for the blatant promotion attempt but you may want to check out our simple (in pretty much every sense) KMS implementation. https://github.com/phaistos-networks/KMS It’s inspired by vault and Google KMS and scales horizontally.
markpapadakis··on A Not-Called Function Can Cause a 5X Slowdown
Great read but what’s up with the 6 banner ads interspersed in the content page ? Maybe a few too many ?
markpapadakis··on Architecture of Nautilus, the new Dropbox search engine
I am very much looking forward to forthcoming posts describing the actual architecture and specifics -- this is a great high-level overview, but I hope and expect they will expand on this expose soon.
markpapadakis··on Ask HN: What open source project, in your opinion, has the highest code quality?
I study codebases as a hobby. I highly recomend Seastar, Folly, Aeron and Disruptor, SQLite, PostgresSQL, LMDB, Tensorflow, Hashicorp’s vault, and the Linux Kernel projects as prime examples of high quality codebases.
markpapadakis··on Falling in love with Rust
It was mostly sarcasm really — I thought the word play was fun:) kidding aside, I am intrigued by Rust. I am glad it gets the attention that it deserves.
markpapadakis··on Falling in love with Rust
Watch the Apple - Think Different commercial at https://www.youtube.com/watch?v=cFEarBzelBs , and pay attention to the script/words. It's as if it was purpose-written for Rust.

In fact, there's even a callout to Result<> "..but the only thing you can't do, is ignore them" . :)

markpapadakis··on Ask HN: How to organize personal knowledge?
I use Dropbox to store everything. There are no files that I care fore that are not managed by Dropbox (or iCloud, for my photos/videos). I use iA Writer to create lists and documents for everything in Dropbox folders. I also have a folder for PDFs/papers (also on Dropbox). There are also folders for my study nodes ( I study codebases ), notes on personal stuff, notes on Programming, on pretty much everything. I use Spotlight (on macOS) to instantly locate what I need among the 100s of such files.

I ‘ve tried all kinds of ideas before settling for this setup, and they all felt forced or just too much trouble for what they were to me. Text files, Dropbox, and Spotlight have been a perfect combination for my needs for years now.

markpapadakis··on Shamir's Secret Sharing
Vault(https://www.vaultproject.io/) and Phaistos KMS (https://github.com/phaistos-networks/KMS) both use SSD for sealing/unsealing, where a master key is created, 'divided' into multiple keys and a minimum number of such keys are required to unseal the service.
markpapadakis··on YAML: probably not so great after all (2017)
I personally dislike with a passion every language or grammar that depends on white-space identation, especially if the designers were extremely opinioned to the degree that you can only use spaces (or even, a specific number of spaces per identation level) and not tabs.

It's not as a big of a deal in Python because as others have mentioned, you usually don't end up writing large functions to begin with, and tab identation is supported. It still means that if you try to send a code snippet to someone over email or Slack, IM, etc it may not work because whitepsace may be trimmed etc. With YAML, it's way worse for reasons the author outlined.

Something may look good on paper (or in screenshots) but practical considerations need to factored in when designing a grammar, and there are myriads pretty-formatting utilities that could be used to that end if one cares for that sort of thing (see also: clang-format).

markpapadakis··on Ask HN: Is 'search' a solved problem?
Search means a lot of things, but even if we limit to mean web-search, as most people understand it, there is a lot more to it than the actual technology that matches queries to documents.

IR is for all intents and purposes a solved problem -- in fact it was solved a long time ago, and I highly recommend the seminal book “Managing Gigabytes”. I also recommend https://github.com/phaistos-networks/Trinity/wiki/IR-Search-... this page(disclaimer: I am maintaining it) for some interesting/important links to IRC technologies, developments, etc. While some novel ideas come out from time to time, the fundamentals haven’t changed -- progress there is incremental and mostly specific to different encoding schemes or ways to execute queries faster by using JIT or more cache-aware datastructures, etc.

Managing and queries documents based on keywords and boolean operators is one thing, and Lucene/Solr, and Trinity (https://github.com/phaistos-networks/Trinity) among other technologies can be used to take care of those challenges. But that’s the easy part (assuming you can do this fast enough, because you almost always can’t afford long-running queries):

- User Interfaces: Not just how results are presented, but also how users can construct or input queries. What options can be come available for filtering matches? - Ranking: Precision is key, and rather simple formulas (tf/idf, BM25, etc) generally don’t work well for many/most domains. Furthermore, ranking is almost always not just about relevancy. It factors in static context scores (e.g document “popularity”), personalisation biases(how likely is it for user to mean Soccer or American Football for [football]),and other signals, fused together somehow to determine the final ranking of matched documents. - Scale: Getting everything right is one thing, getting everything right at massive scale is whole different game. What may work on small scale(algorithms, technologies, services) may not work at all when you scale out. - Everything else not directly related to search but either important or fundamental to a good experience/business: from matching queries to ads, to analytics, to autosuggestions, to training ML models to power all that, etc.

Web search is not a zero sum game. Bing makes over 3nb / year and while it may not have a chance to catch up with Google anytime soon, that’s a great business right there. Ditto for DDG. There are also companies that offer a different or better experience and access to datasets google doesn’t yet.

So, all told, search may be solved only in terms of the basic IR technology that makes it all work, and arguably a lot better than it used to be in terms of user interfaces, ranking, etc, but it will take a lot longer until those other aspects of web search may be considered ‘solved’.

markpapadakis··on Ask HN: What, if anything, do you listen to while coding?
I prefer silence or isolation most of the times. When I can’t get that or I am in the mood for music, I usually listen to random video games music, or tavern music in games ( on YouTube ).
markpapadakis··on Ask HN: Has anyone switched from being a night owl to being a morning person?
I used to stay up as long as I can. I appreciate darkness and quiet, and this was naturally when I could get that. Nowadays though I sleep at around 10am, and I wake up at 06am. It’s almost as quiet, and almost as dark and I get to read, and shower and maybe have breakfast before rest of the family is awake at around 07:30. This turned out to work better for me all things considered. I am also 42 now so this may be in part because I am not that young anymore and the shifted priorities, habits, and needs are better served by this new schedule.

( I never drink coffee )

markpapadakis··on Nobody's just reading your code
Studying codebases is one of my rewarding hobbies. I ‘ve learned more from that practice( https://medium.com/@markpapadakis/interesting-codebases-159f... ), than from most technical books I ‘ve read (by the way, most technical books aren’t particularly great).
markpapadakis··on N.Y. landlord ordered to pay $6.7M for destroying graffiti
Looks like the owner didn’t do this gracefully and didn’t care much about the significance of the Art to its creators and others who appreciated it likely way more than he did. That’s not very nice. On the other hand, if it is indeed his property he should have the right to do what he wants with it. Unless there was a contract of some kind, not a lawyer, I don’t understand how the judge could have sided with the artists who themselves were on a private property and ended up making claims to said property. Aren’t judges supposed to be impartial and make decisions based on the letter of the law ? Though I suppose the law in the us may protect the artists in this case.
markpapadakis··on MyRocks – A RocksDB storage engine with MySQL
Re: myISAM qualities: 1. Writes are infrequent; every 30 minutes or so we insert/update rows. Requests rate is very flow. This is for a specific report -- and said report was infrequently requested. Only updated by a single producer/process. All that means that locking wasn't a concern for us.

2. We didn't need transactions - if the producer would fail while it was executing the REPLACE statements, we 'd start it over and it wouldn't be a problem (idempotency)

3. We didn't care for crash recovery either -- if aything would go wrong, we 'd rebuild those tables (we only cared for 2 weeks or so worth of rows, rebuilding them wouldn't take long).

I think you ignore that, despite MyISAM's deficiencies, it's really fast if you don't care for the aforementioned properties/warranties provided by more modern engines. And it was -- for our dataset, it was almost twice as fast as InnoDB.

We have been running mySQL in production since release 3.x; we moved to it from mSQL. It may not mean much, but we know it mySQL well, at least some of our folks do.

As I said in another reply, we didn't use native table partitioning because we didn't get the expected benefits in a different use-case/dataset, but we certainly should have considered it.

Thank you for the suggestions though :)

markpapadakis··on MyRocks – A RocksDB storage engine with MySQL
Thank you -- we didn't use table partitioning because we didn't have good results in the past, though we should probably have tried it, and maybe it'd have worked well. We just kind of gave up on mySQL use for that specific problem after failing (again, maybe its entirely because of our inability to use mySQL "properly") to get good results and moved on to something else.
markpapadakis··on MyRocks – A RocksDB storage engine with MySQL
It's not that incredible; for most use-cases there is often an optimal(or "better") way to implement a solution that's specifically designed and implemented to support it, as opposed to relying on systems and designs that are broadly applicable/useful (e.g RDBMS).

We just figured out exactly what we wanted, stored the data in chunks, compressed(we used 2 var-int encoding schemes, and snappy compression), indices(skip-lists) for each file and each chunk and for queries, we parellize access to as files required across multiple OS threads (scatter-gather). In fact, I am sure we could have gotten better performance if we wanted to spend more time on that problem. It wasn't novel or particularly interesting or hard anyway. Just something that needed to be done to help us solve a problem.

markpapadakis··on MyRocks – A RocksDB storage engine with MySQL
Yes, we 'd insert new rows every 30 minutes or so. Also, we did try partitioning one of the 3 tables (which made sense) into multiple tables -- it didn't improve the situation at all.

Again, there are probably tunables and practices specific to MyRocks that we just didn't consider. We just didn't want to pursue this any further.

markpapadakis··on MyRocks – A RocksDB storage engine with MySQL
I suspected that folks would think this is about rolling your own thing as opposed to relying on existing solutions, and I thought I was clear what I said was not about that — I just thought it would be somewhat valuable to someone how MyRocks compared to InnoDB and MyISAM for our use case. That’s just one datapoint.

Doesn’t mean others will have a similar experience(in terms of performance and scalability) with us. Obviously. YMMV.

markpapadakis··on MyRocks – A RocksDB storage engine with MySQL
FWIW, we used MyRocks (went from InnoDB to MyISAM, to MyRocks) for some mySQL tables holding a few dozen million rows a few months ago.

InnoDB wasn't working out for us (INSERTs and SELECTs were too slow). MyISAM fared a lot better; INSERTs were a lot faster(obviously?), and SELECTs some 50% faster. But it too wasn't as great as we hoped it'd be. It got to where it was too slow for production use (some operations would take over 4s, which was a deal breaker for us).

We then switched to MyRocks. It was great initially -- much smaller on-disk footprint, INSERTs were fast, SELECTs were fast enough. But two weeks later it also got to where it was too slow. Slower than InnoDB and MyISAM even; also, our mySQL server would often starve for memory, and restarting it was the only practical way to “fix” it.

We would DELETE from those tables every few hours because we were only interested in the last 2 weeks worth of rows, so the dataset size was constrained/bounded, so slow downs weren't a result of tables getting larger.

In the end, we just gave up on MySQL, wrote our own thing that stores and accesses data on-disk directly. Disk footprint is over 2 orders of magnitude smaller, and all operations take constant time(no more than 20ms at 99pc), whereas in the past we ‘d get around 3s at best, and 10 or maybe 20s on average(not even at 99pc).

This is not about us doing anything “better” or about rolling your own alternative to generic datastores. It’s about MyRocks performance’s initially being good, but deteriorating very quickly - and how it compares with InnoDB and MyISAM, at least how it did for us. Also, using an RDBMS wasn't likely the right choice for what we were doing anyway, but for various reasons that's what we used.

It was in beta at the time, and I am sure there are tunables we could change to maybe get a better performance, but we didn’t really bother with any of that.

markpapadakis··on Maintaining an Independent Browser Is Expensive
The parser. Once for properly parsing HTML content, for a search engine, and the other for a framework used for saving pages (For an Instapaper like service). So it wasn't a big deal, but in both cases, a DOM was constructed and operations were executed against the DOM.

Nothing about that or the JS engine is impressive really. That was my point, more or less. Of course, there's a difference between building something that works, and something that works exceptionally works etc etc(all the stuff on top of that), but all told, I still don't think building a browser justifies assembling such a huge org, even if there is no reliance on third-party technologies.

markpapadakis··on Maintaining an Independent Browser Is Expensive
I never worked on a browser, though I did implement the HTML5 spec, twice, and I wrote a JS engine(compiler, runtime). I know that it was a lot easier because I could count on infromation and other advances that were available to me, and anyone else really, at the time, the kind of resources and information I probably wouildn't have had access to it("whats javascript?") 20+ years earlier.

Browsers have evolved - because standards did, and expections along with them, but in the past they too were quite complex(see also Netscape's play: browsers, email, usenet client etc, suit) and again, it was harder then that it is nowadays to build such systems.

markpapadakis··on Maintaining an Independent Browser Is Expensive
On the flip-side though, the tools, expertise/skills and means for building browsers some 25+ years ago don't match what's available today.
markpapadakis··on Maintaining an Independent Browser Is Expensive
I think this is just too many people. No matter how you spin this it’s just too many. Do Google or Microsoft have half as many working on their browsers? I’d be very surprised if they do. It should also be interesting to look back and find out how many people Netscape employed at its peak. I would love to know the distribution of those 1.2k people ( engineers, marketing, whatever ). Of course, it’s likely that I am too naive or clueless.
markpapadakis··on Stop Using Excel, Finance Chiefs Tell Staffs
I was toying with this idea in the past, where cells would support, in addition to primitives(numbers, strings, whatever) and functions, a new type which would be identified by a reverse domain name notation path/key (e.g org.hr.employees.cnt), and optionally a comma separated (key-value) pairs as arguments.

Whenever a value for such a unique (key, arguments) cell would be required, the application would collect all such that need updates -- or, you could click a button to update them all --, create a message that would contain all paths and arguments, and auth info (each user would require to authenticate with some centralized system, so that the auth key would be used for that RPC, and that the service could check for authentication and authorization ), call out to some service and get values for each of those cells (value could depend on the authenticated user, or could even be denied depending on privileges ).

Furthermore, the application would cache each such cell value (auth key, resource path, arguments digest, value) and it would periodically update them, or do so on demand (so that you can use it say on a plane, and refresh them when you land). That’s the gist of it, but there’s more to it and there are some optimization opportunities I am ignoring here, but it could work.

Then again, for all I know, there are already such services, or apps that kind of do this sort of thing.

markpapadakis··on Ask HN: Why is iOS 11 such a mess?
I am not concerned with the replacement costs -- it's just where I live, getting the battery replacement will likely mean that I have to ship it to another city(Athens), or, maybe, find some place in a different city in the island (Crete) who may be able to do it. Its a hassle and an inconvenience for me.

What is odd is that is seems to be related to rendering; sometimes when I scroll on either direction on Tweetbot or on Safari, the iPhone will stall for maybe 500ms and the battery will drop by 20% or so instantly. Other times, I will app-switch to Slack and it will do the same. Reading content in black background/white text on the other hand, doesn't seem to be draining the battery. So I think its safe to assume this is an upgrade related problem. Also, I have disabled spotlight indexing and pretty much everything else I could, and it's also been quite some time since the upgrade for any upgrade tasks to be running to completion still.

markpapadakis··on Ask HN: Why is iOS 11 such a mess?
I tried that to no avail. Reset and install, no change whatsoever
markpapadakis··on Ask HN: Why is iOS 11 such a mess?
Half a day, while far from ideal, was long enough for my needs. However, as soon as I upgraded, and not a minute earlier, the battery problem manifested. Now, it could be a coincidence and it is indeed just a problem specific to the battery, but its at best suspicious - the timing is. I will replace the batter anyway, and I am getting a new iPhone next week, I just don’t appreciate being “forced” to upgrade -- conveniently when the iPhone X came out, no less.
markpapadakis··on Ask HN: Why is iOS 11 such a mess?
I use the iPhone 6s. It was responsive, and battery would last at least half the day just fine. Right after the upgrade, batter would last for maybe 1 hour, tops, 30 minutes on average. No minor upgrade fixed the problem, and apparently its a widespread one too. I am getting either the 8+ or the X, but I still think this is ridiculous at best.
markpapadakis··on Amazon now has a billion-dollar ad business
what if that advertising revenue goes towards offering even lower prices to customers? would that be an acceptable tradeoff? say, out of 1$ generated from ads on amazon.com, 80c set aside for offering lower prices, and 20c for running costs etc? I don't know if that's the case, just curious if it'd be OK if that was the case.
← PreviousPage 2 of 5Next →