HNHacker News
TopNewBestAskShowJobs

cachemiss

278 karma · joined April 15, 2015

submissionscomments
cachemiss··on We're moving on from Firebase
Correct, a common mistake people make is conflating these things. I wrote this several years ago about MongoDB:

One thing that helps is if people stop referring to things as SQL / NoSQL as what ends up happening is various things get conflated.

When talking about stores, it's important to be explicit about a few things:

1. Storage model

2. Distribution model

3. Access model

4. Transaction model

5. Maturity and competence of implementation

What happens is people talk about "SQL" as either an NSM or DSM storage model, over either a single node, or possibly more than that in some of the MPP systems, using SQL as an access model, with linearizable transactions, and a mature competent implementation.

NoSQL when most people refer to it can be any combination of those things, as long as the access model isn't SQL.

I work on database engines, and it's important to decouple these things and be explicit about them when discussing various tradeoffs.

You can do SQL the language over a distributed k/v store (not always a great idea) and other non-tabular / relational models and you can distribute relational engines (though scaling linearizable transactions is difficult and doesn't scale for certain use cases due to physics, but that's unrelated to the relational part of it).

Generally people talk about joins not scaling in some normalized form, but then what they do is just materialize the join into whatever they are using to store things in a denormalized model, which has its own drawbacks.

As to the comment above you, SQL vs NoSQL also doesn't have anything to do with the relative maturity of anything. Some of the newer non-relational engines have some operational issues, but that doesn't really have anything to do with their storage model or access method, it just has to due with the competence of the implementation. MongoDB is difficult operationally not because it's not a relational engine, but because it wasn't well designed.

Just like people put SQL over non-tabular stores, you can build non-tabular / relational engines over relational engines (sharding PostgreSQL etc.). In fact major cloud vendors do just that.

cachemiss··on ClickHouse Cloud is now in Public Beta
Cassandra is not a columnar database, columnar in this sense is about the storage layout. Values for a column are laid next to each other in storage, thus allowing for optimizations in compression and computation, at the expense of reconstructing entire rows. Postgres is a row store, meaning all the columns for a row are stored next to each other in storage, which makes sense if you need all of the values for a row (or the vast majority).
cachemiss··on Comparing ClickHouse to PostgreSQL and TimescaleDB for time-series data
Having used both TSDB and ClickHouse in anger I have some thoughts on this:

They are both fantastic engines, I really like that both have made very specific tradeoffs and can be very clear in what they are good and bad at. Having worked on database engines, I can appreciate the complexity that they are solving.

My most recent use is with ClickHouse, which is great and I think a complete game-changer for the company. However there's a lot of issues (that are being worked on, the core team is great, though there are a few personalities that are a bit frosty to deal with). All of these comments come with love for the system.

1. Joins really need some work, both in the kinds of algorithms (pk aware, merge joins that don't do a full sort etc.), and in query optimizer work to make them better. We have analysts that use our system, and telling them to constantly write subqueries for simple joins is a total PITA. Not having PK aware joins is a massive blocker for higher utilization at our company, which really loves CH otherwise.

2. Some personalities will tell you that not having a query optimizer is a feature, and from an operational standpoint, it is nice to know that a query plan won't change, or try and force the optimizer to do the right thing. However, given #1, making joins performant (we have one huge table with trillions of rows, and a few smaller ones with billions) is really rough.

3. The operations story really needs some work, especially the distribution model. The model of local tables with a distributed table over it is difficult to work with personally. It would be nice to just be able to plug servers in without alot of work, like Scylla, and not have two tables that you have to keep schemas consistent with. THere's also just some odd behavior, like if you insert async into a distributed table, and only have a few shards, it'll only use a thread per shard to move that data over. It would be nice if there wasn't as much to think about.

4. Following #3, there's just too many knobs, maybe if they had a tuning tool or something that would help, but configuring thread pools is difficult to get right. I suspect CH could use a dedicated scheduler like Scylla's, that could dispatch the work, instead of relying on the OS.

5. The storage system relies a lot on the underlying FS and settings on when to fsync etc. I suspect if they had a more dedicated storage engine (controlled by the scheduler above), things could be more reliable. I still don't fully trust data being safe with CH.

6. Deduplication - This is a hard problem, but one that is really difficult to solve in CH. We solve it by having our inserters coordinate so that they always produce identical blocks, using replacing merge trees to catch stragglers (maybe), but it isn't perfect. A suggestion if possible is to try and put the same keys into the same parts, so they'll always get merged out by the replacing merge tree (I understand this is difficult).

The CH team is great, and these will be fixed in time, but these were the problems we ran into with CH.

TSDB was really solid, but we never used it at a scale where it would tip over. Our use case is really aligned with Yandex's so a lot of the functionality they have built is useful to us in a way that TSDB's isn't. (Also, being able to page data to S3 is amazing).

cachemiss··on I/O Is Faster Than CPU – Let’s Partition Resources and Eliminate OS Abstractions [pdf]
Generally these tend to be systems that work with machine generated data, my experience is with sensor data generated by automobiles (automated car efforts).

Naive solutions tend to either summarize the data, store as logs and then run batch processes to index in some form (or leave unindexed and just brute force the computation), or limit the incoming data rate to whatever could be indexed.

These can work for some use cases, but make it very difficult to operationalize these data sources (i.e use them to make real-timeish decisions).

Even human generated data sources (fb / twitter etc.) can generate something close to that data rate.

cachemiss··on How The New York Times Uses Software to Recognize Members of Congress
I've considered something like that, but instead of trying to figure out crimes, it would produce a score for bills.

A corruption score for bills, almost like a facebook for bills "This bill is friends with Exxon". It would figure out who spent the most getting the bill passed, and who they bought off to get it.

Just a simple thing for people to point to when they say things are corrupt. Granted in today's environment, that score would be 100% most of the time, but it would be interesting to have some idea just who bought the bill.

cachemiss··on Three Paths in the Tech Industry: Founder, Executive, or Employee
Honestly, if someone could mitigate the risk for this sort of thing, it's a huge deal.

Personally, I hate living in the tech hubs, I'd love to be able to move back to my hometown, it was a great place to raise kids, didn't have the same crushing traffic / monoculture and it was close to my family. I'm where I'm at due to the jobs, nothing else.

cachemiss··on Three Paths in the Tech Industry: Founder, Executive, or Employee
I am deeply concerned now that my random comment might end up as a VC funded startup.
cachemiss··on Three Paths in the Tech Industry: Founder, Executive, or Employee
Most of them I'd wager. Issue is, and it was for me, is that moving out of an area where I can walk down the street and get another great job to an area where the only great job is the one that this hypothetical company is offering just wasn't worth the risk. I have a wife and kids, and so do most specialized senior engineers.
cachemiss··on Three Paths in the Tech Industry: Founder, Executive, or Employee
It depends on what you are doing. If you are trying to build something like a database kernel, you want people who have done it before, very specialized. There aren't lots of them outside of the tech hotspots.

I don't necessarily mean there aren't smart people outside of the hotspots, but there isn't the specialized skillsets that get built from working at the Microsofts / Googles etc. edit: In the number that you need.

Champaign obviously has lots of smart people (my parents went there), but you may have issues finding specialized senior people there.

cachemiss··on Three Paths in the Tech Industry: Founder, Executive, or Employee
We aren't going to agree on this, around remote employees. I believe there is value to co-located teams, not everyone agrees and that's fine.

The other point is, most startups do not need that talented of engineers, they think they do, but they don't. You don't want to pay bay area salaries if you don't have to, and in most of these areas, you don't have to.

If you are somewhere else but the bay, and you are paying people bay area salaries, but you aren't in SV, so you aren't raising VC like SV companies, I don't think that's a winning move.

cachemiss··on Three Paths in the Tech Industry: Founder, Executive, or Employee
If you are bootstrapping, and you are starting a company in an area that does not need the top tier of engineers (which is most of them, regardless of how they talk about hiring the best), I'd consider starting someplace low cost.

I think college towns in the midwest / rust belt are untapped resources, or areas that used to have strong technically focused companies that moved away. I've personally seen founders bootstrap and launch successful companies (B2B with real revenue, not Uber for cats) in areas with little tech presence.

The ability to not have to shell out huge salaries and equity was a real winner, there's also less distractions. You can get strong engineers (no, not SV / Seattle top end engineers, but people that can throw together a reasonable website) for < 100k in these areas, and they won't bounce around as much. You aren't competing with AmaFaceGoogSoft here, you're competing with HR companies and random consulting houses. There's disadvantages to being outside of the tech bubble, but advantages too. That being said, for an employee, you should at least try and do some time in SV / Seattle etc.

cachemiss··on Why SQL is beating NoSQL, and what this means for the future of data
Caveat, this comment isn't directed at you (I agree with your comment), but rather the points around what you are saying.

One thing that helps is if people stop referring to things as SQL / NoSQL as what ends up happening is various things get conflated.

When talking about stores, it's important to be explicit about a few things:

1. Storage model

2. Distribution model

3. Access model

4. Transaction model

5. Maturity and competence of implementation

What happens is people talk about "SQL" as either an NSM or DSM storage model, over either a single node, or possibly more than that in some of the MPP systems, using SQL as an access model, with linearizable transactions, and a mature competent implementation.

NoSQL when most people refer to it can be any combination of those things, as long as the access model isn't SQL.

I work on database engines, and it's important to decouple these things and be explicit about them when discussing various tradeoffs.

You can do SQL the language over a distributed k/v store (not always a great idea) and other non-tabular / relational models and you can distribute relational engines (though scaling linearizable transactions is difficult and doesn't scale for certain use cases due to physics, but that's unrelated to the relational part of it).

Generally people talk about joins not scaling in some normalized form, but then what they do is just materialize the join into whatever they are using to store things in a denormalized model, which has its own drawbacks.

As to the comment above you, SQL vs NoSQL also doesn't have anything to do with the relative maturity of anything. Some of the newer non-relational engines have some operational issues, but that doesn't really have anything to do with their storage model or access method, it just has to due with the competence of the implementation. MongoDB is difficult operationally not because it's not a relational engine, but because it wasn't well designed.

Just like people put SQL over non-tabular stores, you can build non-tabular / relational engines over relational engines (sharding PostgreSQL etc.). In fact major cloud vendors do just that.

cachemiss··on How Did Marriage Become a Mark of Privilege?
Correct, there's relatively easy ways to de-risk things. My wife and I dated for 3.5 years and lived together for 2 before we got married. We are in a high income bracket, and both of us thought quite a bit about the person we wanted to married, and made sure we had good conflict resolution skills, respect for each other. We agreed on kids, religion , money and sex (make sure you agree on these or you are going to have a bad time).

Chances of our marriage ending are very low according to the statistics. People always view marriage as this binary thing where you are either being completely and utterly scientific about things, or you are just being naive and jumping in. I love/d my wife deeply when we got married, but I was also aware of her flaws and my own. If you don't have any empathy or like living your life for yourself, don't get married (and probably don't date), it won't end well for you.

cachemiss··on Amazon is hiring the most MBAs in tech, and it’s not really close
No, generally you report to a former engineer, even up to a VP level in most cases, in AWS its usually former engineers until you get to s-level.

edit: those people may have ALSO gotten an MBA

cachemiss··on Thanks, but I have accepted another offer
At the higher end of the market (i.e you can get hired at the majors in a senior role or equivalent at a smaller company), I generally see three types of people:

I see people that truly understand their market value, have data backing it up, they are professional (I always tell people how I appreciate the time / offer, and I'm polite to recruiters), demanding about information (I have little tolerance for ambiguity in compensation, define everything, this is how I pay the mortgage and save for retirement / my kids college).

Some managers think that is spoiled (not that you are necessarily saying that), but this is a business transaction, I ask for market and turn it down if it's not that, or if the compensation is ambiguous.

Then I see those that don't know their value and just sort of take whatever because they don't like interviewing / negotiating. I don't see many of those at the high end.

The third group are who you are talking about, total babies. They complain about recruiters, or how much everyone wants them, place crazy demands on companies (I worked with a guy who literally had a rider like Van Halen in his LinkedIn profile, he ended up being a total prima donna who got fired because everyone hated him). Most people in this third group think they are in the first, and they aren't worth recruiting, even if they are strong technically.

Sometimes people believe this means startups / small companies can't compete, but the issue is, startups / small companies really aren't offering any sort of market comp, they don't give enough equity to employees to make the risk worth your while. You don't necessarily need to pay me what AmaGoogFaceSoft do, but you need to give me a significant chunk of equity and give me enough visibility into how that is cut up (preferences etc.) in the company for me to value it, otherwise its worth 0, even if I believe that the company will be successful. If you can't trust me enough with that sort of transparency, then it's not gonna work out. With a huge public company, I know what my stock is worth roughly.

cachemiss··on Amazon's Whole Foods Price Cuts Brought 25% Jump in Shoppers
Amazon as a company believes margins are an opportunity, either for them, or for someone else. They don't believe in raising prices unless forced to by something external. You aren't correct here, they believe in making money on volume and breadth of services / products.
cachemiss··on Ask HN: Should I provide salary history after an accepted offer?
Anecdote incoming, but I've never seen that in 12 years in the industry (asking for an existing slip). And I've worked from everyone from AmaGoogFaceSoft types to HR companies to startups.
cachemiss··on Publishing with Apache Kafka at The New York Times
That's a bit of an oversimplification. Production grade RDBMS systems have far more guard rails, testing and work put in to them than Kafka. It's relatively straight forward to lose data in Kafka, I've done it (its usually control plane bugs, not data plane).
cachemiss··on Ask HN: What is the biggest untapped opportunity for startups?
Unfortunately, my employer and team are directly involved in this space. We may not go this way (due to the effort), but its something we may tackle.

It's not necessarily new computer science, just a clever (if I can be so bold) way to tackle edge computing in the context of a streaming engine.

cachemiss··on Ask HN: What is the biggest untapped opportunity for startups?
I'd also add, there's a component of streaming analytics that isn't solved either.

One of the points I've tried to make at various companies (we've worked at the same one before) is that streaming solutions and batch solutions need to be fused into a single execution engine.

A streaming system on its own (operating on temporal windows) is not nearly as useful as on that can be joined to a storage engine with data at rest. It also needs to be disk based, so windows can be large, which most people do not want to take on. It also needs to be extremely parallel, and efficient.

Thousands of requests a second per server is not even in the right ball-park (which is lots of current execution engines now). Operating at line rate is generally table stakes IMO. The operations on the stream should be parallelized automatically, up to petabytes a day of input. Humans don't have the necessary context to do the partitioning up front, especially with streams that change.

The issue is (and I've tried to come up with designs to address this, though not in practice), is that co-locating the data at rest, with data that is moving through the system is a tricky problem, especially with complicated joins.

They can be the same engine (and should), but traditional database engines tend to have a problem with streaming queries, since they are just repeatedly executing a query against every new record. They are expressible, just not efficient. There is room to innovate in this space, but most people building these engines either solve the parallelism problem naively, or not at all.

There's also the problem of driving this computation to the edge, which is also something I have a solution for in a way that no one is doing, but have not yet met a company willing to take this level of effort on.

All the points you make about the kernel are apt, as are the points about the distribution algorithms. Also, the protocols used aren't nash safe, so at scale most of these systems become an operational juggling act under pressure.

All streaming systems that I know of do not know enough about the underlying data to gracefully rebalance and co-locate, since they all tend to embody the map/reduce paradigm, which is oblivious to underlying data distribution, at least in current practice.

There is available computer science to solve all these issues, I think some of the spatial algorithms out there can also be applied to the streaming space, especially in join evaluation.

cachemiss··on Ask HN: Failed interview, feeling unemployable and depressed – what do I do?
Generally people are looking for a preferred solution, solving the problem, but not in the preferred way is usually not enough.

This isn't good or anything (though, sometimes there's the really obvious super slow way, and its appropriate to ask for something better), its just the way it is.

This is especially pronounced with junior interviewers, I've gotten dinged for missing capitalization on one variable in an otherwise flawless exercise.

Most interviewers at these company, especially those with little empathy or training, are looking to rule you out, and are looking to find a problem with whatever it is you do. Sometimes they tell themselves that its because they want to keep a high bar, but its usually just so they can feel better than someone else, they aren't good at judging problem solving ability, only that you arrived at the solution that they had in mind.

I've had offers from the Big 4, and I've totally bombed interviews with them too. I prepared a lot before hand, and its mostly just luck of the draw, if I get a set of coding questions I've seen before, or is similar to what I've seen, I pass, if I don't, I don't. I'm the same engineer either way.

I wish a lot of people in our industry would quit with the alpha nerd crap, you aren't that important, and you're pissing on people that could help you build your project. I think its egged on by the "A players all the time" mantra at large tech organizations, so when you take insecure nerds, and puff their egos up, this is what you get.

cachemiss··on MongoDB queries don’t always return all matching documents
I actually mean reliable. Its probably different now, but at launch, the defaults were fsync'ing every 30 seconds or so. It would literally just apply the change to an memory mapped buffer and just fsync it once in a while.

They did that so they could look good in benchmarks, and it's why they recommended so strongly that your memory completely fit in RAM or else things would fall apart (pro-tip, any system that recommends that has a poorly designed storage engine).

They also screwed up the consistent side of things as well.

cachemiss··on MongoDB queries don’t always return all matching documents
My general feeling is that MongoDb was designed by people who hadn't designed a database before, and marketed to people who didn't know how to use one.

Its marketing was pretty silly about all the various things it would do, when it didn't even have a reliable storage engine.

Its defaults at launch would consider a write stored when it was buffered for send on the client, which is nuts. There's lots of ways to solve the problems that people use MongoDB for, without all of the issues it brings.

cachemiss··on Spark 2.0 Technical Preview
(Guessing you meant to respond to me)

Excellent, I'm glad that the "big data" world is starting to look at database literature in terms of how it does execution, as there is much to be learned.

Most of these systems are extremely inefficient (looking at you Hadoop), when they don't really have to be. Efficient code generation should be table stakes for any serious processing framework IMO.

cachemiss··on Spark 2.0 Technical Preview
Seems similar to this:

http://www.vldb.org/pvldb/vol4/p539-neumann.pdf

cachemiss··on This Tech Bubble Is Bursting
While it certainly is more efficient to automate those employees away, in my experience they are rarely assigned to new work, they are simply let go. Whether or not that's good or bad morally or for the economy I won't speculate on.

Most companies are actually quite overstaffed for various reasons, freeing up X employees in department Y doesn't mean they slide over to department Z, Z already has more than enough employees usually.

cachemiss··on We are ruthless on code reviews
I always use "we", and I almost always phrase things as a question.

Example: "So we are doing X here, which I think will probably do Y, which could have adverse affects Z, are we sure we want to do this?"

cachemiss··on We are ruthless on code reviews
That's a failure of your team lead. Tough code reviews are important, but I don't tolerate people being petty or nasty. We actually train new employees on how we like to do reviews.

If I see people being given the benefit of the doubt, I call out other senior engineers. If I see people ganging up on a new person, I do the same.

A quote I've always heard, is generally when someone talks about how "brutally honest" they are, they enjoy the brutality more than the honesty. I filter for the latter, and come down hard on the former.

cachemiss··on Heroku Kafka
To clarify, my feelings towards Kafka are from the POV of someone who has had to build a managed service on top of it, which is not the common use case (for which many people seem to be happy with). Other people may have more positive experiences.

In my experience, Kafka is a solid system when you work in its wheelhouse, which is a relatively static set of servers / topics, that you add to slowly and deliberately. If you can't use something like Kinesis, then its a good choice.

In Kafka, programmatic administration is generally an afterthought. They have APIs for doing things, but they generally involve directly modifying znodes. Simple things don't work or have bugs, deleting topics didn't work at all until 0.8.2, and even now has bugs. We've seen cases where if you delete a topic while an ISR is shrinking or expanding, your cluster can get into an unrecoverable state where you have to reboot everything, and even then it doesn't always get fixed. Most of the time you are expected to use scripts to modify everything (there's a wide variety of systems out there that try to build mgmt on top of kafka).

Its dependency on Zookeeper is a pain, and limits scalability of topic / partition counts. Rebalancing topics will reset retention periods because they use the last modified ts of the segment files to check for oldness, meaning if you rebalance often, you need extra disk space laying around. ZK has some bugs with its DNS handling, which affects Kafka if you try and use DNS.

It has throttling, but its by client id, what you'd like in some cases, is to say that a node has X throughput, and have the broker be able to somewhat guarantee that throughput, and create backpressure when clients are overwhelming it. Otherwise your latency can go through the roof. You also want replication to play nice with client requests, and it doesn't (if you add a new broker and move a bunch of partitions to it, you'll light up all your other brokers while it replicates, and cause timeouts).

Its replication story can cause issues when network partitions come into play.

It's highly configurable like many Apache projects, which is a blessing and a curse, as your team has to know all the knobs, both consumer / producer / broker side.

The alternative if you are at a company with the resources to do so (mine is), is to build something that fits your use case better than Kafka, or to use a hosted service like this, or Kinesis.

cachemiss··on Heroku Kafka
Kudos to Heroku. As someone who has had to make Kafka into a managed service, I know what a pain it is (I'm not a Kafka fan for a lot of reasons) to administer in a cloud environment.
Page 1 of 2Next →