HNHacker News
TopNewBestAskShowJobs

plamb

163 karma · joined January 27, 2011

submissionscomments
plamb··on Breaking the trillion-rows-per-second barrier with MemSQL
SnappyData employee here. In general this is called the "HTAP" industry (Gartner's phrase: Hybrid Transactional/Analytical Processing).

SnappyData: https://www.snappydata.io, MemSQL: https://www.memsql.com/, Splice Machine: https://www.splicemachine.com/, SAP Hana: https://www.sap.com/products/hana.html, GridGain: https://www.gridgain.com/

are some of the technologies within it

plamb··on Hazelcast 3.8.3
Worked on GemFire prior to Pivotal acquisition (before Geode) and currently work for SnappyData.

You can imagine GemFire/Gridgain as an apples-to-apples comparison. Both are "enterprise" in-memory data grids originally intended for managing data in low-latency OLTP applications which later added analytics/OLAP features. Geode/Ignite are the open source options for these two IMDGs and also a good apples-to-apples comparison. (Hazelcast also has enterprise/OSS verisons I would compare accordingly)

I can't speak to the current comparison between these systems, but I can compare them to SnappyData. SnappyData deeply integrates GemFire with Spark to bring high concurrency, high availability and mutability to Spark applications. In the world of combining Spark with a datastore over a connector (cassandra, hive, mysql, mongo etc) to enable "database-like" features in Spark, SnappyData has taken the next step of integration. In Snappy, the database (GemFire) and the Spark executors share the same block manager and VM so the systems no longer communicate over a "connector." This, along with our database optimizations, provides the best performance for Spark applciations in what I like to call the "Spark Database Ecosystem."

As such, comparing SnappyData to GemFire/Hazelcast/Gridgain does not make much sense unless you are trying to use Spark in conjunction with these systems. In that case, the main difference I would point out is that SnappyData will necessarily perform better as any of them would need to use a connector to interact with Spark. The better comparison would be between SnappyData and Ignite, as Ignite contains a direct Spark abstraction called "IgniteRDD." That said, the majority of the comparisons/benchmarks we've run have been against MemSQL+Spark and Cassandra+Spark, so I don't have much to say about Ignite vs SnappyData.

User manigandham mentions SnappyData's Approximate Query Processing features (called Synopses Data Engine) which is unique within this space, but a discussion of which would take this too far afield.

plamb··on SQL in CockroachDB: Mapping Table Data to Key-Value Storage (2015)
SnappyData employee here -- This is essentially what we did. The main difference is that we already had a decade old transactional K/V store, that, over time morphed into a more full fledged in-memory database. That is what we integrated with Spark versus rolling a new database. The SQL layer in this database (GemFire/Geode) already had a number of optimizations we could use to speed up Spark SQL queries, even over the native Spark cache.

Like some of the other comments in this thread, the idea was to provide all the guarantees of a OLTP store (HA, ACID, Scalability, Mutations etc) with the powerful analytic capabilities of Spark.

plamb··on Joining a billion rows 20x faster than Apache Spark
Appreciate these comments, the site did not go through much testing before being deployed. Overflowing was modified to eliminate horizontal scroll on mobile but it looks like there were some vertical issues as well. We will get this fixed
plamb··on Joining a billion rows 20x faster than Apache Spark
Our impression was that when Databricks released the billion-rows-in-one-second-on-a-laptop benchmark, readers were pretty awed by that result. We wanted to show that when you combine an in-memory database with Spark so it shares the same JVM/block manager, you can squeeze even more performance out of Spark workloads (over and above Spark 's internal columnar storage). Any analytics that require multiple trips to a database will be impacted by this design. E.g. workloads on a Spark + Cassandra analytics cluster will be significantly slower, barring some fundamental changes to Cassandra.
plamb··on Joining a billion rows 20x faster than Apache Spark
Will look into this
plamb··on Joining a billion rows 20x faster than Apache Spark
The font in the embedded gists or the font on the page?
plamb··on Aerospike: Architecture of a Real-Time Operational DBMS (2016) [pdf]
SnappyData: https://github.com/SnappyDataInc/snappydata
plamb··on SnappyData: A Hybrid Transactional Analytical Store Built on Spark
There was a talk sort of on this topic at Spark Summit in June... Here is the part where 2.0 is mentioned: https://youtu.be/PViLT2E2WPQ?t=407
plamb··on SnappyData: A Hybrid Transactional Analytical Store Built on Spark
It takes about as long as it takes to start up a Spark cluster, and you can interact with it entirely through Spark APIs if that's what you're comfortable with. It can also be used in "Split Cluster Mode" so you can use your existing Spark build instead of what's embedded within SnappyData if you prefer.
plamb··on SnappyData: A Hybrid Transactional Analytical Store Built on Spark
Github repo: https://github.com/SnappyDataInc/snappydata
plamb··on Source: Microsoft mulled an $8B bid for Slack, will focus on Skype instead
Microsoft wouldn't let me use the version of skype i had on win 8.1 anymore unless i upgraded it, then wouldn't let me upgrade it unless i upgraded to windows 10. I opted to try the skype web client and had to install some skype plugin. Chrome quite literally stopped working the next day, though I'm not sure if it's related to the skype plugin. At any rate, Slack's user experience has been head and shoulders above Skype.
plamb··on SnappyData: OLTP and OLAP Database Built on Apache Spark
No we were not referring to snappy compression when we came up with the name.
plamb··on SnappyData: OLTP and OLAP Database Built on Apache Spark
Also (correct me if I'm wrong), the stability of SnappyData will depend more on GemFireXD (related to Apache Geode), the in-memory database that has been integrated with Spark to form SnappyData, then it will depend on Spark. GemFire has been in development for over a decade and has a multitude production use cases.
plamb··on SnappyData: OLTP and OLAP Database Built on Apache Spark
Hi hsshah, we talk about some benchmarking we did in our technical paper (page 10, "Experiments") http://www.snappydata.io/snappy-industrial. This is not a real-life scale deployment, but may be useful.
plamb··on SnappyData: OLTP and OLAP Database Built on Apache Spark
I'm an employee for SnappyData, just letting you know I'm going to make sure these questions get answered.
plamb··on Ask HN: Is the noidea app the same?
hmm i wonder if something is wrong with my cache or browser... I'm pretty sure when I click through it's the exact same app. Will post solution if I figure it out.
plamb··on Ask HN: What would you do differently if you could go back to college again?
I would have double majored or at least minored in computer science to go along with my analytic philosophy degree
plamb··on Announce Day - YC W12 Applicants Chat on Wompt
Also having issues; after logging in through google no people appear and I can't type in the chat box; tried twitter and got some sort of socket error after hitting 'sign in'.
plamb··on Start-Up Chile is a Great Experience But Be Careful Too
(startupchile participant) I highly, highly recommend homechile. Fernando, the manager, speaks good English, was a big help in finding an appropriate place and has helped us multiple times with translation when issues like having our internet randomly go out have come up. Providencia is a nice centralized location and near the office. If I could do it over I probably would have tried to pick a place in Manuel Montt (Providencia) as it is walking distance from the office and also near most of the nightlife.
plamb··on Start-Up Chile is a Great Experience But Be Careful Too
As a current member of the program, I would argue that it is better suited for a company in idea-phase, where the founders likely still having a savings account available (for the two month money issue Ken describes).
plamb··on Ask HN: How is the Portland web dev/startup scene?
If you idle on freenode, hanging out in #pdxwebdev and #pdxhackathon is a good way to meet some developers here in Portland.
plamb··on Ask HN: Can anyone recommend some philosophy reading?
The books that made me decide to be a philosophy major:

Zen & The Art of Motorcycle Maintenance

Ishmael

Days of War, Nights of Love: Crimethinc for Beginners

The Way of the Peaceful Warrior

plamb··on I am starting a company in Chile
Curious as to what you mean by successful. Are you saying the 22 other startups declared failure/died? Are you talking about attracting investment in Latin America? I'm guessing a lot of those companies came back to the US to search here.
plamb··on Ask HN: Who got into YCReject?
I don't believe anyone that's applied would know by now if they've made it to the interview. Here are the dates of notification + interview:

Notification of interviews: Thursday, May 12 Interviews: May 23-26 Final Selection: May 26

plamb··on Ask HN: How did you learn distributed computing/systems?
This site is probably the best place to start: http://highscalability.com/

I learned most of what I know from working for http://www.gemstone.com

plamb··on Ask HN: What are you working on?
We are building web iPhone & Android apps that consume and filter the social media content bars/restaurants are pushing. Ultimately we'll be combining this with a bunch of location based services, recommendations and group communications. --> http://www.barbird.com (also free on iPhone/Android)
plamb··on What makes someone a "Co-founder"
I agree except that 'taking an operating role in the business' is a big grey area.

For example, I could see someone meeting (1) & (2) but merely providing some of the graphics for your first prototype. This may be crucial to your first prototype, but it is an extremely replaceable skill-set and it is not crucial to the longevity of the company. This person should not be considered a co-founder of your business. Now say the same person helps out with product development, takes over marketing and demonstrates that he can one day manage people and you've got someone who is less replaceable and more of a co-founder. Therefore, I think you need to make really clear what you mean by 'operating role in the business' before you make decisions about co-founders

plamb··on Show HN: location-based messaging
Clickable: http://crumbsapp.com/

I think it is a cool idea but that its success will come down to how non-geeks adopt it. If you can get that large pool of non-geeks, whose interaction with the web is dictated by Facebook & Google, to feel incentivized to place notes to their friends then I think it could be a hit. There are also potentially interesting opportunities for local business owners to get 'spontaneous business' via notes placed around their business.

How would you describe your differences from a service like http://geoloqi.com/ ?

plamb··on Ask HN: Have you looked back at your YC submission?
It keeps me awake every night. We submitted on Feb 15th... It's amazing how much changes over the course of a month.
Page 1 of 2Next →