176 karma · joined September 8, 2010
Larridin is the measurement layer for AI in the enterprise. We’re building AI-native systems that correlate AI behavior to business results, expose absorption bottlenecks, and enable enterprises to invest with conviction. Backed my top tier investors such as a16z, Google and Bloomberg and with early traction!
We want high-agency software builders (3–10 yrs) who think AI-first, move fast, and want ownership of large parts of product and business. This is an in-office role in the Bay Area for engineers who want to own outcomes, not just tickets.
DM' me on Linkedin to apply: https://www.linkedin.com/in/ameyakanitkar/ This is also primarily backend focussed role.
Join our team and help redefine the future of workplace productivity with AI. We are 16z backed company, early stage!
Join our team and help redefine the future of workplace productivity with AI. We are 16z backed company, early stage!
Its interesting to see new era of browser wars playing out with Chrome firmly in the lead, but Microsoft attacking with Edge where ChatGPT on bing only available on edge browser and no where else. Would be interesting to see how Google would respond.
Both these companies are being challenged with a new crop of Databricks and the confluents of the world.
Although social media in itself is vital - we are seeing more usage of it correlates with negative behavior pattern (as is excessive tv watching).
Should govt intervene?
They do intervene for substance addiction...
Note execution environment for such jobs is Azkaban executor server itself, so you have to take care of resource management (eg. one job taking all RAM on the machine will affect other jobs running on the same machine)
The the UX side it can improve - but it's rock solid in what it does! Well done!
+1 Twitter should build this!
1. How does it scale horizontally? if I am consuming from say topic with 100 partitions - do I have to deploy this on library on multiple nodes? or expectation is we need to embed this inside spark streaming/ storm/ samza/ any other real time processing framework?
2. Is RocksDB instance local to your consumer/ thread/ node? or can state be stored across nodes? What happens when joins are needed across the partitions?
I have been doing real time processing for a long time now - and this does capture most of the typical things you end up doing with kafka with spark/ storm etc- so this does look a significant step forward.
I am just trying to understand where it exactly fits.
If you want to build twitter today - its definitely going to cost more. It has built a brand - and hundreds of millions of loyal fans.
Question is - is it worth the valuation thats put on the company. Given now they are cash flow positive - that they have lost $2bn in the past is immaterial.
Part of the appeal is exclusivity (besides of course its hard to scale).
I'd rather see YC at 200-300 companies a year - and ensuring these companies continue to grow/ succeed after the batch is over rather than churning out 1000 more.
Log / Stream Processing - Kafka \n Scalable Storage - HDFS \n Data Processing - Spark, Map reduce (in that order) \n Historical Analytics - Hive/ Spark SQL \n Real Time Processing - Spark Streaming, Storm \n NoSQL - Cassandra/ HBase \n NoSQL (In memory) - Redis \n Search - Elastic Search \n
Some more honorable mentions: kibana on elastic search - for analytics visualization \n druid - for analytics \n
Above are the basics - if you add them you will have 90% of the standard stack for big data.
its ok - if only small portion of tests are valid at this point but if they show progress with data - this validates they are on the right track.
Uber tinkering with price too much has not gone down well with uber drivers in bay area either.
Can you also elaborate, how you read HFiles and serve it out from Terrapin servers? Are you using similar functionality as HBase? (With block cache like design if yes how do you keep both in sync).
Your blog is missing this interesting detail.
Very typical use case for recommendation systems etc. We face similar problems with latencies on HBase (At Groupon).
So this solution seems interesting. Would be good to have comparison of other solutions Pinterest tried before building this. eg. loading data into Cassandra instead of HBase etc.
In nutshell - very specific use case - but the one which comes across very often