HNHacker News
TopNewBestAskShowJobs

crorella

997 karma · joined April 5, 2013

↑
submissionscomments
crorella··on Tesla removes parking sensors, the results are predictably terrible
Worsening products to improve stakeholder value. But they forgot that the value itself is generated by good products that people actually want to use. Very shortsighted.
crorella··on Ask HN: Best resources to learn more about GPT, Chinchilla, LLaMa?
Thanks!
crorella··on DMCA Takedown of iptv-org/iptv on GitHub
What would you recommend instead?
crorella··on Epstein Documents, Fully Searchable
It seems there are portions of the content which is just a bunch of symbols, at least that's what the results are showing (ie. "±ÆÆßÚ ÔÏ œ ‹± ß±´ Wø™ª ø7ß ÆªøT±7 ¨± 檥1ª™ª ¨Wø¨ ø7ß ÔÎ ±J ß±´Æ \Æ1±Æ ")

If I click the associated PDF document, I can see there is legible text, so I think it might be an issue with the ingestion of the document itself.

crorella··on John Carmack Leaves Meta
each team or project normally has a group where you can ask questions, provide feedback or ask for features.
crorella··on Robot treats 500k plants per hour with 95% less chemicals [video]
This is very interesting, as a Data Eng, I'd love to see what they are doing with the huge volume of data they collect.
crorella··on The Lie That Facebook Sold You
I haven't checked recently, but last time I visited the site there were some ppl that is consistently posting high quality content, but it is like 5-10%
crorella··on The Lie That Facebook Sold You
If you unfollow everyone you don't see ads ever again. You just have to visit your friend's profiles one by one (or whatever group of friends you have created) to see their unaltered feeds with no ads.
crorella··on SEC charges owners of New Jersey Deli with a $100M Valuation
Relevant in the Trump case, as they are being charged for the same thing and they were complaining of political persecution since they were the only one being charged for that. Well, not anymore.

EDIT: Context - https://www.wsj.com/podcasts/the-journal/people-of-the-state...

crorella··on Ask HN: Is Humane Vaporware?
do you have a link or something with more context about them and what they are doing?
crorella··on Project Naptha
hehe, I was thinking the same, then I read your comment.

I wish they give support to Firefox in the future, getting text (and even modifying it!) is something I need to do often.

crorella··on HN is up again
it is back! I wonder what happened.
crorella··on Bitcoin drops below $20k, Ether cracks $1k – what this means
What it means? It means people realized there is no inherent value in a tech the tries to solve already solved problems in am extremely inefficient way.
crorella··on Coinbase is rescinding already-accepted job offers
What else to expect from a company whose entire business rests around a ponzi scheme?

The way of backing up to people who already left their jobs was very unprofessional.

crorella··on YouTube's Database “Procella”
I find lots of commonalities with other systems aimed at processing big datasets. The only thing I find different is their management of the metadata at a centralized place, which can bring lots of network IO savings since you don't have to query the stripe stats (like in ORC) in order to enable stripe skipping.

The risk I see with that is that you have to be extra careful when processing updates to the metastore, as sometimes there are competing processes writing metadata about the same partition and if not done properly you'll end up with a corruption that might cause problems (ie. you skip a whole file because the metadata was incorrect, causing a correctness problem down the road).

crorella··on Ask HN: Why are people in real life so different?
I think overly confrontational people gets isolated sooner or later. There are certain times and places to talk more serious or pungent themes… like parties or sharing a beers or whatever, but not every single time you meet.
crorella··on Tell HN: People underestimate the effect of colleges in making lifelong friends
I am extremely skeptical about this, I think college is better at that. In 17 years in the workforce I've made 0 friends, while I have 7-9 very good friends from college.
crorella··on Open Source SQL Parsers
> I wonder if you could get most of the way there by exposing the workload (across different pipeline stages) to a materialized view recommender

yes! that's something we are trying to do, since we have a way to create signatures for SQL statements and subqueries are just SQL statements then we can get all the "signatures" a query use and compare if other queries are using the same signatures. Then just sort those queries by number of times used and put some other perf metrics like IO/CPU needed to compute it and you get a good starting point.

Microsoft did something similar with Azure, using bipartite graphs, their solution was more advanced as they also baked in constraints like "the materialized view can't be more than X GB in size" but the end result is the same. (https://www.microsoft.com/en-us/research/uploads/prod/2018/0...)

crorella··on Open Source SQL Parsers
Interesting facts about the tool, we created the parser with speed and efficiency in mind, it is able to process ~5M queries in less than an hour, since many of the SQL statements are associated with scheduled pipelines, I came up with a way of normalizing them and generating a signature we use to skip them if they have not changed, that saved a lot of processing when parsing.

We also created some other datasets that tell you how the tables are normally joined and another team created a ML model that now we use in an internal tool that automatically recommends the best join keys when you join 2 or more tables (since we have lots of historical info on that already stored).

We also had to create a very efficient table profiler to evaluate the candidates, because in some cases there are columns that are widely used in equi-where conditions, but their cardinality is very high, making them bad partition columns.

One guy created a parser that actually gets the most common values used to filter each column, I guess we could use that in the future to materialize some views; my original vision of the project was to continue with partial aggregations for common computations. The thing is there are several pieces of code that are pretty much copy/pasted and reused in many pipelines, so why not materialize those and rewrite the SQL of the subsequent pipelines to leverage the materialized version? huge savings there.

We are planning to present it in VLDB or a similar forum, there are some aspects of it that we need to 'clean' if we want to open source it.

Other parts of the system include the candidate evaluation and the module that computes the expected savings for the best candidate selected during evaluation; this system in particular has a lot of specific Presto and Spark logic that might need to get more general if we want to open source it.

crorella··on Open Source SQL Parsers
Yes! you can then put a UI on top of that and answer stuff like "who will get affected by my change on this table?" or.. "there is a bug in my pipeline and now I need to notify downstream consumers"
crorella··on Open Source SQL Parsers
haha, we did just that in my company, but aimed at distributed tables (like HDFS) for their partition and bucketing schemata.
crorella··on Open Source SQL Parsers
We created a custom purpose parser to get some interesting facts about the SQL statements running in our warehouse, a couple of millions of queries per day. With that information, we created a tool that finds the best partition and bucketing schema for each table, based on how the query patterns of downstream pipelines access them.

Since some of the tables are multi-petabyte with thousand of downstream consumers, the savings have been in the millions of USD, because of CPU savings mostly, but also there is the benefit of improved wall time and data arriving much earlier to dashboards and reports.

crorella··on My First 80 Days of VR for Exercise
slightly unrelated, but it seems tripp [1] is also a nice approach to meditation using VR

1. https://www.tripp.com/

crorella··on Logging at Twitter
I work in a similar area, have some questions:

- How do you deal with standardization across the different events you log? How do you handle standardized naming, representations (formatting) and types?

- What about validations? do you validate the payloads in any way? if so, at what layer? Client before logging? In Kafka?

- Any privacy capabilities? How do you make sure the data is not accessed by processes/people that's not supposed to access a datum?

Thanks!

crorella··on NFT misconception: JPEG aren't stored on the Blockchain
probably adding the hash means hundreds of $$ in storage
crorella··on Lesser-known Postgres features
Great article, a lot of useful commands for those that work programming in the DB.
crorella··on NFT misconception: JPEG aren't stored on the Blockchain
exactly, anyone could alter the content the URL points to.
crorella··on The benefits of staying off social media
I unfollowed everyone and every page/group on FB and now I can just visit the profile when I want to see how they are doing, my feed is empty and since no ads are displayed when you visit a specific person's profile it is a very clean experience. 0 addiction.
crorella··on Spending $5k to learn how database indexes work
It amazes me that things so basic and fundamental like understanding the way indexes work are often overlooked or not leveraged
crorella··on Show HN: OtterTune – Automated Database Tuning Service for RDS MySQL/Postgres
Hey! This is really cool! We did the same but for Presto and Spark by following a similar approach (mostly focusing on Partitioning and Bucketing, but the intermediate metadata datasets ended up being leveraged by multiple teams in the company) Looking forward for the VLDB presentation.
← PreviousPage 3 of 6Next →