HNHacker News
TopNewBestAskShowJobs

hashhar

1,785 karma · joined June 18, 2016

Software Engineer working on Trino (formerly Presto SQL) with an interest in everything open-source and distributed.

https://trino.io/

https://www.linkedin.com/in/hashhar

I can be found most places as hashhar.

submissionscomments
hashhar··on OpenTelemetry at Scale: Using Kafka to handle bursty traffic
In Kafka the "queue" is dumb, it doesn't lose messages (it's an append only durable log) nor does it resend anything unless the consumer requests it.
hashhar··on OpenTF is now OpenTofu
OpenJDK is now the open-part of the JDK which multiple groups provide builds and customizations for.

What I think you are talking about when you say OpenJDK is called Adoptium now (https://adoptopenjdk.net/releases.html). Earlier called Adopt OpenJDK.

Similarly Eclipse folks named their project Temurin and so on.

hashhar··on Generative Image Dynamics
Maybe because it could be construed as "self-promotion" especially because you didn't bother to tell that you are the author of those tweets.
hashhar··on Can LLMs learn from a single example?
This was an awesome explainer - thanks a lot. It helps clear up a lot of jargon I keep hearing in very precise ways.
hashhar··on Rust crate rg typosquatting/redirect to ripgrep
> Adding of namespaces only moves the squatting problem from squatting crate > names to squatting namespaces. People like nice namespaces too. What if > someone grabs an official-looking namespace like "aws", and what if that's > a legit project?

Solved using domain ownership via TXT record for example.

> Using usernames for namespaces makes typosquatting even worse, because many > usernames are odd and hard to remember correctly (would you remember digits > in winapi's owner handle? Is it BurnSushi, BurntSushi, BurnedSushi?)

Fact is that people don't write their package dependencies by memory, they usually go to the readme of the project and use the instructions there.

hashhar··on A note to young folks: download the things you love
yes, see IPFS and Magnet links at https://www.reddit.com/r/DataHoarder/comments/i2btuu/utzoo_a...
hashhar··on macOS 13.5 no longer allows setting system wide ulimits
The reasonable solution for that would be to fail calls to select() if more than FD_SETSIZE fds are being held instead of nerfing all applications to do the setrlimit dance - some of which may not be actively maintained or even if they are would take time for fixes to be available and distributed.
hashhar··on macOS 13.5 no longer allows setting system wide ulimits
That doesn't work now.

The change in limits doesn't persist across restarts anymore like it used to in the past.

hashhar··on The Shader Permutation Problem – Part 1: How Did We Get Here? (2021)
Welp, no wonder I had to redo my engineering mathematics course.

Thanks for the correction.

hashhar··on The Shader Permutation Problem – Part 1: How Did We Get Here? (2021)
IIRC with combinations you get factorials (n!) while with permutations you get n^n.
hashhar··on Systemd auto-restarts of units can hide problems from you
You'll find Altered Carbon very very interesting then. Similar concepts being explored there for human "resleeving".

https://en.wikipedia.org/wiki/Altered_Carbon_(TV_series)

hashhar··on Lazygit: Simple terminal UI for Git commands
Boy have I got the thing for you. git absorb - https://github.com/tummychow/git-absorb

The way to work with it is:

    git add file1
    git commit -m "Fix some bug in file1"
    git add file2
    git commit -m "Add a feature flag for file1"
    # Oh no, first commit was borked, need to fix it
    git add file1
    git absorb
    # this will now automatically create one or more fixup commits
    git rebase -i --autosquash $(git merge-base master)
    # voila, tidy history and each commit works
hashhar··on Ask HN: Are people in tech inside an AI echo chamber?
What you're describing is very close to what the thought experiment of Chinese Room (https://en.wikipedia.org/wiki/Chinese_room).

But yes, it's unfortunate that when the next tokens are joined token and laid out in the form of a sentence it appears "intelligent" to people. However if you instead lay out the individual probabilities of each token instead then it'll be more obvious what ChatGPT/LLMs actually do.

hashhar··on Anything can be a message queue if you use it wrongly enough
Wow, this reminds me of https://www.youtube.com/watch?v=JcJSW7Rprio (Harder Drive: Hard drives we didn't want or need - suckerpinch)

Impractical, yet possible ways to store data is an exciting satire genre for me now.

hashhar··on SQL:2023 has been released
This is the initial PR which implemented support for this - https://github.com/trinodb/trino/pull/6550
hashhar··on SQL:2023 has been released
For complex things like MATCH_RECOGNIZE (and CASTs) whose syntax and semantics differ across underlying systems unfortunately the result will be that some data is going to be pulled and be processed in Trino - so it'll be slower than native. If you are only dealing with a single data source (unless it's not an RDBMS, say files on S3 or an API) I'd say Trino is not needed and would slow you down.

The rule of thumb is that Trino aims to provide a uniform layer over whatever sits underneath. So operations which when "pushded down" to the source result in same results as when executed within Trino do get "pushed down" - i.e. executed on the source. But in cases where the results might differ or it's complex to push-down the operation the operation runs within Trino - i.e. pull data from source (minimum needed data) and then perform operation within Trino.

Note that it's not an all-or-nothing case, e.g. a query like:

    SELECT n.name, r.name
    FROM postgresql.tpch.nation n
    LEFT JOIN postgresql.tpch.region r ON r.regionkey = n.regionkey
    WHERE n.name > 'A'
        AND r.name = 'ASIA'
will result in the following query to Postgres:

    SELECT n.name, r.name FROM tpch.nation n LEFT JOIN tpch.region r ON r.regionkey = n.regionkey WHERE r.name = 'ASIA'
The rest of the query (n.name > 'A') would be applied in Trino to results fetched from Postgres because the collation in Postgres will affect results if we push complete query down to Postgres and may not match results when entire query would be processed by Trino. With single data source this is not easily appreciated but e.g. if the second table in the query came from SQL Server then you'd want to have a consistent comparision logic regardless of source of table.

This is a very simplified example though.

hashhar··on SQL:2023 has been released
Exactly. It provides an API using which you can build connectors to whatever systems you want. e.g. Here's a connector for GitHub API https://github.com/nineinchnick/trino-rest/tree/master/trino....

It doesn't have to be SQL based systems on the other end - the most used connector with Trino is to query files on object storage (S3/GCS/Azure Blob).

Disclaimer: I'm one of the maintainers of the project.

hashhar··on SQL:2023 has been released
https://trino.io/docs/current/sql/match-recognize.html
hashhar··on See this page fetch itself, byte by byte, over TLS
Contrasting this with what happens over plain HTTP would be fascinating as well.
hashhar··on Extracting the GameBoy ROM from photographs of the die
Anyone who liked this would probably also find https://www.youtube.com/watch?v=6mOJtrFnawk very interesting.
hashhar··on Moving my PC into my rack in a 2U case
A nice compromise is to buy a tower UPS - the ones which are deeper but not as high - and then place them on a cantilever shelf on the rack. Mine occupies 3U of rack space and I can even place my NAS on the shelf in front of the UPS.
hashhar··on I replaced grub with systemd-boot
For examples of how amazing rEFInd can look - see https://github.com/hashhar/rEFInd-theme/ and https://github.com/EvanPurkhiser/rEFInd-minimal
hashhar··on BorgBackup: Deduplicating archiver with compression and encryption
Some more context.

> - a CLI verification option that will test a given % of your backup objects (at random) every time it's run, to give you opportunistic random verification to try find issues.

FWIW Restic also supports this - see restic check --read-data-subset (https://restic.readthedocs.io/en/stable/045_working_with_rep...)

Also Restic doesn't have any problems with buckets with deletions disallowed as long as you allow them for just the `locks` directory. A policy I used for one of my targets looked like:

    {
      "Statement": [
        {
          "Sid": "AllowAdditions",
          "Effect": "Allow",
          "Action": [
            "s3:PutObject",
            "s3:GetObject",
            "s3:ListBucket",
            "s3:GetBucketLocation"
          ],
          "Resource": "arn:aws:s3:::BUCKETNAME"
        },
        {
          "Sid": "AllowDeleteLocks",
          "Effect": "Allow",
          "Action": "s3:DeleteObject",
          "Resource": "arn:aws:s3:::BUCKETNAME/locks/*"
        }
      ]
    }
And it even works with buckets with object locking enabled since >= 0.13.
hashhar··on 3M to end 'forever chemicals' output
Tom Scott did a video about bottled-water, tap-water and mineral-water regarding this - see https://www.youtube.com/watch?v=wD79NZroV88

It's interesting because the correct answer is that each country/region has different expectations from bottled water, mineral water.

hashhar··on Read this post ‘unless’ you’re not a Ruby developer
My head hurts! Does unless apply to both branches or just one. The if example is easier for sure.
hashhar··on You will never “fix it later”
Good point. But that works only if the version number is the thing you want to test. And also it's possible that newer versions include the bug or some even newer version introduces a regression re-introducing the bug.

I'm a believer of user-facing testing - i.e. test things like how a user would operate them. So in the context of building a federated query engine our users won't care what version of Postgres they are connecting to - they will care however that the query they wrote doesn't work - doesn't matter what Postgres version.

hashhar··on You will never “fix it later”
One other thing which I find very useful is to write tests which will start failing once the <trigger condition> changes. For example if today the database version I'm testing against disallows some query shape I can:

- write a test that verifies the failure

- add a todo comment in code linking to the test to fix the code when test fails

Then eventually the test will start failing, you can fix the code and the test and profit.

hashhar··on Facebook has a hidden tool to delete your phone number, email
I ran into this myself. It's even more frustrating that they don't warn you about it upfront when closing the account (at least they didn't back when I ran into this).

I was hoping to start over with a fresh account and instead ended up having to create a new email alias in Outlook to use to create a new Apple account - luckily I used an alias the first time as well so the impact wasn't as bad.

hashhar··on The SQL query engine Trino (formerly PrestoSQL) recaps a decade of innovation
At a certain scale it does become very expensive. It's easy math.

When your monthly Athena bill crosses whatever it would cost to have 5 or 10 EC2 machines it'll be cheaper to use Trino. At my previous workplace we moved from ~$40,000/month to ~$18,000/month by replacing Athena.

Athena is a very good tool to start with - unless you have super large scale you'll probably not outgrow it. But when you do there's Trino.

I do contribute to Trino - although I was merely a user when that cost reduction happened.

hashhar··on The SQL query engine Trino (formerly PrestoSQL) recaps a decade of innovation
I must disclaim that I contribute to Trino.

I agree but it depends a bit on what purpose you are using them for. If you mainly use the tool to JOIN some data in bulk and then write output somewhere else (i.e. ETL) - either will serve you fine.

If you write complex queries with multiple filters and want to JOIN across multiple datasets - sure Spark can do that as well but it's not as efficient in pushing down computation to the source.

e.g. A query like SELECT c.custkey, sum(totalprice) FROM orders o INNER JOIN customer c ON o.custkey = c.custkey WHERE o.orderstatus = 'O' GROUP BY c.custkey; when ran on Spark will pull both tables into memory and then perform the join + filter for orderstatus = 'O' and then compute the sum.

While in case of Trino it'll push down the entire query into the remote database (in this case, in other queries it'll push down some parts of the query) so the source database will not need to return gigabytes of data over the network every time the query runs (and hence finish faster as well).

Trino tries to push-down some operations to the remote system which can be done more efficiently there. e.g. filtering on a column that has an index in the remote RDBMS will be faster than pulling all data and then filtering in Trino. Spark doesn't have strong pushdown and has to pull most of the raw data and then apply processing on top of it.

That's one of the main differences. Spark is a distributed job execution framework first while Trino is a distributed federated query engine first and it shows in their strengths and weaknesses.

If you want to run arbitrary user defined transformations on data then Spark definitely has much more to offer than Trino.

← PreviousPage 2 of 21Next →