1,785 karma · joined June 18, 2016
https://trino.io/
https://www.linkedin.com/in/hashhar
I can be found most places as hashhar.
What I think you are talking about when you say OpenJDK is called Adoptium now (https://adoptopenjdk.net/releases.html). Earlier called Adopt OpenJDK.
Similarly Eclipse folks named their project Temurin and so on.
Solved using domain ownership via TXT record for example.
> Using usernames for namespaces makes typosquatting even worse, because many > usernames are odd and hard to remember correctly (would you remember digits > in winapi's owner handle? Is it BurnSushi, BurntSushi, BurnedSushi?)
Fact is that people don't write their package dependencies by memory, they usually go to the readme of the project and use the instructions there.
The change in limits doesn't persist across restarts anymore like it used to in the past.
Thanks for the correction.
The way to work with it is:
git add file1
git commit -m "Fix some bug in file1"
git add file2
git commit -m "Add a feature flag for file1"
# Oh no, first commit was borked, need to fix it
git add file1
git absorb
# this will now automatically create one or more fixup commits
git rebase -i --autosquash $(git merge-base master)
# voila, tidy history and each commit worksBut yes, it's unfortunate that when the next tokens are joined token and laid out in the form of a sentence it appears "intelligent" to people. However if you instead lay out the individual probabilities of each token instead then it'll be more obvious what ChatGPT/LLMs actually do.
Impractical, yet possible ways to store data is an exciting satire genre for me now.
The rule of thumb is that Trino aims to provide a uniform layer over whatever sits underneath. So operations which when "pushded down" to the source result in same results as when executed within Trino do get "pushed down" - i.e. executed on the source. But in cases where the results might differ or it's complex to push-down the operation the operation runs within Trino - i.e. pull data from source (minimum needed data) and then perform operation within Trino.
Note that it's not an all-or-nothing case, e.g. a query like:
SELECT n.name, r.name
FROM postgresql.tpch.nation n
LEFT JOIN postgresql.tpch.region r ON r.regionkey = n.regionkey
WHERE n.name > 'A'
AND r.name = 'ASIA'
will result in the following query to Postgres: SELECT n.name, r.name FROM tpch.nation n LEFT JOIN tpch.region r ON r.regionkey = n.regionkey WHERE r.name = 'ASIA'
The rest of the query (n.name > 'A') would be applied in Trino to results fetched from Postgres because the collation in Postgres will affect results if we push complete query down to Postgres and may not match results when entire query would be processed by Trino. With single data source this is not easily appreciated but e.g. if the second table in the query came from SQL Server then you'd want to have a consistent comparision logic regardless of source of table.This is a very simplified example though.
It doesn't have to be SQL based systems on the other end - the most used connector with Trino is to query files on object storage (S3/GCS/Azure Blob).
Disclaimer: I'm one of the maintainers of the project.
> - a CLI verification option that will test a given % of your backup objects (at random) every time it's run, to give you opportunistic random verification to try find issues.
FWIW Restic also supports this - see restic check --read-data-subset (https://restic.readthedocs.io/en/stable/045_working_with_rep...)
Also Restic doesn't have any problems with buckets with deletions disallowed as long as you allow them for just the `locks` directory. A policy I used for one of my targets looked like:
{
"Statement": [
{
"Sid": "AllowAdditions",
"Effect": "Allow",
"Action": [
"s3:PutObject",
"s3:GetObject",
"s3:ListBucket",
"s3:GetBucketLocation"
],
"Resource": "arn:aws:s3:::BUCKETNAME"
},
{
"Sid": "AllowDeleteLocks",
"Effect": "Allow",
"Action": "s3:DeleteObject",
"Resource": "arn:aws:s3:::BUCKETNAME/locks/*"
}
]
}
And it even works with buckets with object locking enabled since >= 0.13.It's interesting because the correct answer is that each country/region has different expectations from bottled water, mineral water.
I'm a believer of user-facing testing - i.e. test things like how a user would operate them. So in the context of building a federated query engine our users won't care what version of Postgres they are connecting to - they will care however that the query they wrote doesn't work - doesn't matter what Postgres version.
- write a test that verifies the failure
- add a todo comment in code linking to the test to fix the code when test fails
Then eventually the test will start failing, you can fix the code and the test and profit.
I was hoping to start over with a fresh account and instead ended up having to create a new email alias in Outlook to use to create a new Apple account - luckily I used an alias the first time as well so the impact wasn't as bad.
When your monthly Athena bill crosses whatever it would cost to have 5 or 10 EC2 machines it'll be cheaper to use Trino. At my previous workplace we moved from ~$40,000/month to ~$18,000/month by replacing Athena.
Athena is a very good tool to start with - unless you have super large scale you'll probably not outgrow it. But when you do there's Trino.
I do contribute to Trino - although I was merely a user when that cost reduction happened.
I agree but it depends a bit on what purpose you are using them for. If you mainly use the tool to JOIN some data in bulk and then write output somewhere else (i.e. ETL) - either will serve you fine.
If you write complex queries with multiple filters and want to JOIN across multiple datasets - sure Spark can do that as well but it's not as efficient in pushing down computation to the source.
e.g. A query like SELECT c.custkey, sum(totalprice) FROM orders o INNER JOIN customer c ON o.custkey = c.custkey WHERE o.orderstatus = 'O' GROUP BY c.custkey; when ran on Spark will pull both tables into memory and then perform the join + filter for orderstatus = 'O' and then compute the sum.
While in case of Trino it'll push down the entire query into the remote database (in this case, in other queries it'll push down some parts of the query) so the source database will not need to return gigabytes of data over the network every time the query runs (and hence finish faster as well).
Trino tries to push-down some operations to the remote system which can be done more efficiently there. e.g. filtering on a column that has an index in the remote RDBMS will be faster than pulling all data and then filtering in Trino. Spark doesn't have strong pushdown and has to pull most of the raw data and then apply processing on top of it.
That's one of the main differences. Spark is a distributed job execution framework first while Trino is a distributed federated query engine first and it shows in their strengths and weaknesses.
If you want to run arbitrary user defined transformations on data then Spark definitely has much more to offer than Trino.