HNHacker News
TopNewBestAskShowJobs

Max-Ganz-II

636 karma · joined September 24, 2022

Blog here; https://www.redshift-observatory.ch/slblog/index.html

Redshift investigative white papers here; https://www.redshift-observatory.ch/white_papers/index.html

Redshift replacement system tables here; https://github.com/MaxGanzII/redshift-observatory.ch

submissionscomments
Max-Ganz-II··on Coinbase awarded a $500k bug bounty
I Googled for AWS bug bounty programs.

I found nothing.

Do you have a URL of any kind, for more information about this, including contacts?

Max-Ganz-II··on Coinbase awarded a $500k bug bounty
Thank you everyone who replied.

I feel like I have a better understanding now of the situation and what happened with H1, and I feel better about it.

I will now see if I can figure out how to wipe everyone's RS clusters instantly with a single command, so I can report something on H1 after all :-)

Max-Ganz-II··on Coinbase awarded a $500k bug bounty
Yes and yes :-)
Max-Ganz-II··on Coinbase awarded a $500k bug bounty
Just as an aside, I don't think your query is directly logged.

I've not actually checked, so I don't know, but knowing how logging works on RS, I think the cluster crashing will mean your killer query is not logged.

Your session will have been logged by the time the cluster crashes. OTOH, maybe you were logged in for some time first, or there's connection pooling, or you slipped the query into an existing connection's query stream, and so on.

Actually, thinking about it, I think you could reduce the problem to a single query, rather than two, which would help cover tracks.

Max-Ganz-II··on Coinbase awarded a $500k bug bounty
Good point.

The problem I know of is a bit different, in that it is a direct and immediate server crash. It's not a denial of service by making the cluster slow. It's run-query, crash-server.

You are right of course that any normal user can issue crazy queries which hog resources, and hammer performance.

Max-Ganz-II··on Coinbase awarded a $500k bug bounty
I have to say, I did not have a positive first experience with H1.

Probably mainly my misunderstanding, but H1 did not help in any way.

I opened an account, filed a report - I can easily crash Amazon Redshift as an unprivileged user. Provided the DDL/SQL to do so - dead simple, two statements, issue them and boom.

I received a reply, something like, "we have closed the report, if you can demonstrate a working issue we'll investigate further".

I was confused, replied and asked for explanation. No reply.

I tried going to their Support, 403 - doesn't work via Tor browser - no use for an anonymous report.

And that seems to be it - end of road.

I don't understand, no replies, no support, and I've disclosed valuable information and I have no idea what H1 have done or are doing with it (if it's been made public, for example).

(I asked on HN for advice. One line of reply was that this is not an exploit, but a bug, which I can see. OTOH, when I filled in the severity rating form, there was nothing in that where I was evidently going against the grain of what was expected, so I'm not wholly sure. Any further advice in replies now gratefully received.)

Max-Ganz-II··on Advice sought regarding HackerOne and vulnerability submission
This makes sense and I see it, but there is one matter which gives me pause; when I came to submit the issue, there's some kind of standard severity rating system, which you have to fill in. Questions like - can this issue be exploited over a network, or do you need local access? how many privileges are required? what's the consequence of the issue - degraded service, complete denial of service? questions like this. The issue here fitted into this framework of questions - there was no point where it didn't fit, and the severity rating system was evidently thinking in different terms.
Max-Ganz-II··on Advice sought regarding HackerOne and vulnerability submission
No. It's not a way to get in, but it is a way to crash a system (which should not be crashable) if you are in.
Max-Ganz-II··on Advice sought regarding HackerOne and vulnerability submission
I may be completely wrong, but that seems a bit narrow a definition - attackers move laterally once they penetrate organizations.

I can also imagine disgruntled or malicious employees or contractors.

Max-Ganz-II··on Advice sought regarding HackerOne and vulnerability submission
Being able to create a table and issue a query means being a normal user on the cluster - so you have an account and can log in. It's enough to issue `CREATE USER [blah] PASSWORD [pwd];`, which is the minimum create user command.

So we're looking at malicious or disgruntled employees, and also normal queries where users do not realize what they're doing will kill the cluster.

Max-Ganz-II··on Why is Snapshot Isolation not enough?
Amazon Redshift very recently has made this the default choice.
Max-Ganz-II··on [dead]

  Why, man, they did make love to this employment;
  They are not near my conscience; their defeat
  Does by their own insinuation grow:
  'Tis dangerous when the baser nature comes
  Between the pass and fell incensed points
  Of mighty opposites.
I am sorry to see anyone shot, and here quite possibly die, but I think his party were elected on the back of Russian money; Putin did everything he could to influence the election. To ride corruption to power, against popular will, upon matters of great import, is to run grave risk.
Max-Ganz-II··on Development Notes from xkcd's "Machine"
I love this XKCD.

I have a bookmark for it, and whenever I want to kick back a bit, perfect.

However, I've noticed that it has a tendency to go blank (i.e. fail and stop working) when multiple colour spectrum triangles are in play and in particular when they are on top of each other.

I'd really like to see some additional objects to place into the machine.

One other problem I have is that the "perma-link" button doesn't seem to do anything. When I come back to the URL, my machine isn't there.

Max-Ganz-II··on Mishaps in Redshift Temporary Tables
1. Temp tables are a bad idea, last I knew, because they increase the amount of data in the system tables (every time you create a temp table, you create records for it, it's MVCC, so they don't go away until VACUUMed) and system table bloat has been a serious problem for Redshift, since it seems only to go away on reboot. With a client now, lots of tables, takes 60 seconds to issue a system command. A previous client had to reboot daily. I simply advise clients not to use them.

2. Temp tables do not participate in k-safety, so you avoid that performance cost.

3. I suspect I've seen something odd happening with compression with temp tables. I've not investigated.

4. Using `CREATE TABLE AS` is, for me, verboten. Absolutely forbidden. This is because Redshift selects column encodings, and does a very, very poor job of doing so. Never let Redshift select column encodings (or sort keys, or distribution keys, or rely on auto-vacuum, or auto-analyze, and above all, never use AutoWLM).

Max-Ganz-II··on Stack Overflow and OpenAI are partnering
Regarding usage, I was on SO.

I specialize in Amazon Redshift.

I've written a lot of PDFs about Amazon Redshift - serious stuff, deep technical investigations and explanations, published along with the source code which produces the evidence which the PDF is based on - and when people asked questions where I'd written up the answer, I pointed them at the appropriate PDF.

After some months, I received a direct message, which looked to me to be a pro-forma, a standard message sent in this situation, from the staff that I was promoting my site and I should not do so. It was well written and polite.

That's fine - I have no problems with that, it's their web-site.

What I did not like, however, and what came over as slimey, was that the staff had also deleted every post I had made.

This was not mentioned, at all, in the well written and polite message, which then of course became disingenuous. If you're going to do something serious like that, you need to tell people, not let them discover it for themselves.

This was for all posts, where I'd explained something directly or pointed to a PDF - presumably it's a standard action SO take in this situation.

I deleted my account and left.

Max-Ganz-II··on Ask HN: Freelancer? Seeking freelancer? (May 2024)
SEEKING WORK

Digital nomad, relocate and remote both good.

Technologies : Amazon Redshift

Web-site : https://www.redshiftresearchproject.org

In particular, now offering cluster cost reduction. Fee is and only is one month of the saving made.

Max-Ganz-II··on The Reddits
Rather the same here.

A month or two after my main account was banned (see my other post in this thread), I logged into a second account which I'd not logged into for months.

Upon logging in, I discovered the account was "permanently suspended", and the reason for this was, and I quote;

"Your account has been permanently suspended for ."

Max-Ganz-II··on The Reddits
My account on Reddit, and so as I am the founder also the sub r/AmazonRedshift, on 2023-09-30 looked to have been banned by an automated system.

The sub appeared to be working normally, I posted about the Amazon Redshift Serverless PDF, and then Reddit began behaving oddly.

After some investigation, and some guesswork, I concluded my account had been silently shadow-banned, and the sub banned (and then shortly after, deleted).

(Shadow-banning means when you log in as yourself, you see all your posts, and you see them in the threads where they were made. If you view Reddit when logged out, you then see all your posts have been deleted.)

Two years of posts and the sub disappeared, instantly, abruptly, without warning, reason, appeal process or notification, and Reddit is trying to lead me into thinking my account is still active. Make of that what you will.

Having had that experience, I concluded Reddit is not a safe place to invest time in.

Max-Ganz-II··on Free data transfer out to internet when moving out of AWS
I specialize in Amazon Redshift.

Speaking for and only for Amazon Redshift, as I have little knowledge of other AWS services, I hold AWS's blogs, messaging, Support communications, TAMs, the lot, as relentlessly positive and to my eye deliberately and knowingly obfuscating all weakness. I regard information from AWS regarding Redshift as safe to read when and only when you already know what's going on / the underlying truth. Otherwise you will be misled, and to your cost at AWS's benefit.

By the sounds of it, the messaging over this change in data policy is the same.

Max-Ganz-II··on All 7 planets could fit between Earth and the Moon
This is not quite what it seems.

What happens as you pile mass into a planet is that the planet becomes dense, not large, and this is because of gravity.

Jupiter has more than twice the mass of Saturn, but is only moderately larger in diameter.

You can keep dumping mass into a planet, and it just won't get much bigger, until you have enough mass that fusion kicks off, and then suddenly the now-a-star inflates, because it becomes extremely hot and then you have something the size of the Sun.

Max-Ganz-II··on [dead]
I had a sub on reddit, r/AmazonRedshift, for two years, which about two months ago was deleted, without warning, reason, notification or any mechanism for appeal, I think by an automated system (the owning account was shadow-banned, which is to say, you're not notified of it, and when you browse your own content it all looks normal, but no one else can see anything you ever posted).

I had a second Reddit account I used for non-Redshift stuff, which I've not used now for a couple of months. The two accounts to my knowledge are wholly unconnected.

I logged in just now to have a look. The account has been permanently banned, as of two weeks ago.

The reason, and this is exactly what is written in the automated message, is;

"Your account has been permanently suspended for ."

Max-Ganz-II··on [dead]
The ongoing (six to ten weeks now) issue with disk-read-write performance has been resolved now in all regions.
Max-Ganz-II··on Redshift Research Project: Amazon Redshift Serverless [pdf]
Redshift Serverless is not serverless. A workgroup is a normal, ordinary Redshift cluster. All workgroups are initially created as a 16 node cluster with 8 slices per node, which is the default 128 RPU workgroup, and then elastic resized to the size specified by the user. This is why the original RPU range is 32 to 512 in units of 8 and the default is 128 RPU; the default is the mid-point of a 4x elastic resize range, and a single node, the smallest possible change in cluster size, is 8 RPU/slices. 1 RPU is 1 slice. With elastic rather than classic resize, the greater the change in either direction from the size of the original cluster, the more inefficiency is introduced into the cluster. As a cluster becomes increasingly larger, it becomes increasingly computationally inefficient (the largest workgroup has 128 normal data slices, but 384 of the lesser compute slices), and increasingly under-utilizes disk parallelism. As a cluster becomes increasingly smaller, it becomes increasingly computationally inefficient (each data slice must process multiple slices' worth of data), and incurs increasingly more disk use overhead with tables. The more recently introduced smaller workgroups, 8 to 24 RPU (inclusive both ends) use a 4 slice node and have two nodes for every 8 RPU. In this case, the 8 RPU workgroup is initially a 16 node cluster with 8 slices per node, which is resized to a 2 node cluster with 4 slices per node - a staggering 16x elastic resize; the largest resize permitted to normal users is 4x; an 8 RPU workgroup, with small tables, uses 256mb per column rather than the 16mb per column of a native two node cluster. Workgroups have a fixed number of RPU and require a resize to change this; workgroups do not dynamically auto-scale RPUs. I was unable to prove it, because AutoWLM is in the way, but I am categorically of the view that the claims made for Serverless for dynamic auto-scaling are made on the basis of the well-known and long-established mechanisms of AutoWLM and Concurrency Scaling Clusters ("CSC"). Finally, it is possible to confidently extrapolate from the `ra3.4xlarge` and `ra3.16xlarge` node types a price as they would be in a provisioned cluster for the 8 slice node type, of 6.52 USD per hour. Both Provisioned and Serverless clusters charge per node-second, but Serverless goes to zero cost with zero use. With the default Serverless workgroup of 128 RPU/16 nodes (avoiding the need to account for the inefficiencies introduced by elastic resize), one hour of constant use (avoiding the need to account for the Serverless minimum query charge of 60 seconds of run-time), without CSC (avoiding the question of how AutoWLM will behave), costs 46.08 USD. A Provisioned cluster composed of the same nodes costs 104.32 USD; about twice as much. Here we have to take into consideration the inefficiencies introduced by elastic resize, which become more severe the more the cluster deviates from 16 NRPU, that Serverless uses AutoWLM, with all its drawbacks, and which is a black box controlling the use of CSC, with each CSC being billed at the price of the workgroup, and the 60 second minimum charge. All Serverless costs (including the charge for Redshift Spectrum S3 access) have been rolled into a single AWS charge for Serverless, so it is not possible to know what is costing money. It would have been much better if AWS had simply introduced zero-zero billing on Provisioned clusters. This would avoid many of the weaknesses and drawbacks of Serverless Redshift, which wholly unnecessarily devalue the Serverless product, as well as avoiding the duplicity, considerable added complexity, end-user confusion, cost in developer time and induced cluster inefficiency involved in the pretence that Serverless is serverless.

https://www.redshiftresearchproject.org/white_papers/downloa...

https://www.redshiftresearchproject.org/white_papers/downloa...

Max-Ganz-II··on Redshift Research Project: Amazon Redshift Serverless [pdf]
NOTE AND CORRECTION 2023-10-04

I misunderstood how Serverless billing works. Having read the documentation, I understood - incorrectly - that pricing was per-query. The docs talk about pricing being per RPU-hour, but billed on a per-second basis, where if a query runs for less than 60 seconds, it is billed for 60 seconds, and this is for a serverless product. Therefore pricing is per-query - as it is with Athena, and with Lambda.

In fact, it is not.

Pricing is per workgroup-second. I still have not found this stated in the docs; I was pointed to a re:Invent talk where an AWS developer presented a slide which made this clear.

When I was working on the investigation, I was always running a single query at a time, so billing looked right.

This change fundamentally changes the pricing proposition offered by Serverless; the original pricing conclusion was incorrect by an order of magnitude.

I have now rewritten the content regarding billing and have republished. The abstract below is the new abstract. The revision history explains what happened. Credits will credit the reader who pointed the issue out, once they let me know if they want a credit or not.

Max-Ganz-II··on Redshift Research Project: Amazon Redshift Serverless [pdf]
In fact I believe my account has been deleted, but I've not been informed. I can still log in and view my account, and my posts in the threads they were made in - but if I view Reddit when I'm logged out, all the posts I've ever made have been deleted.

If you have been using that sub, please instead keep an eye on the RRP blog, or use the forums on the RRP site, which are here;

https://www.redshiftresearchproject.org/slforum/index.html

Max-Ganz-II··on Redshift Research Project: Amazon Redshift Serverless [pdf]
An aside : posting about this PDF to r/AmazonRedshift, a sub I founded about two years ago, caused what looks like an automated system to ban the sub.

No other information is given, other than a ban has occurred, no links or information to routes to appeal, or find out what happened, or why.

https://www.redshiftresearchproject.org/slblog/2023-09.html#...

(I see now the sub has disappeared from my profile, too. Two years of posts, gone - instantly, no warning, no reason, no information, no notification and no appeal process of any kind, so far as I can see. Reddit appears to be a risky platform to invest time into.)

Max-Ganz-II··on AWS us-east-1 down
I kicked off a Redshift cluster in every region, they've all run and completed, except for `us-east-1`, which is stuck creating the cluster. Been about an hour now.
Max-Ganz-II··on Two new Amazon Redshift versions out (system table diffs)
Funny though the function is "member_of_role". You're supposed to be granting roles :-) but in RS roles are not really roles, as they are in Postgres, you still have normal groups and normal users, as they were before. Roles in RS are really enhanced groups.
Max-Ganz-II··on Two new Amazon Redshift versions out (system table diffs)
Yes. The function links do not yet work. I improve the diffs every now and then - I added the list of function diffs a month or so ago, but getting the links to work is a second round of work, because right now functions are linked to in the URL by their OID, but the OID varies between Redshift version, so I need to link via their name and arguments.

The dense paragraph of explanation at the of the the page, at its very end, says the function links are not yet working. I was thinking to convert the list to a few bullet points, to improve clarity.

What I will probably add next though is permissions - who has access to the object (view, table, function). Mainly they're rdsdb only, but actually knowing is useful (and it lets you spot any slip ups).

Max-Ganz-II··on The part of Postgres we hate the most: Multi-version concurrency control
In what way? I didn't see anything obviously improper when I learned how serialization isolation worked.
← PreviousPage 2 of 5Next →