ZA̡͊͠͝LGΌ causes "Invalid MD5 checksum on messages"
github.com
github.com
This part gets me every time!
Never fails to make me chuckle
Edit: The function in the fix does in fact appear to be a straightforward implementation based around a regex. https://github.com/mathiasbynens/he/blob/master/src/he.js
Spot the error:
1) HTML is not a regular language
2) Regular expressions can only parse regular languages
3) Regular expressions can not be used to to parse arbitrary HTML documents
4) If you try to use regular expressions to parse HTML documents, you are doing it wrong
The error is between step 3 and 4: It's true that you can't write a regex that parses arbitrary HTML.
However, very often we do not have to deal with arbitrary HTML. Very often we have to deal with documents that only use a subset of HTML and they can be parsed by regular expressions just fine.
This is false, regular expressions as popularly implemented can parse non-regular languages.
For example, /^(.*)\1$/ is a non regular language.
Isn't terminology great?
You refer to a misleading overloading of the term when referring to the stack machines needed to realize 'regular expressions with backreference memories', which expands the language class to which they correspond.
Citation : https://www.amazon.com/Introduction-Automata-Languages-Compu...
I agree, assuming the emitter of the document to be processed is a) correct (as in bug free, and as in knows the expected HTML subset and the corner cases of its ad hoc parsing method, and will never make a mistake nor assume the other end has an actual HTML parser) and b) has no ill intent.
In practice, it is neither.
> people thinking they are smarter than they really are
I genuinely think it is the quippy expression of someone who has been burned way too much by the practical side of this that they prefer to frantically laugh at themselves as much as the issue out of despair of people trying to be smart with regexes. They chose to actually not be smart at all, and just use an HTML parser to parse HTML documents.
In the second case, all you might want is to extract the content of the first <h1> tag out of that error page. That's predictable enough of a task that a Regex might be able to handle it, especially if at that point you've already iven up on a full success and you're just salvaging a prettier error message than "system error".
the reason why is that standard nfas and dfas cannot implement arbitrary counters for the purposes of keeping track of recursion depth or "how many open parens are currently open right now in the parsing?"
there are a few subtle points here:
1) you CAN parse and validate non-regular languages by using regular expressions for subsets of them that ARE regular and then validating the whole. this is exactly what most parser generators generate.
2) you CAN use regular expressions to manipulate text in regular languages if you are careful. however, you should be extremely cautious, as non-regular parts of languages are often those that are used to support escaping or quoting or recursive sections of languages and subtle bugs in these sorts of things can sometimes have unintended consequences in terms of security.
3) you CAN use regular expressions for simple search and replace but see #2. if the language allows escaping or quoting, your regexes will not respect it on their own. this may have the unintended consequence of things you expect to have escaped not be escaped or vice versa. depending on other assertions in your project, this may or may not matter. in the worst case, it could result in a hard to track down or security critical bug.
4) regular expressions are implemented by simulating nfas and dfas. sometimes when implementing these simulations, programmers take advantage of the fact that they're programming and just slap counters in there giving their implementations the ability to support some non-regular languages.
as a rule of thumb, if i'm dealing with untrusted input, i'll use a proper parser.
so yes, you can use regular expressions, but you should understand their limitations if you do so. (which also means understanding your input language)
That really depends on who "we" are and what you mean by"very often".
I used to develop web crawlers, HTML parsers, document analysis infrastructure and various other things that come into contact with "content" for web crawlers at various search engine companies. If you assume people can produce valid, or even half way sane HTML, you'll be disappointed. As for how you parse insane HTML: with difficulty.
A "straightforward regex" is so large/complicated the solution actually uses a JS code generator.
Yes, decoding is only part of XML parsing, and you technically can use a regex for the coding.
But 9/10 times sometimes rolls a regex to help parse XML they fail miserably. Including AWS engineers.
Rarely better than using an existing general XML parser and then validating the additional constraints, especially in terms of implementation and maintenance cost.
The question is about tokenization, not parsing.
const decodeEscapedXML = (str: string) =>
str
.replace(/&/g, "&")
.replace(/'/g, "'")
.replace(/"/g, '"')
.replace(/>/g, ">")
.replace(/</g, "<");
Seems like it's in multiple places in the code base too, I think all those clients are automatically generated.https://github.com/aws/aws-sdk-js-v3/search?p=1&q=const+deco...
So that's not even a good fix, think he should have fixed
aws-sdk-js-v3/codegen/smithy-aws-typescript-codegen/src/main/resources/software/amazon/smithy/aws/typescript/codegen/decodeEscapedXML.ts
So please don‘t, I need these buggy hacks ;)
Hopefully they will pay someone to do that.
At least I can read SQS messages now.
I kind of enjoy working with Rust, and have written a few minor utilities to interact with AWS in it for work. I found and fixed a few bugs in the library for it, Rusoto. I observed along the way that the project is sorely in need of people to spend time on management and maintenance. I could do that, but...
I can't ignore that AWS makes some absurdly huge amount of money for Amazon. I don't begrudge them that, but I get paid pretty well at my day job already. Why do more of that work for the benefit of AWS for no money? They ought to pay me for it.
Hell, they have more than enough money to hire a team of professional experts in every language under the sun to maintain their AWS libs. Especially considering that the net effect of higher-quality AWS libs in more languages will result in more money being spent on AWS services and more lock-in. There's no excuse for having such terrible code in a mainstream language like Javascript.
It's broken code. It's bad code.
Never, ever submit bad code. You're just making even more work for everyone.
Its far worse than doing nothing at all. And if you want to highlight the problem, just raise an issue, point to the problem.
https://stackoverflow.com/questions/1732348/regex-match-open...
SO's sense of humor died around 2014.
The `aws-sdk` v2 API is huge. It and its transitive deps are 64MB. In one TS project that I worked on, adding it added 1.5 seconds of transpilation time.
I thought v3 was supposed to be a modular improvement, but even just the `@aws-sdk/client-s3` library is 24MB.
Then I see bugs like `v3 has a dependency on react-native` and I really wonder how the process works internally to release these things.
(I work around this mostly by using native-image.)
You can't even stream an object to S3. https://github.com/aws/aws-sdk-js-v3/issues/2348
Subsumed within rather than superceded, AFAICT—in fact, I’ve had to use notionally outdated documentation (actually, I think it was actually Stack Overflow answers, but...) for the former to fill gaps in that for the latter, so I’m relatively certain of the relationship.
AWS SDK APIs tend to mix poor design with poor documentation, often being extremely leaky abstractions on top of the HTTP APIs, which themselves aren’t masterpieces of either design or documentation.
Did my time with HL7. One of those cases where a standard didn't serve any real purpose. It gave the appearance that it might work, but then bred a whole industry of HL7 <-> HL7 translators because nobody's HL7 is the same. If you need a translator/hub, everyone might as well have their own protocol.
I also work with X12 which is far more consistent (because if it isn't, you don't get paid).
But the format itself is -- believe it or not -- even worse.
The image it produces can't render Zalgo either https://imgur.com/a/pOUu7m1
I see your point and agree with the general sentiment, but this is no leftpad.
I came very close to commenting on the PR, but since it was pointed out elsewhere that the fix was to the wrong file, I suspected it was going to be closed wontfix anyway
So...your argument is that XML parsing/decoding is useless/too easy and every project should reimplement an XML parser/decoder?
If so, I have 0% trust in your judgement.
Moreover, they're already using a `fast-xml-parser` for doing XML parsing. Presumably it doesn't have an unescape function, so they're taking on a dependency on another XML parser (and keeping the old one!) just to get the one function.
It's kinda dumb that fast-xml-parser doesn't full parse the XML and leave content+attribute values in raw forms.
The docs for fast-xml-parser show how to combine it and he.
Had they done that, they wouldn't have this bug that completely breaks the SQS client for me.