Best Practices Are Not Always the Best
blog.bloomca.me
blog.bloomca.me
I find some programmers are very adamant that there's always a right way and a wrong way to do something and you just have to examine the pros and the cons. Well, what about when it's between a legacy codebase that no one on the team has even spent half a day looking at and an idea that someone read on a blog post that they never bothered to try? You can't honestly begin to identify pros and cons in a situation like that. But this is how our industry is now. So you kind of have to take claims of "best practices" with a grain of salt.
I don't even really engage in discussions on "best practices" anymore. Not because I'm cynical, but because they just lead nowhere. Write some code that runs, show it to me and we'll talk.
IMO, what trips people up is when wisdom says "don't use Oracle products" it's completely accurate, but not that helpful when you inherit a huge mess.
Even if I agree with you (I don't know because I've never been in the situation where that was a thing I was considering and weighing against other options) what I'm arguing here is that if someone asks how to do that then you should either tell them along with the admonishment or not say anything at all. But don't go around being a dick by yelling "haha that's the dumbest thing ever and I'm not going to tell you because of best practices"
https://stackoverflow.com/questions/1732348/regex-match-open...
The reaction on StackOverflow was because they were getting dozens (hundreds) more or less identical copies of the same question. It was the same with "parse email via regexp" but I understand that more recent versions of the RFC make this doable.
I call BS.
"Don't have a continuously running service parse arbitrary HTML with regular expressions" would be a bad thing.
Parsing specific, given, HTML files, with known structure to the dev, with regular expressions (e.g. as part of a one-off scrapping script) is totally, absolutely, fine.
If my HTML file is like:
<html>
<body>
<a href="xxxxx">foo</a><br>
<a href="xxxxx">foo</a><br>
<a href="xxxxx">foo</a><br>
<a href="xxxxx">foo</a><br>
<a href="xxxxx">foo</a><br>
<a href="xxxxx">foo</a><br>
<a href="xxxxx">foo</a><br>
</body>
</html>
I can parse it just fine with regex.And even parsing is the wrong word: I can extract the information I want is a better term. I won't be parsing anything in the AST sense.
Can you determine if it's <html><body>Yes</body></html> Vs <html><body>No</body></html> well sure and even more complex operations also 'work'.
However, if you control what's going on then don't use regular expressions it's wasteful overhead. If you don't control it then you can't tell if it will stay like that, which is the trap. Things that don't change the meaning like swapping the order Attributes which does not change the meaning will often force an update to your regular expression.
PS: It's also why these are called best practices not physical laws.
Not really, if you control what's going on it's a very fast option -- use awk, sed, or your scripting language's regex lib, and get the values you need. End of story.
Nobody cares if you can micro-optimize it with lower level string parsing that doesn't use a regex engine -- and the regex engine might end up being faster that that anyway.
>If you don't control it then you can't tell if it will stay like that, which is the trap.
No, but I already covered that in my original comment. One can do it to extract values for a specific, known html file, I said.
Besides,
(a) a lot of changes can still be caught by regex -- e.g. the order of attributes don't matter if the regex checks for that specific attribute alone (e.g. a href in an a tag).
(b) other changes will also fuck any non-regex, parser-based lib. E.g. if the nesting changes or an element is moved to another section. It's not like you don't have to update scripts written with e.g. BeautifulSoup or some HTML parser when the document changes.
>It's also why these are called best practices not physical laws.
That's up to those who advocate them to understand: and don't hand them out as if they were physical laws "NEVER PARSE HTML WITH REGEX".
I mean you are creating a brittle process that's hard to verify it actually worked, unless you have access to the data some other way in which case it's mostly pointless. Regular expressions are not going to tell you if one of the files just happens to have bad data or any number of things that's going to cause problems.
Because that code works fine for 5 years until it blows up after something upstream changes. From a RoI calculation this might be perfectly fine if you're still around and remember how to fix it - it's a 5 minute fix after all! But if you've left, or if that system became old and crufty and fragile you may have just burned 3 days tracking it down.
If someone is asking this question, it is entirely appropriate to tell them "don't ever do that" - they obviously lack the experience and big picture thinking to know when it is appropriate to deviate due to business reasons. If you leave it at that I agree you're being a dick - whenever shooting someone's idea down you had damn well better have some appropriate alternative options.
I believe the point here is exactly that nothing will change, because the regex in question isn't for a service, it is just for that specific file, right now, today. I've regexed specific HTML files myself too, because even though I am very comfortable with XPath and Beautiful Soup and tree representations in general, regexes are even easier on a static file like that.
For re-usable tooling though, after some time you tend to avoid design patterns that you've personal witnessed break down repeatedly and cause issues.
Well, we are, because to quote coldtea, the person you directly replied to, "'Don't have a continuously running service parse arbitrary HTML with regular expressions' would be a bad thing. Parsing specific, given, HTML files, with known structure to the dev, with regular expressions (e.g. as part of a one-off scrapping script) is totally, absolutely, fine." In context the first sentence clearly means that running the service like that would be a bad thing (hooray English and it's deep ambiguities in double negatives).
I still wouldn't necessarily be too upset about telling a junior dev to be suspicious of REs or avoid them (for instance, parsing HTML with REs typically requires using non-greedy matches, and there are some subtleties around that), but if you know what you're doing on a quick job it's fine. If you're a senior dev and you still haven't figured out what's likely to entrench itself and what really is a one-off job, well, that's your real problem. I haven't been surprised about what gets entrenched in quite a while.
Three or four years later, I was having lunch with the guy I'd done that for when he got a phone call. Turns out they were using that program of mine to print labels monthly now, and one of the completely-safe assumptions I'd made years earlier had bitten them. Fortunately, he knew how to resolve it easily, but I learned a very important lesson that day.
From my perspective, this other guy was using a program outside of its design spec. That can be fine, but such a thing is prone to the problem of safe assumptions suddenly not being safe. You wrote the program for the problem you needed it for, and for which, it worked fine. Unless part of the problem at the time was "we expect to use this program for a couple years", there is no reason to make it overly robust.
Code designed to be used once has a habit of sticking around.
The only thing regex has going for it is you might already know it. But, this is one of those time when learning something is worth it.
PS: XP is doing a depth first search, which can be EXtreamly wasteful.
You are trying to make estimates and decisions for project you know literally nothing about. That is about worst practice of them all.
In case the question was asked by someone lacking the experience - well then it is absolutely ok to learn regexp by trying to parse info few downloaded files of toy scrapped site. Having newbie fight like this just to learn makes no sense.
This is true, but only because you're implying that the reason (failing regexp on HTML input) has already been identified. Maybe the parsing job fails silently and then people down the line (who have no idea that the data was passed around as HTML at some point) wonder why no data is coming in anymore, or why they're seeing incomplete data. That could take weeks to trace back to the parsing problem.
In most cases of best practices however, such a mathematical argument doesn't exist. So, it's more of a battle of (not so) informed technical opinions and experiences, where drawing a clear line is impossible.
Reading list:
You should also say WHY not to do that.
This lets whoever asked decide whether the reasoning applies to their own unique situation. Perhaps they are doing a one-off scrape of a single file with consistent structure like in the other comment.
The asker can also apply the same reasoning to similar situations. They won't come back to ask if they should use regex to parse JSON tomorrow.
And if you can't explain why something is "best practice", maybe it shouldn't be "best practice".
This is counterproductive unless you also say what to do instead.
In all seriousness this is a problem I've run into before as well. I once reported a bug and the lead developer of the project responded saying he didn't think the bug existed. Setting the hostname in the config to the server's IP address was irreversible even if you change it later -- the site would redirect back to the raw IP and outgoing emails would have the raw IP in them. Many other users were commenting about the issue with no response from the maintainer, and I ended up finding a workaround using one of the "developer options" which was really just a one line ruby config change. Then, the author wrote like a one paragraph response saying that you should never set up the software like that and not to use that workaround because it's a "developer option". This bug went unfixed for about 2 years after that. And this is a relatively common and high profile piece of open source software. The whole thing was just bizarre, especially how positive my experiences with other open source projects (some third party Spring packages and Go Dep in particular) have been.
I've run into these "no I'm not telling you because of best practices" people in every single "community"
I’ve always designed away the sharp edges. Because I don’t want the 4:00a wake up call from panicky customers.
Which is why all my devs also did rotations in tech supp, QA/test, build monkey, etc.
People were flat out refusing to answer very straight forward questions because "what are you trying to do", "that's not safe", "you shouldn't have to do that".
Other fun times; trying to clear the dirty bit on an ntfs file system, and probing into details only "library writers" should have to worry about.
Whether you like it or not, I and many others instinctively use grammar, spelling and punctuation as a signal about the mind behind the post. Plus, it makes your post easier to read - posts that aren't just a continuous flow of text with no break are easier to read.
My point is - there is an easy way to give your post a greater chance of influencing the thinking more people.
I quite frequently run into the problem of the answer to something simply being "that's not best practice", "that's insecure" or even "that's user-hostile".
I don't /care/. This is getting installed on a Raspberry Pi and installed inside of a touchscreen coffee table and towed across the country. It doesn't matter that this non-internet-connected device may be vulnerable to other websites accessing its webcam. It doesn't matter that hiding scrollbars and cursors is normally user-hostile.
Go ahead and make the point that it's not best practice, insecure, or user-hostile... But don't make that your entire answer.
> We apply best practices most of the time to save time, but we usually tweak them for our concrete situation. Be sure to spend most of your energy on your core competence instead of bikeshedding over processes.
Sometimes the best practice in question does have value and can be articulated. But in those cases the articulated reason makes a better argument and you don’t hear the phrase “best practice” quite as often. When it is just cargo culting, the only defense is to repeat “best practice” over and over again.
The danger comes in that when you want to supplant the best practice with a new practice, you have to genuinely understand the problem in your context, and be able to articulate a full explanation about why the new practice has an advantage.
But there is definitely cargo culting around best practices - see using Redux. Core features of your app shouldn't be left to just acceptable defaults, but instead you should grapple with those problems and choose the best possible answer you can at the time.
We're experiencing that exact problem - we're introducing for the first time a company wide best practice, which isn't even defined yet and in a constant state of flux, but it's become a shield the strongest evangelists stand behind "why would you want to do that? It's not best practice?" - the "best practice" might not last the month, and it's not clear why that's the "best practice", but it's become a magic seal of approval.
Although, I think it's not simple to foresee long-term implications of our decisions. And, best practices do fill that hole. At the same time, programmers should keep an open mind to use a solution if it's a radically simpler choice.
[1]: https://medium.com/@aboodman/in-march-2011-i-drafted-an-arti...
Not following best practices is more likely to result in inefficiency. But the key word here is "practice".
Practice dancing a lot and you will be able to dance well, in the form of dancing you practiced. There are other forms of dance and thus other forms of practice. Practice for lindy hop is going to differ from practice for tango, or jazz. And the best practice for lindy will result in the best lindy dancing, as it has been developed over time and has the best outcomes.
But if for some reason you need to do lindy hop in, say, a really cramped space, the practice will have to change. Best practice doesn't mean only practice.
They called that the "best practice".
"Best" is in the eyes of whoever's calling it a "best practice".
It's often more useful to look at anti-patterns and avoid those (still with a grain of salt) than it is to go looking for a pre-existing "best" solution to your specific problem.
If everyone is doing it. how can it possible make you the best?