https://github.com/Svensson-Lab/pro-hormone-predictor/blob/m...
https://github.com/Svensson-Lab/pro-hormone-predictor/blob/m...
https://bioperl.org/articles/How_Perl_saved_human_genome.htm...
(And of course all the articles about excel renaming genes, and then people trying to clean up the mess…)
I see we’re being fast and loose with the term “artificial intelligence”?
Natural language -> fully working one, I don't mean some email validators but way more complex stuff. Although, I've recently had a case which was too much even for regexes in any form or spec, then sort of grammar-based parser needed to be done from scratch.
— Thomas Szasz
I think that's how we still use it.
https://www.rand.org/content/dam/rand/pubs/research_memorand...
See the Eukaryotic Linear Motif resource: http://elm.eu.org/elms.
Regex to me, is pattern finding and abstraction taken to the extreme. I like the challenge
/
(?<=^|[ \t])
(?<currency_prefix_with_space>
(?<currency_prefix>
€|EUR|\$|USD
)
[ \t]?
)?
(?<number>
(?<integral>
-?
\d{1,3}
(?:[\.,]\d{3}|\d*)
)
[\.,]
(?<fraction>
\d{2,3}
)
)
(?<currency_postfix_with_space>
[ \t]?
(?<postfix_or_ending>
(?<currency_postfix>
€|EUR|\$|USD
)
) | (?<ending>
[ \t]|$|\n
)
)/xI would've probably approached it differently, trying to first get the 'inverted' match (i.e. ignore anything that isn't a currency-like pattern) and refine from there. A bit like this one I did a while back, to parse garbled strings that may occur after OCR [0]. I imagine the approach does not translate fully, because it's pattern extraction rather than validation.
The idea of an exclusionary approach sounds interesting as well. I'll have to think about that a bit.
Now I am enlightened.
Instead of
[ie]*-?[A-Z]+
it looks like they wanted [ie]?-?[A-Z]The harder ones I have dealt with are those looking for malformed syntax where the closing mark that might be missing could be several thousand characters after the opening mark, or the opening mark itself might be missing, across a data set that is several hundred million characters. So you need something very complex to find all the distinctive characteristics of the content that is supposed to be enclosed - while avoiding the many similar structures that give false positives. Sometimes the technically easier solution is too slow to run (look ahead and look behind, etc), so you need to pivot and use other regex features. It can take a day or two to get right.
https://stackoverflow.com/questions/1732348/regex-match-open...
Perfect
We don’t need better language tools. Better parsers can, and already have, been implemented in libraries.
I’d have a look at parser combinators.
Is true though, or perhaps "state"? I know I had to come up with an algorithm because regexp alone couldn't do what I wanted (not even advanced features like lookahead, lookbehind, etc.)
Libraries like python's lrparsing [0] let you assign regex's (aka tokens) to variables, then build grammars by combining them using python expressions. For example:
while_statement = Keyword('while') + '(' + expression + ')' + block
statement = while_statement | ....
block = Token('{') + Repeat(statement + ';') + Token('}')
These grammars are more complex, with more rules you have to follow, but also more checking is done when they are compiled. They tend to mostly work once they do compile. So they are what you asked for, but there is no free lunch.On the down side, lrparsing is pure python so it's slower than python's inbuilt regex's.
Python will literally assign the content of your docstring (a triple-quoted string at the start of a function) to a special double underscored ("dunder") property called `__doc__` that any code can access; not just docs generators, but your own code as well (pretty dang useful for generating on-the-fly help output!).
Which also means that if you run into weird behaviour where a function doesn't do what you think it should do, you can just fire up the REPL, import that function, type `print(function_name.__doc__)` and presto, you have the documentation right there, specifically for the exact version you're using. You don't even need to leave your IDE if it comes with an integrated terminal.
edit: not saying it can't be done but our org isn't using auto docs
(lots of fun dunder functions in Python that make metaprogramming a lot easier, but someone needs to tell you about)
Or even easier: `help(function_name)`!
- The python docs are there to give you the information you need.
- Passion pages are there to do a deep dive into all the crazy shit you can do with that information =D
Any string. Triple-quoted is just multiline strings and works anywhere.
It's not just a good trick: it's the right way to do it, and tooling can be (trivially) made to show you only what you need to see until you need to see more.
A great way to do it is to split them up by concatenating them across a bunch of lines, and put a brief explanation at the end of each non-obvious part (to the right, on the same line).
Plus that also lets you indent within nested parentheses, making it that much more understandable.
I'm baffled when I come across a file like this where the code itself is heavily commented, but a gnarly regex is not. Regexes are not strings, they are code -- and with their syntax, they need comments even more.
Many CoffeeScript constructs were adopted into JavaScript and TypeScript, but unfortunately, heregexes weren't among them.
Later languages have added support for the `x` flag, including C# and Rust. There is also a stage 1 proposal for JavaScript: https://github.com/tc39/proposal-regexp-x-mode
I probably use regexes the most in notepad++
Second, not at all. An LLM can tell you how the regex works (hopefully). It can't tell you what each piece means in terms of the program's logic. Or at least not always and not reliably.