I implemented a format-agnostic search that can match patterns across various naming conventions like camelCase, snake_case, PascalCase, kebab-case. If needed, I'll integrate in space-separated words.
I've just published the tool to PyPI, so you can easily install it using pip (`pip install super-grep`), and then you just run it from the command line with `super-grep`. You can let me know if you think there's a smarter name for it.
If you do, email a link to hn@ycombinator.com and we'll put it in the second-chance pool (https://news.ycombinator.com/pool, explained at https://news.ycombinator.com/item?id=26998308), so it will get a random placement on HN's front page.
So if I'm trying to locate the error message "because the disk is full" but it's in the code as:
... + " because the " +
"disk is full")
then it will fail.So really, combining both our use cases, what would be great is to simply search for a given case-insensitive alphanumeric string in files that skips all non-alphanumeric characters.
So if I search for:
Foobar2
it would match all of: FooBar2
foo_bar[2]
"Foo " + \
("bar 2")
foo.bar.2
And then in the search results, even if you get some accidental hits, you can be happy knowing that you didn't miss anything.But the string split thing you mentioned happens a lot when searching for OpenStack error messages in Python that is often split across lines like you showed. My current solution is to randomly shift what I'm searching for, or try pick the most unique line.
(`first\S?name` is usually better, by ignoring whitespace -> better ignores comments describing a thing, but `.` is easier to remember and type so I usually just do that)
lets say someone would make a plugin for their favorite IDE for this kind of search. How would the details look like?
To keep it simple, lets assume we just do the super-case-insensitivity, without the other regex condition. Lets say the user searches for "first_name" and wants to find "FirstName".
one simple solution would be to have a convention where a word starts or ends, e.g. with " ". So the user would enter "first name" into the plugin's search field. The plugin turns it into "/first[-_]?name/i" and gives this regexp to the normal search of the IDE.
another simple solution would be to ignore all word boundaries. So when the user enters "first name", the regexp would become "/f[-_]?i[-_]?r[-_]?s[-_]?t[-_]?n[-_]?a[-_]?m[-_]?e[-_]?/i". Then the search would not only be super-case-insensitive, but super-duper-case-insensitive. I guess the biggest downside would be, that this could get very slow.
I think implementing a plugin like this would be trivial for most IDEs, that support plugins.
Am I missing something?
> So the user would enter "first name" into the plugin's search field.
Why wouldn't the user just enter "first_name" or "firstName" or something like that? I'm thinking about situations like, you're looking at backend code that's snake_cased, but you also want it to catch frontend code that's camelCased. So when you search for "first_name" you automagically also match "firstName" (and "FirstName" and "first-name" and so on). I wouldn't personally introduce some convention that adds spaces into the mix, I'd simply convert anything that looks snake/kebab/pascal/camel-cased into a regex that matches all 4 forms.
Could even be as stupid as converting "first_name" or "firstName", or "FirstName" etc into "first_name|firstname|first-name", no character classes needed. That catches pretty much every naming convention right? (assuming it's searched for with case insensitivity)
Ya. Query tokenizer would emit "first" and "name" for both. That'd be neat.
If you're going that far, and you're in a context which probably has a parser for the underlying language ready at hand, you might as well just convert all tokens to a common format and do the same with the queries. So searches for foo-bar find strings like FooBar because they both normalize to foo_bar.
Then you can index by more than just line number. For instance you might find "foo" and "bar" even when "foo = 6" shows up in a file called "bar.py" or when they show up on separate lines but still in the same function.
* my understanding was simply that the regex would (A) recognize `[a-z][A-Z]` and inject optional _'s and -'s between... and (B) notice mid-word hyphens or underscores and switch them to search for both.
So you's search for "/first\_name/i".
/first[-_]?name/i
Or to use your example, just checking for underscores and not also dashes: /first_?name/i
Backslash is already used to change special characters like "?" from these meanings into just "use this character without interpreting it" (or the reverse, in some dialects).Basically in vim to substitute text you'd usually do something with :substitute (or :s), like:
:%s/textToSubstitute/replacementText/g
...and have to add a pattern for each differently-cased version of the text.
With the :Subvert command (or :S) you can do all three at once, while maintaining the casing for each replacement. So this:
textToSubstitute
TextToSubstitute
texttosubstitute
:%S/textToSubstitute/replacementText/g
...results in:
replacementText
ReplacementText
replacementtext
[1] https://www.gnu.org/software/emacs/manual/html_node/emacs/Re...
:S/textToFind
matching all of textToFind TextToFind texttofind TEXTTOFIND
But not TeXttOfFiND.
Golly!
In my setup, `/foo` will match `FoO` and so on, but `/Foo` will only match `Foo`
The other killer feature of nimgrep is that instead of regex, you can use PEG grammar [1]
[0] - https://nim-lang.github.io/Nim/nimgrep.html
[1] - https://nim-lang.org/docs/pegs.htmlOne minor inconvenience is that the scoring should ideally be different per filetype. For instance, Python would count "foo-bar" as two symbols ("foo minus bar") whereas Lisp would count it was one symbol, and that should ideally result in different scores when searching for "foobar" in both. Similarly, foo(bar) should ideally have a lower different score than "foo_bar" for symbol search even though the keywords are separated by the same number of characters.
I think this can be accomodated by keeping a per-language list of symbols and associated "penalties", which can be used to calculate "how far" keywords are from each other in the search results weighted by language semantics :)
Improving the IDE to find one or the other by searching for one or the other is missing the point or the article, that consistency is important.
I'd rather have a simple IDE and a good codebase than the opposite. In the example that I gave the worst thing is that it's the framework which forces you do use these two names for the same thing.
I didn't miss the point, I disagreed with the point because I think it's a tool problem, not a code problem. I agree with most other points in the article.