RegEx Crossword
jimbly.github.io
jimbly.github.io
This might be fine for hard mode, but as someone who considers themselves a regexpert it's not very approachable as a first puzzle IMO.
A more gradual introduction to the format would be to give a few clues that give you confidence on specific characters, that then let you lock in some other characters in other hints, and so on.
For instance, replacing `.*H.*V.*G.*` with `.{3}H.*V.*G.*` would go a long way because you could confidently place an `H`. And say that intersected with `(DI|NS|TH|OM)*` on the `H`, you could then place a `T` from the second clue because of what you learned from the first clue.
It could just be that I'm missing something or not as good at regex as I thought, and please let me know if that's the case. Either way though, when I'm trying a new kind of puzzle I'd like to feel like I made some sort of progress after trying for 5-10 minutes, and here 2 chars does not feel like progress.
But since this is my first introduction to this kind of puzzle, I need some anchor points at the beginning so I feel like I have something to work off of.
I'm not even asking for a whole row, just an easier set of known chars at the start of the round so I have a hint at which of the 39 constraints I should start with.
To be honest I'm not even saying this puzzle should change so much as I am looking for a different puzzle to dip my toes.
I don't have any interest in starting the puzzle if I don't feel like I can put a foot down somewhere. It's like trying your first Minesweeper game, making two random clicks, and getting two `7`s. Where do you go from there? Or learning Sudoku from the hardest difficulty level, without having built up a library of patterns from the easier difficulties.
(XHH|[^XH])*
is equivalent to just .*
EDIT: It looks like I read the regex wrong. I guess I need to do the puzzle after all!It is "If any of X or H appear they appear in a sequence exactly matching XHH".
[^C]*[^R]*
LOOKS unsimplified, but it actually means that if the string contains an R, the preceding characters CANNOT be C.Doesn't it really just mean a string can't contain a C followed by an R?
I agree it definitely feels imposing when you first look at it, but stick with it. Look for spaces that have a very small set of possibilities, and then try to map out what neighboring spaces could have for each possibility. If you really feel stuck, take a screenshot and mark it up.
I had the same experience when I tried it yesterday. There doesn't seem to be any good "starting point", like there is on a traditional crossword puzzle or sudoku.
I think from a game design perspective it is really interesting though: Like sudoku, this kind of puzzle gives you a wide range of options to archive different player experiences and difficulty levels: You can make easy levels by mostly using constant-width regexes and non-conditional letters and you can slowly increase the difficulty level by making regexes less constrained and more ambiguous. A designer could even craft specific "paths" through the puzzle by combining easy and hard regexes.
Finally, a designer could gradually introduce more complex regex features (or other patterns, like multiple constraints) over successive leves.
(I guess I could PR on github when I get some time later today)
https://web.mit.edu/puzzle/www/2013/coinheist.com/rubik/a_re...
The ones on https://regexcrossword.com/ are nice but this is the original, for me at least!
https://play.google.com/store/apps/details?id=de.chagemann.r...
I solved this in about 1.5 hours by starting at the top left, entering any string that satisfied at least one condition, then moving on to the next condition and “fixing up” any previous entries. I was fearful that I might arrive at a nearly correct solution that I would have to massively backtrack from, but it didn’t happen - I only needed a few short backtracks. I think the large number of constraints helps a lot.
https://news.ycombinator.com/item?id=26439598
for proper credit to the original author (and place of publication).
Thanks for posting the answers. I think I'm done working on it for now, but that one was really bothering me.
Thanks though. I tried again on my laptop and got it - was a bit easier with a larger display than on my phone.
I think I saw at least 3 places where that technique could be used, probably more.
There isn't a technical reason why the traditional crossword format can't have clues in the shape of regex right?
As far as I can tell, the puzzle is fully constrained.
From the start, the bottom-left cell as well as the right column can be solved just from the starting hints, and then you can derive the top-left cell per the backreference.
In particular, you can reframe the right column's hint as (AB|OO|OR), and only one of those satisfies the bottom row's hint.
Something as simple as "OO\nDO" should be valid. But also "OO\nDD" and myriad other solutions.
edit: and by full match (since this seems to be the source of some confusion in another subthread), I explicitly mean anchored with your \A...\Z or whatever you want to use.
They just went the extra mile to make an special puzzle.
(and on top of that, we're doing a full match.)
- the pattern is found in the input
- the start of the match is the start of the input
- the end of the match is the end of the input
By this definition, 8 spaces does not match the pattern.
R*D*M* does not specify anything that has to be found in the string. Nor does any pattern of ()* or []* no matter what you put between the parens or brackets.
In all of those cases, any possible string matches the regex starting at the beginning of the string and ending at the end of the string.
Your clarification doesn't clarify anything for me.
That is literally what is happening: https://github.com/Jimbly/regex-crossword/blob/master/crossw...
Plugging the first space into a DFA described by the regex is an immediate failure - there is no exit from the initial state initiated by the space character. It's a non match.
Regex engines will say they do match because they are by default checking for for substrings that match the language (such as the initial empty string of each line of grep), not for strings that match the language.
*edit: added last paragraph.
Some regex notations include "^" and "$", and some don't. A lot of software (the grep command, for example) uses the kind that does support "^" and "$". This puzzle uses the other kind.
Essentially, when a notation includes "^" and "$", it allows writing cleaner more concise patterns. These notations add an implicit "." at the beginning and end of every pattern unless you use "^" or "$" to turn that off.
As for how you're supposed to know this, the puzzle tell you, but there is a very strong clue, which is that many of the patterns have a leading/trailing ".". This would be totally superfluous in one type of notation, so it must be the other kind.
Here are some patterns from the puzzle's notation:
.*H.*H.*
(DI|NS|TH|OM)*
F.*[AO].*[AO].*
and here are how they'd look in a notation that uses "^" and "$": H.*H
^(DI|NS|TH|OM)*$
^F.*[AO].*[AO]Edit: Fix HN oddity
% python3
>>> import re
>>> re.search(r'R*D*M*', ' ')
<re.Match object; span=(0, 0), match=''>
>>> re.fullmatch(r'R*D*M*', ' ')
>>>Tip: Ctrl-Z is your friend ;)
Any ideas how it was invented?
See this comment for the link to the original, including the author's name and the puzzle in context with its official solution:
https://news.ycombinator.com/item?id=26439598
(I wish that other comment would get upvoted to the top -- this was written by an identifiable person for a specific identifiable puzzle event, so it's not like mysterious anonymous Internet folklore or something.)
The final clue I had to complete was that tricky .*(.)(.)(.)(.)\4\3\2\1.*
Agreed, I definitely found myself doing that.
> The final clue I had to complete was that tricky .(.)(.)(.)(.)\4\3\2\1.
Interesting...I think I had the main guts of that one worked out when I was about 1/3 of the way through, but admittedly I did take a screenshot and mark it up in order to keep track of the various possibilities. It was (...?)\1* that tripped me up because I had an overly broad interpretation of its mechanics.
I wonder if it would be possible to record a bunch of 0-100% runs and then to make some sort of visualization to demonstrate different approaches...
Reload if you don't see the regexes turn bold when you click in a hex!
Really neat puzzle though!
Personally I'd even pay good money for a series of these style of puzzles.
Aren't crosswords supposed to spell out a word? This feels like random characters. Unless i'm supposed to rotate it or something?
It took me close to three hours to solve this, hahaha
Would be nice if they had a couple hand mande puzzles instead of random ones.
Python, for example, has a fullmatch method.[0]
libicu's matches() function returns true "if the pattern matches the entire string, from the start through to the last character."[1]
PCRE has various flags that change what it means for a regular expression to match, including PCRE2_ANCHORED and PCRE2_ENDANCHORED. Used together, these options would require a full match with no change to the regular expression itself.[2]
0. https://docs.python.org/3/library/re.html
1. https://unicode-org.github.io/icu/userguide/strings/regexp.h...
In Mystery Hunt puzzles, which this originally was, "we have to interpret this in a way that would allow there to be a meaningful and unique solution" is not only a perfectly legitimate form of reasoning, but often necessary!
It's not really exactly the same kind of reasoning, but in a puzzle I wrote a year before this for the same event
https://www.mit.edu/~puzzle/2012/puzzles/into_the_woodstock/...
you could look at it and say "hiragana is only ever allowed to be used to write Japanese!!!" but insisting on that rule (much as it applies in most situations) wouldn't give the puzzle a meaningful solution. :-)
Maybe a closer equivalent would be that in this year's Mystery Hunt, there was a puzzle using a set of variant Hashiwokakeru (Bridges) logic puzzles. There were hints about which rules were changed but it wasn't stated whether the rule changes applied individually (one puzzle each) or cumulatively (when rules get changed, they don't change back afterward), or some other way. So, it was necessary to make assumptions about what was meant and see whether they allowed a solution. That's typically considered fair and appropriate throughout Mystery Hunt-land.
But this puzzle doesn't recognize any of those as valid.
Assuming an empty 7-cell row, the regex in question will match 7 times, once for each character. A result of multiple partial matches is not equivalent to getting just one perfect match, which is what the puzzle requires. Internally, the puzzle enforces this requirement by making each Regex rule require a full line match (see line 118: '^' + rule + '$').
Personally speaking, I felt like that was the most intuitive way to interpret the mechanics of the game, so I'm willing to give the programmer a pass when they slightly modify the regex rule prior to evaluation.
[1]:https://github.com/Jimbly/regex-crossword/blob/8b178f32eba37... [2]:https://i.imgur.com/PS0BmJ6.png
R?(CR)*MC[MA]* is not matching for the full line RMCMMMMMMMMM
Why?
Are you sure you're going the right direction? It should be going from the bottom left to the right (in the order of the the regex text).