> What ensures that the “rest of the string” is, in fact, a suffix of the input?
This is not something you want to guarantee.
takeWhile, dropWhile and many are three parsers which might consume 0 bytes. Backtracking via some kind of try parser would probably be unhappy as well.
> I assume the purpose of this design is for compatibility with lazy input sequences
Probably not.
> I would want to know the position and length of the parsed output
I don't understand exactly what you mean by these, but these would like be two 'parser' functions (which don't advance through the string), having types:
positionOfParsedOutput :: Parser Int
lengthOfParsedOutput :: Parser Int
I don't think these two are implementable directly from the given definition of Parser.
I'm quite far through a language implementation project, and I'm very happy about having rolled my own parser combinators.
The core I used was:
newtype Parser s a = Parser { runParser :: s -> Either ByteString (s, a) }
* Either is better than Maybe because it gives you room for an error message.
* s is polymorphic instead of String, and here's why:
Your job becomes insanely easier if you split parsing into a lexing and parsing phases. But a lexer is just another parser, which different inputs and outputs. Look!
lexer :: Parser String [Token]
parseExpression :: Parser [Token] Expression
But back to
positionOfParsedOutput and
lengthOfParsedOutput which I said weren't implementable from the above. Simply replace the
String type from your definition with a tuple or a struct, in order to carry the extra information.
data ParseState =
ParseState { remaining :: String
, pos :: Int
, len :: Int }
newtype Parser a = P (ParseState -> Maybe (ParseState, a))