Dateparser: Python parser for human readable dates
github.com
github.com
arrow.parser.ParserError: Could not match input to any of
['YYYY-MM-DD', 'YYYY-MM', 'YYYY'] on '01-06-17'
Here's a list of (US English–centric) test dates that dateparser ((ddp.get_date_data(date)['date_obj']).date()) and delorean (delorean.parse(date, dayfirst=False, yearfirst=False).date) both parse correctly, nearly all of which arrow fails on: 01-06-2017
01-06-17
2017-01-06
01/06/2017
01/06/17
2017/01/06
Jan 6, 2017
Jan 6 2017
2017, Jan 6
2017 Jan 6
2017, January 6
2017 January 6
January 6, 2017
January 6 2017
January 6nd, 2017
January 6rd, 2017
January 6st, 2017
January 6th, 2017
January 6nd 2017
January 6rd 2017
January 6st 2017
January 6th 2017
2017, January 6nd
2017, January 6rd
2017, January 6st
2017, January 6th
2017 January 6nd
2017 January 6rd
2017 January 6st
2017 January 6th
01//06/2017
01//06//2017
01--06-2017
01--06--2017
01/06-2017
01-06/2017
I like Delorean's API better than arrow's (strictly personal preference) but think dateparser's language detection is interesting.[1] http://delorean.readthedocs.org/en/latest/quickstart.html
What's the correct parsing of that date? Is it 2001 or 2017 or 1917 or ... is it June or January ...?
https://github.com/codygman/git-date-haskell
Had to do a small bugfix from bitrot, but this is a nice package that wraps the git date handling code which the author claims[0] was the only sane implementation he could find.
0: http://stackoverflow.com/questions/9831956/parsing-fuzzy-dat...
https://docs.google.com/spreadsheets/d/1dKt0R247B8Mx5sFXd7ht...
>>> from dateutil import parser
>>> parser.parse('')
datetime.datetime(2014, 11, 24, 0, 0)
It gets worse with fuzzy parsing: >>> parser.parse('something meaningless', fuzzy=True)
datetime.datetime(2014, 11, 24, 0, 0)It seems that dateutil has just not been receiving much love from its developers lately.
Does it handle proper grammar for singular values (i.e., 1 week vs. 1 weeks)?
Can you elaborate what you mean by "proper grammar for singular values"?
For example as I write this HN url says that it is 8 hours old. Without knowing the exact format how can I extract these sort of dates out of random text/html documents?
https://github.com/bear/parsedatetime/blob/master/parsedatetime/tests/TestNlp.py
may be what you're after.