247 karma · joined June 7, 2014
I build machine learning stuff for fun & work. From Paris.
A really nice to have would be to have inline Latex mathematical formula, but it may be too specific for your app.
I just discovered Neocities BTW, it sounds very interesting!
I'm a data scientist with experience in various types of machine learning projects. I have a statistical background, and strong SWE skills.
I have experience in structured data, NLP, images, deep learning, recommendation systems, predictive maintenance, etc.
Why should you hire me? - I can understand your problem, torture your data and find the right model - or I will try something new, out of the box, if you need it - I know how to work with best practices for the code and how to integrate into a workflow - I love to deliver quality work
Drop me a mail! matthieu@databiz.io
P(B|A)>P(B) <=> P(A|B)>P(A)
These are the kind of questions you have to ask yourself. Defining a metric is hard, and there is no good shortcut.
- Both R and Python support strings, factors, and complex objects in a dataframe. What is NOT supported by feather?
- Feather is "not for long term data storage". Will it be standardize in a distant future?
- Do you plan to integrate it into Pandas?
Nice one nevertheless!
> Turns out unrolling tight inner loops really speeds them up! 38% speedup from that alone. Makes sense, it's doing one eight of the branching
https://twitter.com/notch/status/595509436040491008
Even billionaires may want to micro-optimize! Bonus point: he is developing for a "20 years old platform"
This rotted apple should not hide the good part of HFT, which is to reduce spread and inconsistencies between markets and to generate profits from this (positive) action. HFT took the place of traders, who were paid a lot for doing that stupid task.
>>> ds = u'2014-01-09T21:48:00.921000+05:30'
>>> %timeit ciso8601.parse_datetime(ds)
100000 loops, best of 3: 3.73 µs per loop
>>> %timeit dateutil.parser.parse(ds)
10000 loops, best of 3: 157 µs per loop
A regex[1] can be fast, but the parsing is just a small part of the time spent. >>> %timeit regex_parse_datetime(ds)
100000 loops, best of 3: 13 µs per loop
>>> %timeit match = iso_regex.match(s)
100000 loops, best of 3: 2.18 µs per loop
Pandas is also slow. However it is the fastest for a list of dates, just 0.43µs per date!! >>> %timeit pd.to_datetime(ds)
10000 loops, best of 3: 47.9 µs per loop
>>> l = [u'2014-01-09T21:{}:{}.921000+05:30'.format(
("0"+str(i%60))[-2:], ("0"+str(int(i/60)))[-2:])
for i in xrange(1000)] #1000 differents dates
>>> len(set(l)), len(l)
(1000, 1000)
>>> %timeit pd.to_datetime(l)
1000 loops, best of 3: 437 µs per loop
NB: pandas is however very slow in ill-formed dates, like u'2014-01-09T21:00:0.921000+05:30' (just one figure for the second) (230 µs, no speedup by vectorization).So if you care about speed and your dates are well formatted, make a vector of dates and use pandas. If you can't use it, go for ciso8601. For thomas-st: it may be possible to speed-up parsing of list of dates like Pandas do. Another nice feature would be caching.