High-Performance Log Parsing in Haskell: Part One
blog.safaribooksonline.com
blog.safaribooksonline.com
How can one say one parse is high performance and another isn't if no comparison is made? e.g. the author could use pylogsparser and benchmark it:
https://pypi.python.org/pypi/pylogsparser/0.8
Then we can get an idea of whether the project set out to meet its goal of being quicker than python.
In any event, I'd be interested to see how luajit-lpeg compares to their haskell impl. It even has a nice online test tool:
[1] that's just measuring the time to parse log files into some sort of structured data, not necessarily to do anything with it
After playing with Haskell for the last two years, I figured that it would be pretty damn hard to write conventional LAMP/JS webpages in Haskell, but I still wanted to try to solve some of my real world problems with it.
So I'm parsing the log files that my LAMP stuff produces.
So far it's a very fun and frustrating waste of time, but given my recent progress, it will probably become a slightly less fun and frustrating time saver.
I just recently cleared up some confusion and type errors coming from ByteString/Lazy ByteString/Text/Whatever conversion issues (which made up 95% of the recent frustration, as every other language I'm using right now has one single sort of string (which more than often serves to contain ints ;))
... and database CRUD and JSON typeclass instances have appeared as if by magic.
Awesome!
And now I can celebrate by unleashing that stuff tens of megs of logs, opening top, and then seeing the executable instantly eat up all my RAM, before slowly conceding space to the postgres server.
Hopefully your post will help me implement streaming.
Cheers!
Really?
You should check out:
http://adit.io/posts/2013-04-15-making-a-website-with-haskel...
Some tutorials will have a few routes and a DB, some other will have an intricate rendering and templating architecture and basically no data storage solution, some are more or less a hello world with a session manager... And there's Yesod, which seems to do everything out of the box, but which has such a colossal amount of dependencies that I never got anything to build beyond fresh yesod-init setups, which still manage to fail on my current setup, despite using Stackage and sandboxes.
Among what I've seen in Haskell, ORMs and web frameworks are the two types of libraries that carry the largest and most cumbersome monad transformer stacks around; making a website involves having to typecheck a huge monad transformer into another huge monad transformer. This is not easy.
Not to mention that, in the world of open-source web developement, nearly all the documentation, expertise, copy-pastable code examples and standard practices are in dynamic OOP languages.
Perhaps the ease of RoR, PHP and Node has let everybody forget how much stuff is involved in making modern webpages : )
You know people use Scotty in production right? It's more than cute, though I agree that most Haskell library/framework documentation could use some work.
> And there's Yesod, which seems to do everything out of the box, but which has such a colossal amount of dependencies that I never got anything to build beyond fresh yesod-init setups, which still manage to fail on my current setup, despite using Stackage and sandboxes.
Hm, did you try only using Stackage or sandboxes? If your up for another solution, I know Halcyon[0] is supposed to be very frictionless.
> Among what I've seen in Haskell, ORMs and web frameworks are the two types of libraries that carry the largest and most cumbersome monad transformer stacks around; making a website involves having to typecheck a huge monad transformer into another huge monad transformer. This is not easy.
Matter of opinion? I find Scotty and Snaps monad transformer stacks to be pretty simple. This is something that becomes easy with experience I think, just like getting used to using composition over inheritance.
> Not to mention that, in the world of open-source web developement, nearly all the documentation, expertise, copy-pastable code examples and standard practices are in dynamic OOP languages.
Mostly a function of manpower available I'm guessing. I find when you ask for examples or file bugs against libraries/frameworks they are responded to quickly though.
> Perhaps the ease of RoR, PHP and Node has let everybody forget how much stuff is involved in making modern webpages : )
You just have to remember how much polish they have and how it was nowhere near this easy in the beginning.
...
Halcyon, I'l remember that. I still have to try Nix one day, too
...
Matter of experience, I guess, yeah
...
Precisely, there's a virtuous cycle going around popular platforms due to the larger amounts of features and documentation getting written, and in turn, the larger amount of new devs getting into it. Shame those platforms are built on such bad languages... It's cool to see the web evolving in a direction where Haskell can solve very specific problems without disrupting people's stacks.
- one is dead
- another is the raw Hackage page for a session library spun off from another project
- the last is a cute hello world for the framework that underpins Scotty.
No Node.js tutorial would finish without a working session manager.
Here is a Yesod example[4] that uses Auth.
Here is a Snap example[5] that uses Auth.
> No Node.js tutorial would finish without a working session manager.
Yeah, many Haskell library/framework writers tend to see end to end tutorials as hand holding or see something as Auth as so simple it needs no explanation.
It was a roadblock for me personally.
0: https://github.com/agrafix/Spock
1: http://hackage.haskell.org/package/Spock-auth
4: http://www.yesodweb.com/book/blog-example-advanced
5: http://www.christopherbiscardi.com/2014/01/07/getting-starte...
I have a nice lil' prototype going on in Snap, I have Persistent running, and I'm this close to have auth working.
But in the meantime, I have lots of logs to parse ; )
It is indeed a sum type, but I don't think that is why (or true of all sum types, or less true of product types). It sounds like they are conflating "enum"?
As an example of a product type for whom we can define every possible representation: ((),()) has precisely one representation.
As an exampe of a sum type for whom there are infinite representations: (Either String Integer)
My understanding has it that it is called a sum type because 1) for finite types the number of representations is the sum of the number of representations under each tag, and 2) it behaves more generally like a sum (a^x * a^y = a^(x+y) is an equality in arithmetic, (x -> a, y -> a) is isomorphic to ((Either x y) -> a) in programming).
Would be good to talk more about this and quantify too. Very interesting topic.
Also note that there is a feature proposal[1] for Applicative Do notation.
0: http://community.haskell.org/~simonmar/papers/haxl-icfp14.pd...