HNHacker News
TopNewBestAskShowJobs

mohaps

584 karma · joined May 26, 2011

author, artist and bonafide geek (http://mohaps.com). twitter: @mohaps
submissionscomments
mohaps··on Ask HN: What's your funniest hiring interview story?
Interviewer: What are your strengths? Me: Cache Invalidation and Naming Things.

We paused and looked at each other for a minute and then we burst out laughing.

mohaps··on Spiral: Self-tuning services via real-time machine learning
the basic difference in approach was to switch the earlier approach of

if (conditionA && conditionB && !conditionC) cache_it()

to

hey look, an item with featureset X is cacheable while one with featureset Y is not.

This reflects in the API which is just two calls predict() and feedback()

this simplifies the integration code and is easily debuggable even in the face of changes.

mohaps··on Spiral: Self-tuning services via real-time machine learning
To add to Vlad's comment above: The major win was in being able to create a simple API which inherently computes update -> cache invalidation probability taking into account the semantics of the item being cached and its relevance to the query. To be able to do this explicitly via heuristics leads to bespoke effort for each new and modified query.

With Spiral, we were able to approach this top down as a classification problem.

e.g.

If you have a cached query for "Friends that liked my post", the Spiral classifier quickly learns that "Post Last Viewed At" or "Post Last Modified At" is not relevant to this via the feedback from the caching code.

Pre-spiral, this was expressed via a curated blacklist/whitelist which had to be recreated if the query characteristics changed.

mohaps··on Spiral: Self-tuning services via real-time machine learning
One of the authors here. Be happy to answer questions
mohaps··on Show HN: PaleoCodex, a prehistoric species catalog (using wikipedia data)
Thanks! Working on the timeline feature. The temporal data and diet/size tags need some clean up - a tedious work via an offline tool with RegEx and some nltk elbow grease :)
mohaps··on Show HN: PaleoCodex, a prehistoric species catalog (using wikipedia data)
I built PaleoCodex as an online interactive encyclopedia for Dinosaurs and other prehistoric animals for my six year old using wikipedia data.

He is a big dinosaur buff and wants to be a paleontologist when he grows up. His project brief was "Build a pokedex... but for dinosaurs"

Still kind of a WIP (Search is very basic and Text to speech / Listen feature may not work on iPhones, works on iPad/Nexus and Desktop using Chrome and Safari) and very much a weekend hack.

Built using python, tornado, wikipedia data and the skeleton html5 boilerplate. Uses ResponsiveSpeech for text to speech.

Enjoy! (and do send in your comments and suggestions)

mohaps··on Show HN: A header only C++11 LRU Cache template class, with no dependencies
now i remember why I went with the old linked list implementation.

with the std::list, everytime I refreshed a node... i did a copy.

basically to the tune of list.push_back(*iter) where as I wanted to keep it as a simple unlink and relink of the node

copy of the comment: https://news.ycombinator.com/item?id=12391069

mohaps··on Show HN: A header only C++11 LRU Cache template class, with no dependencies
ah, now I remember why I had carried on with the linked list from the old c++ codebase

https://gist.github.com/Quinny/09e34afb1d187e43dea2f3c3b0c04...

the key refresh in std::list required a copy of the std::pair, where as I wished to keep it as an unlink / relink

mohaps··on Show HN: A header only C++11 LRU Cache template class, with no dependencies
thanks. fixed
mohaps··on Show HN: A header only C++11 LRU Cache template class, with no dependencies
yeah, saw the gist posted by quinnftw. that actually is a better way to do what I'm doing. Will update that along with the

cache.remove(const Key& k, F deleteCallback) method.

Thanks for the feedback

mohaps··on Show HN: A header only C++11 LRU Cache template class, with no dependencies
the naked pointers are inside an internal kv namespace and are a convenience/carry over from the previous implementation (I've covered the justification for the way it's written in a comment above).

the TL;DR is that it is written the way it is to achieve constant time remove/remove_and_add_to_end operations which is not an erase+append but more like an unlink/relink operation.

mohaps··on Show HN: A header only C++11 LRU Cache template class, with no dependencies
the kv::List is not intended as a generic linked list implementation like std::list.

The reason for the specific implementation is to get constant time removal and remove-add_to_end operations. Truth be told, it is a hold-over from the previous implementation. I shall be writing benchmarks for this next up and shall revisit the decision on whether or not to use std::list soon-ish :)

The : private NoCopy is just to block the copy constructors.

mohaps··on Show HN: A header only C++11 LRU Cache template class, with no dependencies
yup! that's next up. :)
mohaps··on Show HN: A header only C++11 LRU Cache template class, with no dependencies
mostly: yes! :)

Also, keep in mind that most of the library is templates, which need to go in the header files.

BTW, most of boost is header-only :) e.g. boost/noncopyable.hpp would give you a header only equivalent of the NoCopy class I wrote.

I've always felt weird about having to include some big-arse library for just using a simple container, so for things like this I prefer to write header-only versions.

mohaps··on Show HN: A header only C++11 LRU Cache template class, with no dependencies
header-only usually means there's one or more .h/.hpp files that you can include in your source. There's no source file or library to compile and include

e.g. if I had Hello.h and Hello.cpp, you'd need to either add Hello.cpp to your make/build or build a library and link to it.

Header-only version means all you do is

#include "Hello.h"

and you're all set to go.

mohaps··on I don't think the Port Authority would like it if you looked at this photo
Hmm, I wonder how's the pokemon population out there? That'd drive them nuts too! :D
mohaps··on What's your favorite children's book?
Oh The Places You'll Go by Dr. Seuss
mohaps··on Show HN: Flask-Ask – Amazon Echo Development in Python
Good work on the lib. I think EchoSim is provided by iQuarius media and not Amazon.
mohaps··on Show HN: Pass a URL, get summarized content
Any details about the backend/implementation?

shameless plugs for two similar projects(open sourced both) I did a while back 1) Algorithmic Summarizer: https://github.com/mohaps/tldrzr 2) Readability Clone / Article Body Extractor with summary, significant image and text : https://github.com/mohaps/xtractor

Both are deployed on heroku and the urls are in the github readme files.

mohaps··on G is for Google
ha! Any problem in Software Engineering can be solved by adding another layer of abstraction! :D
mohaps··on Why we maintain desktop apps for OS X, Windows, and a web application
Atom/Electron stuff from github also allows you to package a HTML5 application as a native desktop app : http://electron.atom.io/

kinda like cordova/phonegap for desktop

mohaps··on Chrome address spoofing vulnerability proof-of-concept for HTTPS
but valid phishing can occur with "Please call this number" type of scam! :(

I know my dad would fall for that.

mohaps··on Why Are Geospatial Databases So Hard to Build?
loved the article. I strongly believe that when doing geospatial indices, it is quite important to up front establish the nature and distribution of your data.

Searching for a generic/theoretically all-encompassing solution is quite akin to looking for the perfect pub-sub system for all workloads. :)

My comment was related to lat/lon indexed data clustered around cities. If you know beforehand that you're going to deal with such a dataset (static or dynamic), one can pragmatically decide on some sort of by-convention partitioning (e.g. Partitioning by continent/region which will bring it down to in-memory indices not needing continuous disk access).

I love the non-euclidean bit in the piece you wrote. Anyone who has tried to do a k-nearest neighbor query east of New Zealand will appreciate that bit. :D

mohaps··on Ask HN: Sci-Fi novels?
Personal favorites: "Flowers for Algernon", "Snow Crash" and "Schild's Ladder" #YMMV
mohaps··on Ask HN: Why did the BEEP protocol never get much traction?
yeah. I always find RFC 3117 (basically design notes) very interesting to read : http://tools.ietf.org/html/rfc3117
mohaps··on Ask HN: Why did the BEEP protocol never get much traction?
fair point, erkose. But why didn't the JSON-ified successor to BEEP come about?

As someone who has spent the last 15 years basically trying to plumb bandwidth aware/adaptive chatty AND bulk-data-transfer applications, I always look at the RFC's and see a ton of good ideas.

No fan of XML... but that's a personal bias.

mohaps··on ‘Human computer’ Shakuntala Devi has died
You should look up the tragic case of the Indian mathematician V.N. Singh to get some perspective :( http://articles.timesofindia.indiatimes.com/2013-04-19/patna... Shakuntala Devi had amazing calculation prowess, but to me it was always seemed more in the savant zone.
mohaps··on Show HN: TL;DRizer - an algorithmic summarizer webapp/api in java (weekend hack)
yeah, that's what I ended up doing. :) first time using git. These old bones have been ground up bad by CVS/Subversion :P Here's the announcement: https://news.ycombinator.com/item?id=5535827
mohaps··on Show HN: TL;DRizer - an algorithmic summarizer webapp/api in java (weekend hack)
TL;DRzr is now open source! https://news.ycombinator.com/item?id=5535827
mohaps··on TL;DRzr (Weekend hack of a Summarizer) Open Sourced
As promised on previous thread https://news.ycombinator.com/item?id=5523538 :)

Have fun, be good, enjoy! :)

Page 1 of 2Next →