Little Data: How do we query personal data?
petekeen.net
petekeen.net
The heart rate monitor would know when you spiked yourself with caffeine, combined with the sleep monitor that knows you've only slept 4 hours the past 3 nights.
An alert pops up on your computer: "hey, you need rest. Stop drinking so much caffeine. Eat an apple and take a 20 minute nap, you'll feel better."
Those are just two data sources.
Harvest (time tracking) + RescueTime (productivity) + Runkeeper would be super valuable too. You'd know that you're 10% more productive on days that you run.
This, among other things, is the reason I started building http://personable.me.
Ultimately, I see it as being the recepticle for WiiThings data, Fitbit data, etc., and I'm working on those now. I just finished up (but haven't deployed) Webhooks, with the idea that input + webhooks for output should allow a personal API to interact with other systems when API data is updated.
To that end I've been working on my own analysis tool for personal data. Most of the work has been on a Windows desktop app, and data import is limited to CSV. But data is synchronised to my phone.
https://www.wittenburg.co.uk/Entry.aspx?id=0a505400-5bf6-4a6...
Interact looks neat, but I suppose that this has a slightly different objective. To me, for Personable, the aim is that the API interacts with other services you use. Perhaps for the purpose of making your life easier (e.g., enter a personal API endpoint instead of having to fill out a profile, and let the website read in the data you let them access), perhaps for making the lives of others easier (e.g., your friends want to know where you're at, but don't want to check Facebook, Foursquare and Twitter to see which checkin is the most current), but all of that interaction depends on at least some of that data being net-accessible.
I haven't gotten into what should be public vs. what should be private yet, so, thus far, the advice for Personable is to just treat it like you should treat every other cloud service, and not put anything there you wouldn't want the world to read, but privacy options are forthcoming as well.
There's no inherent reason you can't profit from an open source service, especially if you're able to provide value to those willing to pay for the convenience of not having to run the service themselves, and/or those unable to.
Good luck either way. Closed source may not be the death knell you expect.
I want the data sources to include far more than just health data--GPS, title of active computer window, sms history, file modification times, device and sensor readings, snapshots from your webcam, etc--really anything that you can record.
If anyone is interested, feel free to email me!
- All data is just a stream of "events" from different "sources" for any given "data type"
- Sources are things like my desktop, my laptop, my phone, my camera, etc
- Data Types are things like global position, cpu usage, webcam snapshots, webpage text, etc. Anything really.
- Events are just a timestamp + zero or more other "lightly typed" "fields".
- By "lightly typed" I mean the field is either: binary data, searchable text, a float, or an integer
- Fields give the actual data for the event. If there are no fields, it's a countable event (ex: imagine recording every heartbeat, all you need is a timestamp). For something like global position, you would have latitude and longitude fields. For another datatype called "semantic location" you have just a searchable text string with things like "car", "work", "home" or whatever other labels you have. Fields can also be binary data (like photos, videos, etc). I'll probably try to make lots of data types, each with very few fields, to make it easier to work with the data (ie, it's nice if there is only one numeric field in any given data type, then it's easy to graph over time)
I'm currently collecting all of the events via some scripts that are specific to each data type, and then emailing those events (serialized as json) to my own server, which just inserts the events into a sqlite database specific to that source/data type pair.
The nice thing about this approach is that anyone can easily send an email from any language, and if you don't want to use my backend, you can just send the events to your own email server and then do whatever you want with them once they're there.
Another benefit is that you and I don't have to agree about exactly what fields and data types there are. If you call your gps data "gps" and I call mine "global position", well, who cares? It's our own personal data. If this takes off, I'll make guidelines about what to call what information and how to format it for easier interoperability.
Bet seriously, security is definitely a goal of this project.
(full disclosure, I am the founder of this company)
edit: fwiw, I understand if it won't be.
It works by connecting to services you might all ready use etc.
(full disclosure, I am the Co-founder of this company)
We need to standardize on common formats so that anyone can publish in them (either original data or a bridge from a custom API) and anyone can write code to do analyze and augment it.
Some fields - particularly in the "sciences" - are already linking massive datasets between them, allowing new applications that they had never even considered, but here in the common "web app" world we're still stuck in the world of non-standard APIs with custom and non-extensible JSON formats.
We have that planed down the pipeline:D
Project Description
Personis supports an accretion/resolution approach to reasoning about people, places and devices. It was designed to support user control of both the information held about them and the way that it is used.
Like balancing your checkbook once/month. (You see, we used to have these little books called "registers," and ...)
Much of my quantifiable data is sent to Google's Fusion Table, but I do not feel this is a good long term solution.
My intention is to define a base set of criteria for how a certain database needs to be formatted, similar to another comment here, and then let anybody make their tools available as either for "collection" of "visulisation/analysis". As long as some core fields are standardised, e.g. "quantity" and "date", then the data can be easily analysed. Each individual would control their own database, either on their own host or as a DBaaS, but tools to collect and visualize data could be shared.
I am leaning towards a document store (e.g. CouchDB and Cloudant), as that would allow any tool to push data in without knowledge of the schema. One of the nice things about some of the DBaaS is you can easily create individual username/password or API keys with specific permissions, so a third-party tool could write records, but not necessarily read any of your data. A standardised database would also benefit by having other tools able to utilised it, unlike something like Google's Cloud Datastore (which I do like!) In particular, I am thinking about the CouchDB and ElasticSearch integration.
So, why not just wait for apps like TicTrac or Saga to support every service? I have two reasons. Firstly, many of the other tools to aggregate data seem to have gone out of business. Secondly, there are some services I do not like to give third parties access to; email is one example, as is the ability to log keystrokes on my computer. However, I would like to see the summarised data from these services recorded with the rest of my Little Data.
Another option would be to dump everything to text files and upload them to Google's BigQuery, but I am leaning towards a shared tools / individual database model, as it would probably encourage better collaboration with other people.
If so, please contact me: rpedela [at] datalanche [dot] com