Analyzing Your Browser History Using Python and Pandas
applecrazy.github.io
applecrazy.github.io
raw_data = [line.split('|', 1) for line in [x.strip() for x in content]]
can be simplified to a single loop: raw_data = [line.strip().split('|', 1) for line in content]
Using str.replace here is also non-idiomatic: plt.title('Top $n Sites Visited'.replace('$n', str(topN)))
How about using str.format instead: plt.title('Top {n} Sites Visited'.format(n=topN)) f"Top {topN} Sites Visited" sqlite3 ~/.mozilla/firefox/$YOUR_PROFILE_ID/places.sqlite "SELECT datetime(visit_date/1000000,'unixepoch'),url FROM moz_historyvisits JOIN moz_places ON place_id=moz_places.id ORDER BY visit_date DESC" > hist.txt
(Looks like the timestamp is in UTC here.)----
Does the original post's code perhaps count two visits to the same URL as one? Something like
SELECT datetime(visit_time/1000000-11644473600,'unixepoch'),urls.url
FROM visits JOIN urls ON visits.url=urls.id
ORDER BY visit_time DESC;
might be worth a try. Traceback (most recent call last):
File "hist.py", line 14, in <module>
data.datetime = pd.to_datetime(data.datetime)
File "/usr/local/lib/python2.7/site-packages/pandas/core/tools/datetimes.py", line 373, in to_datetime
values = _convert_listlike(arg._values, True, format)
File "/usr/local/lib/python2.7/site-packages/pandas/core/tools/datetimes.py", line 306, in _convert_listlike
raise e
pandas._libs.tslib.OutOfBoundsDatetime: Out of bounds
nanosecond timestamp: 1601-01-01 00:00:00
The time on my Mac is correct, and datetime gets the correct time?How long is the Chrome database kept for? I use Chrome Canary and the standard Chrome, but only had a few hundred results in each, a lot less than I was expecting. I can't remember the last time I cleared my history either.
Cool graph though anyway!
And, by the way, I’m not done with this series. Next week I’ll try to predict browsing patterns using this dataset.
data.loc[data.datetime == '1601-01-01 00:00:00', 'datetime'] = '1970-01-01 00:00:00'One note you may want to add (or step you may want to tweak) is that the initial sqlite3 command won’t work if chrome is currently open. I just copied the History folder to somewhere else and did the same thing, as per the superuser thread you helpfully linked to.
https://dataorigami.net/collections/all/products/create-mark...
pd.read_csv('hist.txt', sep='|', header=None)It does, use df[column_name].str.split(expand=True), where df is your DataFrame.
Edit: nevermind, it does do that (https://pandas.pydata.org/pandas-docs/stable/generated/panda...)
I haven't had any reason to use pandas yet so it'll be a good excuse.
It's a project that's not very long, requires data everyone has and gives a cool insight at the end.
Nice!
I haven’t made it very far yet, but so far it seems worth recommending.
~/.config/google-chrome/Default
or: ~/.config/chromium/DefaultIirc stock chrome also has a way to view this.
The main purpose of the post was to demonstrate collection and cleaning of data and give an overview of it through a basic visualization.
In the future, I'd like to show how this data is a gold mine of information, using it to predict browsing trends, create a profile of interests, and more.
Mainly to show why ad tracking/selling of user browser history is bad, but also to teach some data science techniques along the way.