Surfer: Centralize all your personal data from online platforms
github.com
github.com
I would prefer a cli tool with partial gather support. Something that I could easily setup to run on a cheap instance somewhere and have it scrape all my data continuously at set intervals, and then give me the data in the most readable format possible through an easy access path. I've been thinking of making something like that, but with https://github.com/microsoft/graphrag at the center of it. A continuously rebuilt GraphRAG of all your data.
It builds an entire ecosystem around your data where it is programmatic rather than just dumping text files. The point of HPI is to build your own stuff onto it and it all integrates seamlessly together into one Python package.
The next stop after Karlicoss is https://github.com/seanbreckenridge/HPI_API which creates a REST API on top of your HPI without any additional configuration.
If you want to get more fancy / antithetical to HPI, you can use https://github.com/hpi/authenticated_hpi_api or https://github.com/hpi/hpi-graph so you can theoretically expose it to the web (I am squatting the HPI org, I am not the creator of HPI). I made the authentication method JWTs so you can create JWTs where it will give access to only certain services' data. (Beware, hpi-graph is very out of date and I haven't touched it lately but my HPI stuff has been chugging away downloading data).
Some of the /hpi stuff I made is a bit mish-mash because it was rip-and-replace from a project I was making so you'll see references to "Archivist" or things that aren't local-first and depend on Vercel applications.
It is based around SQLite rather than Supabase (Postgres) which I think is a better choice for preservation/archival purposes.
It exported 75MB json of ChatGPT "Conversations". I extracted 19MB or raw text from this as a CSV. I then took this into Nomic.ai and embedded all of the text to create a clustered visualization of topics in my ChatGPT conversations.
1. Much tougher data privacy regulations (needed per country)
2. A central trusted, international nonprofit clearinghouse and privacy grants/permissions repository that centralizes basic personal details and provides a central way to update name, address(es), email, etc. that are then used on-demand only by companies (no storage)
By doing these, it simplifies things greatly for people and allows someone to audit and see what every company knows about them, can know about, and can remove allowances for companies they don't agree to. One of the worst cases is the US where personal information is not owned by the individual and there is almost zero control unless it's health related, and can be traded for profit.
It sounds to me like what you're describing under 2 is a real usecase for blockchain contracts?
Store your latest data encrypted on-chain and give every 3rd party you trust a key that corresponds to the relevent part of the data?
Curious about opinions on this.
It is important for privacy activists to understand that „centralised“ is an anti-pattern for privacy.
Instead we need security and control over our data on devices and internet platforms guaranteed by the law.
Still silly, but closer.
I started drawing individual frames of the game I wanted, I remember about 45 minutes into this venture I had an existential crisis about how how many frames you'd need for something like GTA, to show every possible combination.
I had the right idea but wasn't thinking about how to leverage the computer correctly.
The realization that every possible image that can fit on a screen can be stuck in a bitmap helped keep it going. Everything that could possible ever be photographed, just sitting there in the latent space waiting to be summoned.
... And now, I'm amazed that through some basically fancy noise we can type in words and get pictures in under a second. They even almost have the right number of fingers.
This lets me create dashboards to see usage for certain topics. For example, I have a "Dev Browser" which tracks the latest sites I've visited that are related to development topics [1]. I similarly have a few for all the online reading I do. One for blogs, one for fanfiction, and one for webfiction in general.
I've talked about my first iteration before on here [2].
My second iteration ended up with a userscript which sends the data on the sites I visit to a Vector instance (no affiliation; [3]). Vector is in there because for certain sites (ie. those behind draconian Cloudflare configuration), I want to save a local copy of the site. So Vector can pop that field save it to a local minio instance and at the same time push the rest of the record to something like Grafana Loki and Postgres while being very fast.
I've started looking into a third iteration utilizing MITMproxy. It helps a lot with saving local copies since it's happening outside of the browser, so I don't feel the hitch when a page is inordinately heavy for whatever reason. It also is very nice that it'd work with all browsers just by setting a proxy which means I could set it up for my phone both as a normal proxy or as a wireguard "transparent" proxy. Only need to set up certificates for it work.
---
[1] https://raw.githubusercontent.com/zamu-flowerpot/zamu-flower... [2] https://news.ycombinator.com/item?id=31429221 [3] http://vector.dev
1. Use Mobile App APIs.
2. Generate OpenAPI Arrazo Workflows.
1 ensures breakage is minimal, since mobile apps are slow upgrades and older versions are expected to keep working. 2 lets you write repeatable recipes using YAML, and that makes it quite portable to other systems.
The Arazzo spec is still quite early though, but I am hopeful of this approach.
1) Constant updates to existing packages 2) Continued expansion of more sites/apps to export your data from
I don't know anything about trademarks, service marks, etc but I do know that the product name "Surfer" has been in use for about 40 years in my industry, geoscience, by a company in Golden, Colorado. [0]
Maybe you can make a new product in a different industry and recycle the name. I don't know how that works but right now, you're playing in an established product's namespace.
https://surfer.nmr.mgh.harvard.edu/
https://towey-websurfer.apponic.com/
Also many open source libraries have also used the name:
https://rubygems.org/gems/surfer
https://www.npmjs.com/package/surfer
etc
[0] https://www.goldensoftware.com/products/surfer/
Cool name but it was taken way back when I was writing geoscience software. That's been a while.