A novel approach to entity resolution using serverless technology
tilodb.com
tilodb.com
We built TiloDB as the tech team at a European consumer credit bureau when we were faced with the technical challenge of how to assemble hundreds of millions of data sets about tens of millions of people in a way that is scalable and allows fast searching, without breaking the bank.
We tried various technologies, such as graph databases, but none of them could give us satisfactory performance.
So we turned to the opportunities of serverless technology (AWS specifically) to build a new type of entity resolution technology.
In this article we write about the technology breakthroughs that led to TiloDB, and there is also an interactive demo where you can submit data, see it linked, and see other people submitting data in real-time.
We want to spin the tech out into a new company, release it as OSS, and so are keen to hear about potential use cases you might have.
Setting up a business based on a new DB tech that has one user, though, is tricky.
Playing devil's advocate, how do you plan to make money? Who are the users, why do they turn to TiloDB, how do they learn about it, how do they adopt it, how do they be convinced to pay you something for it? Etc
So we want to make the software open source, but restrict a few modules that would be necessary for enterprise customers, such as security and auditing features.
We have quite a few companies lined up that want to do proof of concept trials with us. So far they are mostly big fintech companies that use it for anti-fraud, also AML/KYC companies that need to match and search lots of data from different sources in real time. Also very large companies that need to solve their "data silo" problem.
Adoption - hopefully they start with the OSS version, play with then want to upgrade to the enterprise version.
One area we have less experience is with which type of OSS licence to use.
I see a lot of similarities between Kafka and Confluent. You're looking to spin out a tool that worked well for you, and offer it commercially.
I'd suggest planning more around operating TiloDB as a managed service. You happen to just need lambda, s3, and dynamo today, but the "capturable" value for many customers will be if you also manage it all for them (especially upgrades). You can still offer the open core and let folks run their own, but it sounds like a lot of the goodness comes from the way you run it.
Having said that, licensing is currently fraught in this space. Each major "database" vendor (Elastic, Redis Labs, Confluent) is basically trying to find a way to figure out how to avoid AWS (and other clouds) from just taking their code and operating it as a service.
People have very strong opinions on this topic, ranging from "open-source isn't a business plan" to "AWS is violating the spirit of the OSS community" and many more. My personal advice would be to assess more clearly why you want to be open source (you mentioned community and applications you couldn't imagine) and whether you think open source better achieves those goals than say a free tier or distributing a core binary / container image for free.
What, more specifically, do you want to get out of being open source? Contributions to the core? Contributions to the operational part? More users and feedback?
re OSS strategy - it's more users and feedback that I firstly think of. For instance, people keep telling us that there could be a really great use case in crypto compliance/auditing - tracking related wallet addresses etc. We don't really know enough about blockchain to validate that, but I think an OSS community could.
It seems to me that all the knowledge / experience of your staff is your real asset here, not necessarily the code. While that, in theory, should mean open-sourcing the code won't matter for your business, in practice it means you will be seeding competitors unnecessarily.
Turning your staff into highly chargeable consultants could be a more sustainable business model than trying to herd the cats of the internet into trying to improve your product offering, when most of those guys won't have the experience of your existing team. By offering the code out as open source you are giving a bunch of people a leg up and cutting short your time as the only game in town, which puts extra pressure on sales and might not work out.
For one Graphistry project, we run a single node neo4j with 0.5b nodes/edges, so something in the description isn't adding up for me here wrt perf. Maybe an open benchmark would help?
I do agree indexing matters, as that was night/day for our use cases. For ML workloads, we are looking at vector indexes, which graph DBs do not currently support. The ones in this article are on text and take > 100ms, so I'm curious..
The response times provided in the article are for the whole process of searching and returning the entity. The indexes themself are obviously a lot faster - to be precice we are using DynamoDB for storing the indexes, which most times return results in <10ms. Compared to other databases this may still sound slow, but we know that we won't run into scaling issues in this way and that's kind of what matters currently most for us.
Hope that somehow makes sense what I wrote.
RE:extremes, we see graph DBs OK for small time series (ex: 2 nodes with a bunch of event multiedges), but not full blown time series... where we'd use a tsdb. Some vendors demo this, but always felt like wrong tool.
The many-hop case is interesting! We don't see 1K-hops typically, and I get nervous even at 10-20 on graph DBs we've used. I can imagine in logistics or sciences that happening more, or maybe even some rdf systems. Partition keys start mattering fast, whether a kvdb or a mpp, but I don't have an intuition here. Probably easier to differentiate on, but too niche?
1k hops is also not something we see on a regular basis in our old business, which is much about people moving houses and transactional data from payment service providers. Ppl with money issues seem to move a lot more often and also fraud cases often have a lot of hops.
What licence should a new company like us adopt when we want to build a community but we also want to commercialise the technology, especially when we already have a "enterprise ready" version of the tech?
The batch mode had naturally orders of magnitude higher throughput. We did have real-time single-record mode which was pretty fast as long as the stream of the incoming single-records wouldn't saturate the worker array capacity (here is the difference from serverless as the worker array was limited by whatever was statically configured at the moment as adding/removing nodes wasn't an instant on the fly operation)
Couple years later i worked at another company on a similar, though somewhat simpler, project when it was in the process of total rewrite for performance reason - the old version was really slow - that rewrite failed spectacularly for a lot of reasons. So, yes, performance is a kind of a noticeable factor in the domain.