Show HN: Delver, a natural language interface to your app
delver.io
delver.io
Most likely, you're just converting my questions into queries and parsing it over my data. Because your hompage is not clear if you are letting out just an API or you are giving out results of our data by yourselves. But if you gave out just an API, then it would mean we still have to process the converted queries and create our own interfaces to display the results. So I'm going to assume you take care of the data processing as well.
But, wait? You need access to my data to perform those queries. Let's say I have a million users. This means, you could potentially log every single query AND the result of the query into YOUR databases.
This means,
If I ask you "Which percent of my users pay the highest and are from the United States?"
And you perform a query on my one million users to find that out, you have data in your hands that my competitors or third party advertisers will come around you like sharks for. Whereas, all I can do at that point is just hope that you won't sell my data to them, which puts me and my data in a pretty vulnerable state.
Not to say that this is a bad idea. It's a brilliant idea, but I'm not sure how you are going to earn the trust factor.
I'd love to hear your thoughts on the minimum desirable interaction here - is an API that just translates natural language to, say, SQL useful enough for you to pay for?
Definitely!
If you get the natural language to SQL right, then it is multitudes easier for me as a developer to parse a JSON and display the results to an interface than having to learn the natural language processing part by myself and then try to interact and implement with it!
EDIT: I wanted your email. Nevermind, I found it.
It might be a completely worthless idea, but I would consider offering a "USB Stick" version rather than an appliance. It is much easier to replace, costs close to nothing, yet it can still be implemented as a self-contained blackbox with no customer serviceable parts.
The tricky bit is going to be your enterprise sales channel, you need someone with the right connections.
1) they often have sprawling, disparate data sources, with schemas designed by deranged people in the 80s or 90s. The cleaner (and admittedly smaller) your data, the easier you can integrate, and to be honest, there's probably a hard limit at some point beyond which we couldn't provide value without doing bespoke customisation.
2) there are some fearsome competitors at the higher level of the market, who offer wider business intelligence suites.
This may be fear-based thinking though.
1) The more sprawling and disparate the data sources, the more deranged the schemas, the more money to be had making sense of it.
2) Enterprise consultancies are ripe for disruption. There exist exactly zero satistifed customers of these companies.
"If the 3rd character in "customer ID" is 7 or 8, and the start date is after 2007, then they are a "premium" customer, and the maximum order amount doesn't apply, if they're shipping an order to Michigan, Texas or Florida".
"If customer ID is greater than 80000 and the date is after November 2009, then the real customer number should be reduced by 10000, because we had a problem and needed to reuse customer IDs". (meaning, an invoice for customer 12000 on November 2004 is related to a different customer than customer 12000 for an invoice on November 2010).
"If the employee's start date is >2005, then check table 1 for employee data if their last name starts with A-M, and table 2 if their last name starts with N-Z, otherwise check table 0 (legacy) Oh, and in table 4, if the employee ID starts with L, that means "legacy", so use table 0 to find their information, but remove the L".
These are situations I've run in to in the last few years, and I'm sure many of you have similar WTFs in your experiences. If someone has their data in good, solid, structured formats/tables, natural language syntax might fun/easy/exciting, but those people can also be served by things like Crystal Reports, some books, and a few hours of learning. The companies that most desperately need NLP->SQL probably also have the worst data.
For people with good data, even those with the expertise to query it, they'll still often have end users who want the data. The cycle of 'call IT department, ask for data, wait for data' or 'email SaaS provider, request report, wait for report' can be short circuited in these cases, and I believe that's of value.
https://docs.google.com/a/hotwoofy.com/forms/d/1ixCUouKsq1Q4...
However we wouldn't be able to even disclose some anonymised data, let alone have something communicate with the outside world that was munging our real data. Just the idea from security attack vector stance, effecively allowing any query would be a a deal breaker.
The problem is I can't see much happening in the way of tuning, we would be the clients from asshole land:
Oh yeah when I make a query, I don't get good results back Ok lets have a look, what's the query like Can't tell you that Ok, what's the data like Can't tell you that What can you tell us? System sucks
But obviously if someone else is providing data for tuning the NLP stuff to actually work on our data, if we can run the output as a AST, putting it anywhere we want as we would the output from our DSL, I could imagine the business case for paying a few cents per user, per month.
When you talk about tuning, are you saying you'd be unlikely to have time to train the system after initial setup? We're making an iterative model that'll allow you to add new concepts, new sentence structures etc as you go, and we've thought a few times it would be good to expose a log of queries (especially failed ones), and also allow end users to say 'this is wrong/nonsense' whenever they get results.
Time to train wouldn't be the problem, it would be a case of letting you guys near the data. The lawyers would have kittens. Having a nice tuning tools would be a good idea, as it allows us to do it.
Would the thing nock out an AST, or would it be SQL only? As it stands, one of the benefits of our own DSL (using Irony) is that we can implement the AST in T-SQL or just C# code against POCOs.
Also, in my distant C# days I was very impressed with Irony, glad it's still around.
I realize you might be aiming this application at people who actually can't read the SQL, but you're also saying "You're drowning in tedious reports." implying they actually use SQL and this is to replace it.
So make the examples consistent and give people some insight into how this works and what actual SQL query a question generates.
We need to work to clarify the message a little, because we're in the situation where we're simultaneously targeting IT departments and developers who would benefit from removing the burden of ad-hoc reports, but also their end-users who want the data. I'll give that some thought.
It's nice to know the technology is becoming available to the average programmer like myself ;)
[1]: http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.115...
[1]: http://gate.ac.uk/sale/dd/related-work/Kaufmann_nlp+reduce_E... [2]: http://gate.ac.uk/sale/eswc10/freya-main.pdf
However, I would love to have seen some demo.
I'll post again when we have a demo up.
Not that this is bad, but it is pretty far away from a general natural language interface where you basically have a conversation with your app.
That said, what you call trivia questions, we call decision support, reporting, and other grown up things. :P
It may be harder to see this from their website now, but they provide natural language interfaces to SQL databases.
PS: I am also working on something like this, though not fully defined as yet.
Also interested to know the model - is this a SaaS? if so perhaps some pricing info would be useful, or how it interfaces with your data, security considerations etc.
Business model-wise, we'll be offering per-seat licensing for internal apps, and likely traffic-based pricing for more public apps. We are a SaaS app - the intelligent bits happen in the cloud. The default option is to integrate one of our agents to gather data, or publish data to us via an SDK (or REST API). Because it's a deal breaker for many people, we're working hard to enable scenarios where data never leaves your network, but we know your schemas and ontology, and so the NLP bit happens on our side, but the querying happens privately. Obviously this can have an affect on our ability to do entity recognition, so it's an interesting problem.
We're really hoping to produce a demo integration soon, and I've mentioned Magento elsewhere, or something like Wordpress etc. I'll publicise that when it's available