To Be Continuous
pipelinedb.com
pipelinedb.com
Esper does exactly this - you run streams of events over it and it continuously executes SQL to see if it matches. If so you can:
- run code
- make new streams
- store the results
Esper's been doing this kind of thing for 9 years now.
> Seems good but... such a weird project. Codehaus, svn, not high activity, but consistent, stable releases for five years. Maybe just not the kind of thing webdevs get into? Not sure if that's a strike against or not.
With the demise of Codehaus it looks like they've moved to github:
https://github.com/espertechinc/esper
But oddly they don't seem to have migrated their svn history, and the README implies they don't plan to... I certainly hope they didn't lose it when Codehaus shut down. There was, as noted, already 9 years worth of code changes in that repo. That would be unfortunate.
/usr/lib/pipelinedb/usr/lib/pipelinedb/bin/pipeline-init
Is this intentional?
EDIT:
After playing around with the .pkg file it looks like the packed Payload contains '/usr/bin/pipelinedb/usr/lib/pipelinedb' which is probably the problem. I see broken symlinks for pipeline-init etc in /usr/bin pointing to /usr/lib/pipelinedb, so I'm guessing this repetition of the path above is a mistake.
Also I see a postinstall script creating a symlink from pipeline to psql. This seems like a bad idea as psql is pretty universal already as the name for the PostgreSQL CLI binary, maybe 'pipesql' might be better?
Please shoot me an email (I'm Derek) if you have any issues installing this package. Thanks for your patience!
I was under the impression that the academic projects had proposed StreamSQL as a general language, though since StreamBase's acquisition it now seems to have been branded as TIBCO StreamSQL[2]. Have you guys been part of any efforts to make sure that there is an open language standard?
[1] http://streambase.typepad.com/streambase_stream_process/2013...
[2] http://www.streambase.com/developers/docs/latest/streamsql/
To your point about promoting language standards, we've intentionally kept the syntax as close to SQL as possible in order to keep things simple. The goal has always been to give the broadest range of developers the simplest way possible to develop realtime applications using only SQL.
Truviso got bought by Cisco and disappeared into their internal projects.
StreamBase got bought by TIBCO and is still available today.
In terms of not requiring that raw data be stored, a typical setup is to keep raw data somewhere cheap (like S3) so that it's there when you need it. But granular data is often overwhelmingly cold and never looked at again so it may not always be necessary to store it all in an interactively queryable datastore.
As I mentioned, PipelineDB certainly doesn't aim to be a monolithic replacement for all adjacent data processing technologies, but there are areas where it can definitely introduce significant efficiency.
You can do anything with PipelineDB that you can do with PostgreSQL 9.4, but with the addition of continuous SQL queries, sliding windows, probabilistic data structures, uniques counting, and stream-table JOINs (what you're looking for here, I believe.)
But I don't have a lot of use cases in personal projects, and am unlikely to find a good use-case at work in the near future. What's the 'adoption path' for something like this?
I think a really robust sample data set with example queries (think the neo4j imdb examples) would be a great way to show how powerful and easy something like this can be.
The main tradeoff with PipelineDB and other stream processing frameworks like Riemann, Storm, Spark Streaming, Samza, and others is mainly flexibility for simplicity. Not all streaming computation lends itself to SQL, but in scenarios where it does continuous SQL queries and a relational database can be simpler. But as with all data processing endeavors, you have to find the right tool for the job.
Would it be possible to set triggers or something on the continuous views? Lets say I want to take action (immediately) when a value calculated over sliding window goes above a limit.
It's a bit late here but I'll definitely play with PipelineDB tomorrow.
Awesome--let us know what you think about it!
Found this gem in the docs :D
Unsupported Aggregates: xmlagg ( xml )
:(- State in the data. In many sources we have, processing depends on some internal state, which must be kept along the time. For example some process has started and we will know when it ended, and we must keep its state so we could correctly process the ending event (to match it up). I am not clear how this will work with continuous views. I would say this is actually the major reason of what makes ETL processing non-trivial.
- Processing failure. Let's say something goes wrong and the data processing fails (or it can actually be even planned downtime). How do we know where to restart, to avoid processing data twice or miss data? Does the continuous stream take care of this metadata? And how does it deal with the state information per above? If you do data processing in batches, there is an obvious point of restart. Again, I think the extra complexity that "continuous" approach says is unnecessary relates to the fact that you want to be able to checkpoint the state of processing for various reasons.
If so, this could be an interesting alternative to RethinkDB's changefeeds, as RethinkDB doesn't support joins on the change stream.
Currently continuous views must read from a stream. However, in the very near future it will possible to write to streams from triggers, which would probably give you enough flexibility to model the behavior you want if you could conceptualize a table as a stream of changes.
Our next release (2.1) is due in about three weeks and includes automatic failover/high availability. Feeds on table joins (and other greatly expanded feed functionality) will be in 2.2, which should happen ~6-8 weeks after 2.1.
(Sorry to jump in with a shameless plug; what PipelineDB is doing is super-cool; I also met the founders a few times, and they're awesome, smart, and very driven people -- I'm really excited about what PipelineDB has to offer!)
> diff <(cat a.txt) <(cat b.txt)
My first thought (aside from "Cool") was that the current time would be the tricky thing that can't be incorporated into a continuous view. But even that seems to be handled! http://docs.pipelinedb.com/sliding-windows.html
Looks pretty impressive. :-)
http://www.postgresql.org/docs/devel/static/logicaldecoding.... and 9.5's track_commit_timestamp = on.
The documentation states it as "by default". It's required to change wal settings in config file and write a custom logical decoding output plugin in order to take advantage of that feature.
Currently, PaaS provisers such as RDS and Heroku Postgres don't support and it would not be easy to setup it manually. AFAIK, that feature is intended to be to used for backups.
However, we love Postgres and plan on actively merging upstream releases!
Is the license decision driven by business or is there some dependency that pushes you to GPL? For us Apache 2.0 has been worth it even when other companies use our code in their products.
I'm experimenting with the EventStore pattern for a side project, and I have struggled to implement projections. Could PipelineDB be a way to deliver that?