Osquery: Expose the operating system as a relational database
code.facebook.com
code.facebook.com
The next step would be manipulating the OS as relations. E.g. an insert into the process table allows you to actually spawn a process. It would start to get really interesting from there...
[1]: http://incidentalcomplexity.com/2014/10/16/retrospective/
For Plan 9, it was "everything is a file" and the power of using simple operations like bind mounts to create complex software interactions that would otherwise require monolithic protocol and library stacks anywhere else.
Here, it's... "everything is a table"? I'm not very familiar with table-oriented programming, but what advantages does having the RDBMS be the prime metaphor over the file system really bring? Structure? Rob Pike had some interesting words on that: http://slashdot.org/story/50858
Easy queryability and easy joining of data across a whole datacenter.
This makes it easy to think about system data across all sorts of boundaries that the file metaphor makes somewhat cumbersome.
while file metaphor became big in previously SQL domain of data processing, i.e. the whole ecosystem of HDFS and everything on top of it
One advantage of "everything is a table" is that your structures are well-formatted and there's no risk of problems when you put a space in a pathname. For most implementations of "table", you can also have the data formats be well-typed. This brings reliability and security benefits.
I think there's validity to Rob Pike's argument in many contexts -- for instance, you absolutely won't see me defending the semantic web over the greppable/Googleable one. But in the specific case of text files with a single, well-defined structure, his own argument seems to imply that there's no sense in a second tool having to infer the structure on its own.
(The usual way this is worked around these days is separate files for each field, or files designed to be parseable, which is why Linux's /proc/*/ is such a mess. Compare /proc/self/stat and /proc/self/status, and /proc/self/mounts and /proc/self/mountinfo. Also look around /sys a bit.)
This is great. One of the frustrations I've had with Puppet and Ansible is the lack of a clear model for data. It's quite difficult to know the scope and dependencies and origin of all the variables that one deals with.
If one could update tables and then have that representation be reified to the machines it would be awesome.
This approach could be a good fit for package management. So that packages are updated and run within a transaction and changes can be committed in a single seamless step.
That said, it would be awesome to query specific package versions, or even individual package file MD5s from an SQL interface to check system exposure when new exploits come down the pipe.
"It provides atomic upgrades and rollbacks, side-by-side installation of multiple versions of a package, multi-user package management and easy setup of build environments."
Abstract, I know, but I don't think anyone has looked at this in any detail yet.
http://p2.berkeley.intel-research.net/papers/EvitaRacedVLDB2...
I have some planned projects which require a homoiconic relational language. I was hoping someone else could be inspired to design such a thing so that I don't have to. It looks like someone already did. I am happy to be proven wrong :)
p.s.: Oh! there's a Meijer paper behind it [2].
1: http://ql.io/ 2: http://queue.acm.org/detail.cfm?id=1961297
docker run -t -i imiell/osquery /bin/bash
root@81fbc2076e1c/# osqueryi
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
osquery - being built, with love, at Facebook
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Connected to a transient in-memory database.
Use ".open FILENAME" to reopen on a persistent database.
osquery> select * from processes;
+----------+-------------------------+-----------+-------+---------+------------+---------------+----------------+-----------+-------------+------------+--------+
| name | path | cmdline | pid | on_disk | wired_size | resident_size | phys_footprint | user_time | system_time | start_time | parent |
+----------+-------------------------+-----------+-------+---------+------------+---------------+----------------+-----------+-------------+------------+--------+
| bash | | /bin/bash | 1 | -1 | | 1764 | 18276 | 17 | 18 | 95476444 | 0 |
| osqueryi | /usr/local/bin/osqueryi | osqueryi | 19380 | 1 | | 4312 | 110652 | 225 | 327 | 96321589 | 1 |
+----------+-------------------------+-----------+-------+---------+------------+---------------+----------------+-----------+-------------+------------+--------+
osquery>https://github.com/ianmiell/shutit/blob/master/library/osque...
There are deps on thrift and rocksdb modules defined at the bottom.
Should be useful for those looking to port to other CM tools.
cf:
SELECT * FROM Win32_LogicalDisk WHERE FreeSpace < 2097152- it's cross platform and supports many *nix operating systems
- adding new tables is very well supported via a simple API: https://github.com/facebook/osquery/wiki/creating-a-new-tabl...
- several tools and utilities exist to leverage the power of SQL at scale (osqueryd is a full operating system instrumentation tool which allows you to use SQL to instrument your whole infra): https://github.com/facebook/osquery/wiki/using-osqueryd
All in all, WMI is great, no doubt about it, but osquery has a few unique features which make it a cool, interesting product that you can use all across you internal infrastructure.
Great to have a better alternative for unixes!
[1] http://en.wikipedia.org/wiki/Web-Based_Enterprise_Management...
In this case, it's about easily correlating data to pull more complex information out of specific system data sets.
The usual implementation is simple: take any host monitor (say, collectd) that can export key/value pairs from a host, or take a log stream over the network and pair it with a host monitor/log scraper to create key/value pairs. Then insert into an SQL engine while appending to a log for a historical record (or PTA/PITR/whatever, i'm not a DBA). Separately you can create a database application to query/modify the database as needed.
But we're talking like, a handful of python scripts that don't ever change except to add new search features. This seems like a big departure from the simplicity of that approach. Am I missing something?
http://www.akamai.com/dl/technical_publications/lisa_2010.pd...
HTML version of the paper here: http://www.dmst.aueb.gr/dds/pubs/conf/2014-EuroSys-PicoQL-ke... Paywalled version: http://dl.acm.org/citation.cfm?id=2592802
What I meant is that they don't even say 'works on Linux' like in 'they don't even cover all the most commonly deployed Linux distros'.
The idea that they should test on every distro in order to say that it works on Linux is a bit silly.
On the flip side, while I am more than happy to deploy new development projects to Linux hosts, and even in docker containers, I'm still stuck developing under windows, and having to run a VM to do most of my development isn't something I'm all that in favor of.
That said, Azure is about the only cloud provider that really caters to .Net deployments, that is correct and probably accounts for a lot of it.
Me employer's next generation of applications is being deployed in Linux on Azure, because of their better support for our legacy applications (some will be around for quite a while)... which is part of why I said that "yes" running in windows is important.
If Node and Mongo didn't run in windows three years ago, we wouldn't be migrating to Linux today. I introduced Node in order to improve client-side web resources in a few web projects... once that was in place, it was a natural fit for one-off scripts (importers, timed tasks, etc). From there it became the API service for search (with mongodb behind IIS/ARR). Because of that, and the stability so far, it's our next generation platform. None of that would have happened without being able to run on windows.
I think a lot of people/companies have some of these. Even if you're 90% Linux for example, and 10% Windows, its nice to have a tool that works across the board.
Since this also seems to be targeted at laptops however, I bet Windows is still a large percentage there!
From the wiki, it says it will soon be available on homebrew as well.
for my gory details.
I'll post a shutit script with the lot in tomorrow probably.
The rocksdb dep is already done:
https://github.com/ianmiell/shutit/blob/master/library/rocks...
but bed beckons. In the meantime IYI:
https://github.com/ianmiell/shutit/blob/master/library/osque...
Thrift is another dep:
https://github.com/ianmiell/shutit/blob/master/library/thrif...
A long time back I created something similar to this atop Oracle[1]. It used a Java function calling out to system functions to get similar data sets (I/O usage, memory usage, etc). It was definitely a hack, but a really pleasant one to use.
Be cool to see an foreign data wrapper for PostgreSQL[2] that exposes similar functionality. I'm guessing it'd be pretty easy to put together as you'd only need to expose the data sets themselves as set returning functions. PostgreSQL would handle the rest. Though I guess that would limit it's usefulness to servers that have PG already installed. Having it be separate like this let's you drop it on any server (looks like it's cross platform too!).
[1]: I don't remember exactly when but I think 10g had just been released.
[2]: http://www.postgresql.org/docs/9.3/static/postgres-fdw.html
Well, the announcement says its cross-platform because it runs on two flavors of Linux (Ubuntu and CentOS) and on Max OSX.
I think its interesting to see that MIG is in Go and thus cross platform "by default". It also seems to be more privacy-compliant.
osquery's SQL is sexy however.
That said I'm also wary of a single piece of software that basically give you control over absolutely everything (control everyones laptop, etc. silently and quickly. Thats the best rootkit ever. You wont even detect if its being compromised because its a trusted piece of the OS!)
There is a number of conceptual differences between MIG, GRR and OSquery. MIG does not retrieve data from endpoints, but instead focuses on answering yes/no questions.
For example: find hosts that have a file in /var/www that contains the regex '12345'. MIG will run the search on all endpoints and return the location of files and hosts that match. But it won't return the files themselves. Privacy is preserved, but investigators need to manually retrieve files if needed. MIG takes this approach to keep the search very fast (parallel AMQP runs in seconds across thousands of endpoints). Retrieving data at large scale is too expensive for Mozilla (bandwidth, storage, execution time), and we like to protect privacy.
I really like the SQL approach taken by osquery. I brainstormed something similar last year [1] but did not get around to it yet. SQL is very natural to use for a lot of IT/Security people, and it's great to be able to transfer that knowledge to search tools.
GRR and OSquery are awesome tools. And I can tell their authors are solving some really hard problems. We're competing a bit, but it's OK to have multiple tools with varying goals. And it's great for borrowing each other's good ideas ;)
[1] http://4u.1nw.eu/Presentation_workweek_20130916.pdf slides 10 & 11
It's here if anyone wants to try it: osquery-0.0.1-trusty.amd64.deb (11 MB)
https://drive.google.com/file/d/0B3ROVJqBXqYAOVNTTkhqQzNUa0k...
http://en.wikipedia.org/wiki/Be_File_System
I remember from back in the day that was one of the really cool feature of Be.
edit: reading the source, I may be totally wrong about this.
edit2: https://github.com/facebook/osquery/wiki/building-the-code
Having said all that I'm going to install it and try it out because it's new and shiny.
I guess it does buy community goodwill to throw handfuls of money off the Facebook float...
you could then deploy these frameworks on a bunch of servers and a external monitor can independently query via SQL "How u doing;"
`select p0.name, p1.name from processes as p0 join processes as p1 on p1.id = 1 where p0.id = 1;`
Should never produce two different values. This is impossible to guarantee in all cases if data is updated asynchronously by OS.
I guess my point is that a reasonable definition of consistency could be many things, including a definition that says referenced tables are read only once in the order they are encountered (this should be equivalent to a series of eg. ps and netstat commands stored in variables and then manipulated)
[1] http://multicorn.org/foreign-data-wrappers/ [2] https://github.com/Kozea/Multicorn/blob/master/python/multic...
Otherwise it sounds like you're insinuating that MySQL is somehow languishing or limping along, which is ridiculous.
MySQL is going to SELECT almost any query faster than Postgres in a single-user/desktop installation, it's going to be easier (although granted less strict) to import data, plus it has Sequel Pro.
The idea that MySQL is much faster than PG is largely a relic of the era of MyISAM, which was very fast for trivial queries, but also terrible in basically every other way.
/sarcasm/
This project currently has a proprietary C++ API, so why not standardize with JDBC (at least for those making use of the JVM)?
One idea is to make a PostgreSQL Foreign Data Wrapper (FDW) that would allow using PostgreSQL JDBC (and ODBC) drivers.
At CloudHelix, we did a Postgres FDW to OpenTSDB, which gives a time dimension as well.
That was an issue at Akamai - how to get historic as well as realtime with Akamai's Query system ([WARNING: PDF direct download] http://www.google.com/url?sa=t&rct=j&q=&esrc=s&source=web&cd...)
Interesting stuff though! Maybe a FDW could connect Postgres to osquery, which could allow joining with local tables or other FDW-accessible data.
The FDW approach to OpenTSDB looks like:
select to_timestamp(atime::float), value, hstore(regexp_split_to_array(tags, ',')) as hs from chf_realtime where i_start_time >= now() - interval'1 min' and agg = 'sum' and metric = 'df.bytes.percentused' and tags = 'host=*,mount=/|/data|/ssd' ; to_timestamp | value | hs ------------------------+-------+--------------------------------------------------------- 2014-10-30 01:15:33+00 | 84 | "host"=>"XY4.iad1", "mount"=>"/", "fstype"=>"xfs" 2014-10-30 01:16:33+00 | 84 | "host"=>"XY4.iad1", "mount"=>"/", "fstype"=>"xfs" 2014-10-30 01:15:33+00 | 9 | "host"=>"XY.iad1", "mount"=>"/data", "fstype"=>"btrfs" 2014-10-30 01:16:33+00 | 9 | "host"=>"XY.iad1", "mount"=>"/data", "fstype"=>"btrfs" 2014-10-30 01:15:33+00 | 49 | "host"=>"XY.iad1", "mount"=>"/ssd", "fstype"=>"btrfs" 2014-10-30 01:16:33+00 | 49 | "host"=>"XY.iad1", "mount"=>"/ssd", "fstype"=>"btrfs" 2014-10-30 01:14:55+00 | 63 | "host"=>"XY.iad1", "mount"=>"/", "fstype"=>"xfs" 2014-10-30 01:15:55+00 | 63 | "host"=>"XY.iad1", "mount"=>"/", "fstype"=>"xfs" 2014-10-30 01:14:55+00 | 1 | "host"=>"XY.iad1", "mount"=>"/data", "fstype"=>"btrfs" 2014-10-30 01:15:55+00 | 21 | "host"=>"XY.iad1", "mount"=>"/ssd", "fstype"=>"xfs" 2014-10-30 01:14:55+00 | 21 | "host"=>"XY.iad1", "mount"=>"/ssd", "fstype"=>"xfs" 2014-10-30 01:15:50+00 | 63 | "host"=>"XY.iad1", "mount"=>"/", "fstype"=>"xfs" 2014-10-30 01:15:50+00 | 8 | "host"=>"XY.iad1", "mount"=>"/ssd", "fstype"=>"btrfs" 2014-10-30 01:14:56+00 | 89 | "host"=>"XY.iad1", "mount"=>"/", "fstype"=>"xfs" 2014-10-30 01:15:56+00 | 89 | "host"=>"XY.iad1", "mount"=>"/", "fstype"=>"xfs" 2014-10-30 01:14:56+00 | 55 | "host"=>"XY.iad1", "mount"=>"/ssd", "fstype"=>"xfs" 2014-10-30 01:15:56+00 | 55 | "host"=>"XY.iad1", "mount"=>"/ssd", "fstype"=>"xfs" (17 rows)
edit: and after more reading, looks like the ETL could be done with osqueryd, and a plugin to push output to sqlite if desired -- even better!
Have you seen the other neat functionality about the Akamai query system? They allow (actually require) devs to create callbacks that get exposed as query tables to do debugging (vs letting devs on machines as first or second line of debugging).
It's an interesting paradigm though again the way they use it, like osquery, is only real-time and using the psql/OpenTSDB method allows for history as well as real-time.
osquery (and Akamai Query's) syntax is definitely simpler. We do plan to add a logical schema to map tags in OpenTSDB to logical columns to help with that.
Today we know via Category Theory that tables are Turing complete and are actually quite synonymous with CT itself.
In other words, thinking of computation in terms of tables with rows and columns and relationships between tables is an interesting and promising (given CT) approach to computing that has been discussed in the past but then left largely unexplored.
There may be some application of concepts from category theory to relational algebra, of course, since category theory is incredibly abstract and has some tangential relation with most everything. But it seems a bit too glib to say that tables are "synonymous with CT itself".
Evidently a mathematician at MIT, David Spivak[0], has done some work on Databases as Categories. I found a presentation he gave to Galois[1] and a summary of the talk by E.Z. Yang.[2].
That said, I think "tables synonymous with CT itself" is a bit strong. Rather, the argument is that database schemas are categories if you model it appropriately: tables are objects in the category, and foreign keys are the morphisms/arrows).
[0] http://math.mit.edu/~dspivak/ [1] http://math.mit.edu/~dspivak/informatics/talks/galois.pdf [2] http://blog.ezyang.com/2010/06/databases-are-categories/
Also see this: http://code.galois.com/talk/2010/10-06-spivak.pdf
I mean how do they think people will write and rewrite programs using sql when files are already here?