How to Corrupt an SQLite Database File
sqlite.org
sqlite.org
"The author disclaims copyright to this source code. In place of a legal notice, here is a blessing: May you do good and not evil. May you find forgiveness for yourself and forgive others. May you share freely, never taking more than you give."
Unlike the JSON license clause "The Software shall be used for Good, not Evil", the SQLite blessing doesn't put any restriction on your "freedom to run the program, for any purpose (freedom 0)" or on any of the other freedoms listed by the FSF.
It's no different from putting this in a library:
# This work is dedicated to the public domain.
# I hope you enjoy using it!
No one could argue that if you didn't enjoy using the library (maybe the API is a confusing mess) you would somehow be in violation of the license.
p.s. I upvoted your comment because I think you raised an interesting point!
That's how we interpret it. The problem is that we don't know what the rights-holder has in mind, and what that entity's willing to sue over.
Lawyers can be comically risk-averse. Comical, that is, until you see the kinds of things people actually sue over.
No. That follows from the fact that the author gives up copyright. No copyright -> no license terms.
Wouldn't that be in contrast to the license?
More info: https://en.wikipedia.org/wiki/Destroyer
They also "do documentation right" and this is another accessible, clear example of that.
It's also all public domain. I've used their documentation as an example in the past when driving through documentation improvements and i'll no doubt point to them again!
Not Java, not Python, not Ruby. TCL & C.
Documentation done right for a small percentage of programmers elbow deep in neckbeards maybe. Alienating and weird for the rest of them.
Still a fantastic product ofc!
Even assuming the lack of a library, however, I still find the "but it's all in C!" argument unpersuasive. The idea that C is some deep "neckbeard" stuff is just silly. Part of understanding programming is understanding how your language of choice interacts with the lingua franca that is C, even if you don't know C yourself (which you really should, even if you're working in Ruby or Python or Java on a regular basis). So learn what you need to learn before using SQLite if you have to write your own bindings. That is not a big deal. If it "alienates" you, learn more. It's all out there for you.
Hipp himself admits SQLite wouldn't have been possible without Tcl.
If you are a noob, you probably shouldn't be delving directly into the SQLite documentation. You should be reading the documentation that explains the interface for your specific programming language. Python has sufficient documentation for SQLite.
Besides, most languages have a predefined and fixed interface for opening databases of any kind and running queries on them. The only thing that differs is the connection string or the constructor of the database/connection/reader-object. For example in python this would be PEP249.
This is in stark contrast to something like MongoDB, which barely parses the request before reporting a success, and doesn't make any guarantees of when or even if it will ever save the data (though it usually does).
My understanding is that the kernel (for Windows and Linux values of "kernel") will never lie to you. They will accurately report what the storage drivers told it, and the lying occurs at that level.
Storage drivers seem to be a big source of corruption for MSSQL [1]:
> The most common cause of database corruption (more than 95% of all corruption cases) that we in PSS encounter turn out to be caused by a platform issue, which is a layer below the SQL Server. The most common individual cause is a 3rd party driver or firmware bug.
And as well
> My guess would be that hard-drives have explicit instruction(s)
They do, but the problem is that the storage drivers lie about what actually happened. E.g. Basically they implement write caching for the "flush the write cache" instruction.
[1] http://blogs.msdn.com/b/suhde/archive/2009/04/08/introductio...
But of course doing that is up to the user. Using write-back caching with long sync interval and disabling fsync. It's really nice, gives better than SSD performance with regular HDD. Until you shut it down uncleanly, then you're screwed. But of course any sane person would use this only for temporary or other really non-important data which can be regenerated or lost without problems in such situation.
I'm using such configuration with ERP,BI/ETL (Extract, Transform, Load) tasks. When I start the task, I anyway drop and recreate any tables required for the task. SO I don't really mind if data gets corrupted. That's just life. Doing safe commits would make task very slow.
Only good question is how to balance smartly, in application data / caching, database engine caching and file system caching. In cases where database runs on same server as the processing application.
As someone mentioned, this is contrast to say how MongoDB was shipping not too long ago. They had turned of any acknowledgement for writes. So doing a db write was more like a throw over the fence and pray operation. So it didn't even try to be durable, by design.
Now, computers are built with subsystems from a half dozen different companies, none of which can be bothered to fully document their product's behavior. Product life cycles are so fast it isn't worth working out what a device actually does, because by the time you do you either won't be able to buy it anymore, or someone's new iteration will be 20% cheaper/faster/bigger and you will get killed in the market if you don't switch to it.
␄
PS: We were also happy to have a whole MIP of Vax equivalent (or identical) processing power, and I was slightly reknowned for my ability to code cleverly enough to get 42 disk IO/second out of our main disk drive, so I'd rather not go back, think you very much. I'll get along as best I can with the miraculous and inexpensive rubbish we build our systems from today.
So, if the drives didn't lie about a flush that problem is solved. But if it's on disk at time t, that's no guarantee that you can read it back at time t+1. The drives can physically fail. These "days of old" are before my time, but I really doubt they had magic disks that never physically degraded.
Android automatically deletes corrupt SQLite files!
As it turned out, I was able to fix the problem by dumping the database and creating a new one. Oddly enough, Adobe's repair function apparently didn't try this fairly straightforward (once you knew to do it) repair process.
I ran a forum on SQLite a while back, and things worked extremely well as long as I didn't pretend I knew better than the library.
There's no good excuse to not use PDO and there hasn't been one for years now. Before that, I've implemented a write queue, which seems redundant in retrospect, in case the file lock issue came about, but it never did even though I did hit the write queue a few times.
As for crashes, periodic snapshots of the db and journals (that's very important) is usually the best way to avoid recovery issues.
* Read
* Write
* Callback (to sort out paradoxes when two or more devices update the store independently offline and need to merge)
It's possible to do this with CoreData, but poorly, with a high burden on the developer to learn the entirety of Apple's APIs and no way to alert the user as to what it is doing under the hood, which causes the app to hang for minutes or even forever until the managed object context says it's ready.
I think where they perhaps missed the boat is that nearly every app that needs to synchronize across devices needs this, but I'm having trouble finding a simple SAAS plan. When I was younger I was interested in hosting my own database but now I'm just not. I want to pay a few bucks and have someone else do it.
So on that note a possible startup idea is to host couchbase and charge for it. I think the core of the problem is that app sales are one-time, so a subscription model may not be appropriate. But the bandwidth will typically be so small that it won't matter. So that puts the total value per user maybe in the 25 cent range. How many million new users per year would it take to gross a million dollars.. I can see their dilemma. But, I think there is something to this. Whoever pulls it off could be the next dropbox but for databases. If this all clicks for someone, look me up!