When you have the tools you have the power.
10 karma · joined September 22, 2009
When you have the tools you have the power.
The author of the 6b article kindly provided more information about the time it took for the migration:
http://news.ycombinator.com/item?id=3621561
A couple of limitations of the Percona tool are related to changing data in tables with ON DELETE and ON UPDATE foreign key constraints and possible lock-ups of tables. It is a bi-directional replication tool, so it has to deal with the master-master replication case and as such does not guarantee data consistency.
Still the tool is helpful in many cases and congrats to Percona for developing it. Replication safety is a non-trivial problem in general.
What they really needed is ChronicDB http://chronicdb.com
http://chronicdb.com/blogs/nosql_is_technologically_inferior....
Flexibility of schema definition and flexibility of schema change are two different things. Defining schemas only involves data. But changing schemas involves not just data, but code too.
ALTER TABLE performance will eventually improve. PostgreSQL 9.1 lifts-off the lock-up limitation.
But what a performant ALTER TABLE will not improve is its intrusiveness to applications. When you change the schema you break the app.
That's the problem.
Also, what do BCC and AR refer to?
Exactly! Try versioning data with ChronicDB:
What's worse than migrations? Being unable to turn an application off, since it uses the old schema.
With ChronicDB we support indefinite backward compatibility. Unlike per-record versioning tricks, application code does not need to be aware of migration code.
$ chd version mydb
$ chd revert -d <txn_id> mydb
Neither replication nor backups protect from accidentally deleting data. And restoring from binary logs requires downtime.If the problem is complex, you want to model relations, not ignore them.
(But clearly not in this hairy query.)
> "I find migrations painful and unnecessary."
A schema-less model neither makes a migration less painful nor eliminates it.
In MongoDB, what did you do when the data model changed?
I second that: schema-less is misunderstood.
There's a difference between flexibility of schema definition and flexibility of schema change[1].
Flexibility of schema change, which NoSQL does not solve, is increasingly more important. Not just for large data stores but also for the data development process and release process. To avoid playing the suboptimal schema-change game both the code and the data need to be updated together. Or at least be given the illusion that they have[2].
A probably obvious question most developers must have asked by now is: if we've built great tools to version source changes, how come we haven't built great tools to version data changes?
[1] - http://chronicdb.com/blogs/nosql_is_technologically_inferior...
It sounds that SQL itself wasn't the problem. Were you looking for versioning alternatives in SQL that weren't up to par?
We have been working on building what we hope are the procedures for this with ChronicDB (http://chronincdb.com). But it turned out harder than it seems, and we are not sure it will quite work out. We'd welcome feedback.
But one has to wonder, what happens when you need to upgrade the app? Shutting down the process and destroying the memory image doesn't seem like the best option:
- First, it disrupts connected applications since the process is killed, introducing downtime.
- Second, when starting up again in say version 2, the data that will be loaded in memory still needs to be transformed in the format expected by version 2. This transformation can take time on large data, introducing further downtime.
The challenge would be to eliminate this downtime by combining a solution for both client disruption and state transfer. A data abstraction like a database using SQL can simplify such a solution.
> How do you guarantee that all of the apps that touch that data use the current version of said code?
An approach that may be worth considering is to not require all apps to use the current version: allow multiple versions. For some cases this would work, say if the semantics of the newer version are backwards compatible with the semantics of the older version. If the data semantics are preservable, transforming a schema could happen while each data access request to the schema is actively transformed.
But it clearly wouldn't work in all cases. More work to handle that would be needed.
> Code normalization is as important as data normalization.
True, this is the near show-stopper really. In that case, the best one can hope for is preparing the state of the new version (new data in new schema) and carefully coordinating a quick restart of the old version for the new version.
I would love to hear your thoughts on this. We have been working towards that direction with ChronicDB (http://chronicdb.com) and would welcome feedback.
1. Updates should not only be applied in sequence.
It is better to produce a binary diff between any two versions, and apply only that (one) binary diff. The reason for this isn't efficiency, but semantics. Updates not only fix things, but break things. Meaning, updates corrupt application state (data), both in-memory and on-disk. It can be disastrous to apply an intermediate update that removes state, only to realize that a future version reversed the semantics and needs to use that state (which was available, but is now gone).
Peserving backward compatibility is important, which means the ability to skip some version updates is necessary. To the extent possible, reversing updates is important too.
2. The ideal update system should apply updates live, not offline.
With a model that accounts for updating the entire state of an application, updating live is possible. The reason most updates are not applied live yet is that the model is not descriptive enough to change the entire state of the running application.
Notable state that should be updated, but often isn't, is continuations and the stack. This is why GUI applications need to be shut down to update.
Scheme's call/cc (call-with-current-continuation) solved making changes to continuations and stack state decades ago better than Erlang. Erlang cannot force stacks unroll or continue from arbitrary points.
3. Updates must be produced with source code and programmer input.
Updates should not be produced with binaries as input.
The reason is the need to account for application semantics, which binaries do not expose in the detail source code does. Although automated, sophisticated semantic-diffing based on control-flow can be developed, it is sometimes inconclusive whether an update will break things.
4. It is necessary for programmers to provide live update guidance.
In the cases where producing provably safe dynamic updates is not possible, it is input from the programmer that can clear any conservatism of the safety certification process.
Tools are needed for programmers to reason about the semantic safety of their live updates, integrated in the development process. Including tools that help transform application state between versions.
If your supervisor's background does not intimitate you, and you don't desperately want to become like them, then find someone else.
In my opinion, women possess a unique perspective of what people may want. With 99% of founders being male, it sounds that women founders are likely to introduce unimaginable ideas.
I may have pushed her a bit to apply, and she may have needed the push, but she had no trouble filling a well thought-out application.
(I am dying to see the response)
But it takes over with dependencies of newer versions than necessary.
Requires: python
The automatic mechanism in place is to run ldd, identify the list of libraries used by the executables distributed, and mark those libraries as dependencies. Since the package is build on your newer, Debian 5 package, it requires (automatically) the newer libfoo (like python-2.7), while it would have worked just fine with the older libfoo (python-2.5).
The problem is that about 80% of packages in distributions could be compiled against older versions of dependencies, but never do. Be default, distributions mark in the packages the dependency of their currently installed libraries, ignoring the fact that the software could have been compiled against an older version of the library too.
There's no mechanism for identifying such minimum possible dependencies.
Doing so would have made your software immune to the broken Python installations.
They discriminate against you just because you are solo.