We Need A Standard Layered Image Format
shapeof.com
shapeof.com
We have a Standard Layered Image Format. HURRAY. NEXT PROBLEM, PLEASE.
Open this test image if you don't believe me: https://dl.dropboxusercontent.com/u/824493/layertest.tif
https://sites.google.com/site/elsamuko/new.tif
tiffinfo of this image:
.PDF .SVG .DXF/.DWG
All of these are open standards or basically open standards which are already widely implemented and allow the use cases mentioned (vector and raster combinations and layers) and also many of the ones you have thrown in.
Of course the real problem is that they are far too complex for most people's needs so why bother.
People really don't really often need to share the 'layers' of images combined with vector data (usually you use the layers yourself then just share the final product - a flattened JPG or similar), however if they do the above formats are there for them.
A good example of extending TIFF with tags is Pixar's PhotoRealistic RenderMan. It uses a TIFF flavor as a base for its mip-mapped textures. Such a TIFF contains a pyramid of layers of the image at different pre-downsampled resolutions and a bunch of custom tags, the renderer interprets when it reads such a file.
However, I can use 'Open As...' in Photoshop and open this as a TIFF. I will simply get the highest resolution layer from the file. That's a simple and battle-proven example of why TIFF is indeed a great answer to the problem.
has photoshop plugin aswell. No real support in the gimp though (cinepaint possibly has it though.)
It also has the advantage that if you extract the zip you can now access whatever pieces are inside as separate files even by tools that don't understand the containing format. (eg a bunch of jpg/png inside the zip are perfectly viewable in any image viewer). Trying to do the same with SQLite is problematic as you have to write queries to grab the blobs and then dump them to disk.
Where SQLite as a format does shine is if you need to keep historical information such as an undo/redo log. That makes it good for an application format, but not for an interchange format.
A more relevant example: this is how .docx files work, for very much the same reasons (one XML file at the root, and then a bunch of referenced resources.)
.zip seems to be the universal 'combine some stuff in a single file' format, probably because it is old enough to be guaranteed patent free (OTOH: http://broadcast.oreilly.com/2010/06/is-zip-in-the-public-do... and http://www.pkware.com/support/zip-app-note. http://www.pkware.com/documents/casestudies/APPNOTE.TXT only mentions encryption and patching as patented)
One could even imagine padding files in the archive so there's some leeway in growing the compressed files, it'd likely depend on the format expression.
Any performance benefits sqlite might have are easily outweighed by the ease of interacting with the zip-based format.
Edit: previously I had balked at this claim, but I'm wrong. While I can't find any official notice of its death, and it's still in use in Scene7 and somewhat supported in the open-sourced version of Flex, trawling message boards does indicate Adobe shying away from it in favor of SVG[1][2] and it's not supported in CS6 without an extension.
Still, his problem with FXG (or a format like it) seems superficial. Yes, FXG defines links to other assets instead of combining them into one file, but to me, that's a feature, not a problem: I can them programmatically swap out assets easily by just manipulating the references within the file. It makes it really useful for generating customized assets on demand. In practice, it's like HTML for layered documents. We don't complain that HTML is broken because it links to images instead of including them, do we?
If you don't like that you have to send a folder: zip it up and give it a custom extension if need be, a la browser extensions and Microsoft Office formats.
[1]: http://mail-archives.apache.org/mod_mbox/flex-dev/201303.mbo... [2]: http://mail-archives.apache.org/mod_mbox/flex-dev/201303.mbo...
So to transmit web pages? Sure, we use HTML as an interchange format. You receive everything about the page to manipulate it, including the locations to where the images and other related files are.
In other words, a layered document is not the same in kind as a JPEG or a GIF, which are flattened images. You need to be able to manipulate a layered document in the same way you could manipulate a Word document or an HTML file.
Whether you use a plain text format like XML or a SQLite database, the end result is going to be the same: in a SQLite database, you could place the component images/fonts into the database as blobs, just as you could technically place the component images/fonts in an XML file as data URIs.
But at least with XML, it's editable as plain text and doesn't require the component assets to be transmitted with the layout-defining file. I much prefer that over a format that dictates everything must be encapsulated within.
Actually, we do.
Saving a webpage and getting tons of file is a nightmare while the IE approach of lumping them all together is just so sensible it's hard to imagine a universe where this isn't standard.
Replacing resources is the special case. Sharing, moving or just storing an image isn't. Just get or write a tool that can replace the resources for you and you have completely satisfied the 0.001% as well.
As I and others have mentioned throughout the comments here, the single index file—be it XML like FXG, HTML, or otherwise—and linking to assets within it is pretty common and standard approach to complex documents. Combining them by zipping them up is a trivial addition and, again, common way to solve the "but I don't like having multiple files to transfer" problem.
That, seriously, 99% of the population will never, ever, accomplish in their entire lifetime. Instead, you get files without the resources. Instead, you get files with hard coded local links to C:\xxx.
This of course works for HTML since most "regular people" never ever has to deal with either html files or its resources (but every time they do, it is sure to be a disaster).
Images on the other hand is something that most people deal with in some way or another. It must be something that you can easily share or it will be pointless.
It would be nice if you could actually add a caption to an image - without having to "destroy" the original by converting it to a png just so that you'd be sure that the recipient could handle it.
Everything has its place. Resources work fine for html files if you are the creator, if you want to store a page for later use it is a disaster.
I love the fact that the resources are separate from the document when I write use InDesign or LateX - such a relief from having to deal with word documents etc. At the same time those solutions make it completely inaccessible to most people I know, they are great for their isolated environment but disastrous for anything else.
So, an image with resources... Might work well, terrific even, in your studio but if it is ever going to be used by the masses it is a dealbreaker.
What if PDF only used linked resources, no one would even know what PDF was by now if that were the case.
"Imagine what it would mean if HTML was an encapsulated file format: you couldn't pre-fetch or cache static assets at all, because you'd have to get everything any time the structure of the document (the HTML part proper) changes. That's not a 0.001% problem."
You could of course do all that with single encapsulated file as well...
Yes, they do. All the Microsoft Office formats (DOCX, XLSX, etc.), all the OpenOffice formats, all the iWork formats, ePub, Safari extensions, Chrome extensions, JAR files, and countless other general public-facing formats are exactly that: loose collections of files that contain an index file (or multiple index files) and are combined using zip. People by and large don't seem to have a problem dealing with them: transfers, edits, and other operations are seamless to regular users.
> What if PDF only used linked resources, no one would even know what PDF was by now if that were the case.
Many PDFs do, in fact, rely on this capability.[1] It's generally transparent to the user.
> You could of course do all that with single encapsulated file as well...
No, you couldn't. You'd have to transfer the entire file to determine what's changed. That's the the definition and principle benefit of encapsulation: one file gets transferred at once instead of many in piecemeal.
[1]: http://en.wikipedia.org/wiki/Portable_Document_Format#Conten...
In other words, a single encapsulating file... Which was my point?
The exact implementation is hardly relevant is it?
"No, you couldn't."
Of course you can. Just include hashes of the different resources within the file and only the parts of the file that contain changes needs to be downloaded. Or implement something more general along the lines of the rsync algorithm.
Taking html as an example you could also do everything transparently on the server if you wanted to. Or do everything transparently on the client instead if you wanted to keep a snapshot of each visit as a single file.
But no, I'm not arguing that we should encapsulate all web pages. But on the client side, if the user chose to save a webpage, the result should in most cases be a single encapsulating file.
"Many PDFs do, in fact, rely on this capability.[1] It's generally transparent to the user."
The keyword was only used linked resources. I think my point was, and is, rather obvious. PDF would never be where it is today if it weren't for its ability to encapsulate resources.
EDIT: My point isn't that all files should be encapsulating all the linked content. There must be a point to it. Web pages on a server, hardly beneficial... Image files? Absolutely.
Meanwhile, we'll ignore all the semantic issues around the data. What exactly _is_ a layer? Any given image file? What if some tools only handle e.g. alpha-layers? What if a layer is actually a filter kernel?
How are the layers composited? How are they ordered? Is it a linear order, or is there actually a tree of layers?
The reason PSD works so well is only partially that things are all in one file. It's also that the semantics of its contents are extremely well defined. (OK, if you're willing to consider PSD documentation "extremely well defined")
And Gus is almost completely punting on that part - define _that_ well enough, and it might make sense. Until then, it's just another VFS with a blob of assets that the receipient probably can't parse exactly as intended.
That way, you could standardize on the semantics that are already documented and familiar to other developers, but avoid the difficulties of the file format, which seem to be the issue here.
* disclaimer: I've never had to write a PSD-parser, I'm just going on hearsay.
Converting PSD's to JSON seems like a reasonable thing to do, but converting them to normalized SQL seems like a never ending shit show.
And why all the spite? Gus isn't proposing that Acorn's database schema BE the standard. He says that clearly. He simply wrote a blog post giving an example of how he has used SQLite for storing layered images. (Also, he did happen to describe how he does layer sequencing - did you read the full article?) Any final spec would of course need to define additional semantics. That wasn't the point.
But if all (or most of what) you want is a VFS, SQLite isn't necessarily a good way to make one. As has been mentioned elsewhere in the thread, it's fairly common for modern file formats to be zips of a structured filesystem; although this has upsides and downsides compared to SQLite, it seems much easier to run unzip than look up how to get SQLite to output the binary content of a table row to a file.
The problem is, at least in the case of image editors (and 3d editors), standard formats arguably stifle innovation because you have to break the standard to add anything.
Maybe you decide it would be be better if all vector colors were stored as HSVA instead of RGBA except now no one can read your files. Maybe you'd prefer to store floating point colors with values greater than 1.0 to represent light emission. Maybe the standard only defines circles but you want eclipses. Maybe the standard says circles are defined based on rounded rect with the maximum roundness but you'd like them to be based on a center point and radius. Maybe the standard doesn't support text on a path but you want text on path. Maybe the standard doesn't support linking paths so that if the text doesn't fit on one path it bleeds into another. Maybe you'd like to justify text across multiple paths but that's not in the standard. Maybe you like columns but that's in the standard. Or you want to be able to define areas to be cut into pngs, each area with a specifiable filename. Or you want some of the settings to be per animation frame but the spec never thought about animation so your SOL. Or you want layer fx. Or you want to add a new layer effect. Or you want to be able to embed a PDF as a layer. etc.. etc.. etc..
Photoshop's basic chucky format is well known. It's not that hard to write some code to read all the chunks you care about. Putting those chunks in an SQLite format will not make it any easier to deal with new chunks your code does not understand.
Indeed: that's exactly why Adobe created FXG instead of adopting SVG in the first place[1] (though they are now circling back to SVG).
[1]: http://www.mikechambers.com/blog/2008/09/30/why-adobe-chose-...
If the parsers are only looking for data relevent to them, I see no reason why a standard couldn't be developed.
I might be oversimplifying the problem...
I can use Spatialite's built-in functions to rotate, translate, scale, etc. images; and use SQL to pull out sub-volumes of the stack, edit and composite them.
My biggest stack is almost 2000 images, over 90GB uncompressed data. Working with subvolumes is pretty snappy up to a few hundred megabytes, which is good enough for my purposes. For bigger jobs it should be possible to parallelize some tasks for better performance.
Not entirely on-topic, but the takeaway is I'm thumbs-up for using SQLite to process image data.
What would be the downfalls of passing around a DB as an image file? Does the compression suck? Does the performance suck? What makes this a terrible idea?
EDIT: to clarify, by compression I'm referring to the non-lossy type - I'm assuming that within the db you'd have already processed assets.
Since SQLite aims to be a generic way for applications to store data, it'd be nice if its "open database" and "create database" functions allowed you to supply, say, a 4- or 8-byte magic value that would be stored at a fixed location in the file's header so tools could distinguish your SQLite-based file-format from everybody else's, without having to load the thing up and examine the schema.
There's no reason `file` couldn't be made to recognize special cases of SQLite databases. It certainly handles far uglier cases already.
And SQLite already provides a place to store meta-data about a table: http://www.sqlite.org/pragma.html#pragma_user_version Though the bummer is user_version is only 32 bits.
If the format catches on, they wont. For one, it would have it's own fixed file extension, and second, file(1) will be updated to perform and extra check even if something initially looks like a common SQLite database file.
If you would be willing, what's your opinion on using SQLite as an image format like this? The "When to Use" page specifically sites application file formats as a good use, but would your header byte allocation be done begrudgingly or contentedly?
It also has options for byte order to allow faster read/writes
I'm not in the Photoshop group, but I do make liberal use of SQLite in the projects that I'm involved in at Adobe.
Library table <- pixel data
Layer table <- heirarchy/duplication of library items
Attributes table <- attributes of layers
Also, are the concept of Pages accounted for by the ordering of parentless layers, or does Acorn not support the notion of pages? Would that be its own table or a modification on the Layer table.
For reasons that I don't understand, it's adoption is weak outside of the VFX industry.
Apart from VFX and post-production, OpenEXR has also become the standard file format in HDR applications.
IMHO it isn't mainstream yet because most people don't know what HDR imaging is and why current image formats are inadequate for HDR. For still image editing Photoshop is overwhelming dominant, whereas in VFX pipelines there's a lot of custom tools so you need an open format.
This would seem to be more targeted at digital content creation (and toolchain application interop) than a normal-use file format.
It is possible to clean it up of course - perhaps by vacuuming, or definitely by copying to a new file at database level. But it seems dangerous to me to start with a native format with this behavior and attempt to clean it up later.
File formats don't have behaviors. Programs do.
Lots of binary formats, including PNG, will ignore any extra junk after the end of the file. So if a change makes a PNG file smaller, there's nothing stopping a program from leaving data from the old version at the end. But that's not normal behavior for programs that write PNG, so it's not an issue.
As long as programs adopted the behavior of VACUUMing the database, this wouldn't be an issue either.
<!--...-->
</points>heh
So you could possibly argue that XML (a la SVG) is the right format to represent the vector files within that container, but it wouldn't be appropriate for the container file itself since binary data really bloats XML.
<svg viewBox="## ## ## ##">
<image xlink:href="layer1.png" x="##" y="##" width="##" height="##"/>
<image xlink:href="layer2.jpg" x="##" y="##" width="##" height="##"/>
<rect x="##" y="##" width="##" height="##" fill="red"/>
</svg>
All that's left is to standardize a way to package it. create table layers (id text, parent_id text, sequence integer, uti text, name text, data blob);
become (Vec2 used for notation's sake to make the declaration simpler--just two reals): create table layers (id text, parent_id text, sequence integer, uti text, name text, data blob, offset Vec2, bounds Vec2, rotation real);
This would make it really easy to ignore layers which have no component in the view space, or to compactly represent raster subimages."bounds" is the size of the axis-aligned bounding-box of the image. "offset" i a translation from the center of the axis-aligned bounding-box of the layer image data to the origin, and "rotation" is the rotation (in radians, CCW) of the layer. Order of application is rotation, then translation.
EDIT: Fixed missing def for "bounds".
I get that HN is awfully bare-bones in its appearance, but this I think is a bit too far.
> The behavior of the 'auto' value is user agent-dependent, but should cause a scrolling mechanism to be provided for overflowing boxes. — http://www.w3.org/TR/CSS2/visufx.html#overflow
Apple's rendering is overly pretty to the point of being useless. They dropped support for a perfectly reasonable semantic requirement (make all the content available) which was in the spec literally before their browser existed.
You can argue that websites should be written to conform to the standards-violating behavior in Apple's browser rather than actual Web standards, but we already went through that with Microsoft in the '90s and early 2000s, and I don't think you'll find a lot of people who are eager to return to those days. Browser-makers are expected to conform to the standard, not vice-versa.
I am sympathetic to the fact that this is inconvenient, because I use Apple's software too. But it's Apple that's responsible for the inconvenience, not all the web designers who are correctly following the standard.
So no, Hacker News is in the wrong here. They're relying on the assumption that the user is using a browser that renders scrollbars, and providing a sub-par experience for any browser that does not. This is somewhat analogous to a site that only renders correctly in IE.
How's that going to work for vector images?
Which is true. Except for that little "implementation" detail.
I like the idea of using SQLite as a complex-data-in-single-file format. Hadn't really thought of that for cases like this. But it does nothing to solve the problem of standardization.
I'm not proposing that an image format like this one should use IndexedDB, but if relying on SQLite means being reliant on a specific code base that has to be used by anyone that wants to read the file - then it's probably not such a good idea.
When discussing this I think one should also have in mind eg the move to web apps as well - a reliance on SQLite will make it hard for them to read the format. A standardized simple base format that can easily be supported in new languages would be preferable as I see it - but I'm no expert in the image processing area.
Or, you could just zip that whole thing up and everyone who doesn't have to care thinks it is "just one file", like .war's.
EDIT:
Found the official examination[0] on the performance of blob storage.
How is this a problem? How could is possibly be easier to open a file in sqlite than it is to just extract the file into it's folder structure with any standard decompression software?
I don't know about you, but last time checked there is no easy way to open a file contained in a sqlite database in my editor directly. I for one would much prefer the archive method than using a sqlite database.
Also, I would rather edit xml, than edit a sqlite database. XML can be edited by any text editor. I'd say text editors are a bit more universal and widespread than sqlite editors.
I wonder whether it would be more useful to just represent the whole thing as a big JSON structure with base64 blobs. Each layer can be an object in the "layers" array with the attributes as keys.
WebM is based on mkv as well.
Standards wise, open-source wise, and generally it's a good move (who wouldn't want to get a standard complex image format that would make sense to the common programmer?), but I don't think the finance department over at Adobe is going to like it very much?
http://en.wikipedia.org/wiki/OpenRaster
gimp (plugin), krita, and mypaint