X11 Universal File Opener and XDG Mess
vermaden.wordpress.com
vermaden.wordpress.com
$ xdg-mime query filetype file.docx
application/vnd.openxmlformats-officedocument.wordprocessingml.document
$ xdg-mime query filetype file.xls
application/vnd.ms-excelMaybe you have more recent xdg-utils version?
I run 1.1.3 here.
Maybe what's missing is the libreoffice package? It seems like all these types are provided by /usr/share/mime/packages/libreoffice.xml on my system (which is read by update-mime-types).
Arguably this is a common enough format that you shouldn't really need libreoffice installed to have it, though I would think any other office program would have it. I think /etc/mime.types serves as the ultimate fallback, though it doesn't have entries for libreoffice file extensions on my system.
The Garbage In, Garbage Out principal applies here.
For anyone interested about how the plumbing, here are the specs:
shared-mime-info-spec[0] is for mapping glob patterns, magic numbers, string fragments, etc. to MIME types; and desktop-entry-spec[1] is for defining what an application is, its icon, how to launch it, what MIME types it handles, etc.
[0] https://freedesktop.org/wiki/Specifications/shared-mime-info...
[1] https://www.freedesktop.org/wiki/Specifications/desktop-entr...
On my (Debian) system, the following packages provide the definitions that map *.docx files to the appropriate MIME type:
$ grep -l '\*.docx' /usr/share/mime/packages/*.xml | xargs dpkg -S
shared-mime-info: /usr/share/mime/packages/freedesktop.org.xml
libreoffice-common: /usr/share/mime/packages/libreoffice.xml
... and subsequently: $ xdg-mime query filetype Downloads/something.docx
application/vnd.openxmlformats-officedocument.wordprocessingml.document
$ xdg-mime query default application/vnd.openxmlformats-officedocument.wordprocessingml.document
libreoffice-writer.desktop*Storing the MIME type separately is definitely a good thing. This isn't the only way you do it. You start by checking the MIME type and then fall back to one of our crude detection mechanisms, like guessing what the file type is based on extension.
We already store plenty of metadata with files--permissions and timestamps for starters. The metadata is sometimes lost when you transfer to another system. That's OK... you can repair the metadata.
Given that this metadata is already present in the HTTP protocol, it's present in files embedded in email messages, and given that extended attributes are supported by many different filesystems, I think this solution is completely workable. You could make the argument that since we spend so much time interacting with files over HTTP, that having a MIME type in the metadata is actually the common way to do things.
So you need something to disambiguate.
I can understand that if you've never personally run into this problem, it may seem unimportant. It is extremely frustrating to work with two different files with the same extension on a regular basis.
Care to share some examples? I've very rarely had this problem when some specific program chose to use something like .dat as the extension, but I doubt someone making this decision would have picked a sensible mime type either.
Then you mention http already solves this problem. Great. Now you only need to fix the remaining five trillion protocols used to transfer files that don't.
If you're old enough you might know the unix haters handbook. It also made fun of this problem and smarty-pants explained how the next version of macos will strictly store any metadata, such as mime type, separately in the file system meta data and not in the file(name) anymore. I think they even gave the size of an image as an example. Thirty years later and this is still what macos does; Jpeg files still end in jpg and contain exif data.
Metadata goes in the file, file type is the same. That's the only sane way to get this to work in an interoperable way. It's the least common denominator, since you always have these two, no matter the environment. For anything else, the ship has sailed.
MIME types are flawed in a world that only stores file extensions. On Android, you can only register apps as handling specific file MIME types, not file extensions. And for a non-mainstream file extension like .spc, every file browser app synthesizes a different MIME type (which may be blank in some cases).
There are flaws, like PS1 sound dumps often coming with wrong volumes and echo, but then again most recorded versions of those songs were also generated through inaccurate emulation, then fed through lossy codecs, and some have playback/recording glitches.
Can't extensions be long? We're not stuck with 8.3 anymore; AFAIK, "foo.mycustomprogramnamehere" is a 100% valid name.
The number of video games that have a different .tex for textures ...
The point is that if content types were actually stored separately, it would be possible to not rely on detection. Furthermore, it would also be possible to know about a file's content type without it being present on the user's system.
You see, in xdg, there is a core database of content types and how to detect them, and more can be installed with programs if they care to do that. But if a file's content type is not in that db, it's unknown to your system entirely. So your system doesn't know that it doesn't know. It's very inconvenient.
I might write a proper blog post about this if I get some time to get my thoughts together. I worked for a year or so on this and with the xdg spec, it's all implemented in pretty frustrating ways.
So, what makes it a better hint than file extensions? If you see an extension you don't recognize, why is that any worse of a signal than a mime type you don't recognize?
Does it apply to dynamically generated content? No. Does it apply to key-value stores like S3 and GCS? No.
Could you give an example where one would want to choose the MIME type of some data, from a selection of options?
You will sometimes need to change the MIME type of something. For example, if you use a text editor to create an HTML file, you'll want to change the MIME type to text/html. It's also nice to be able to mark *.h files as containing C or C++ code.
I have a cat. I can derive that it is a cat when the situation requires it. I don’t need to have a metadata label attached to it, in my own handwriting, saying “cat”.
Of course, for some files it is possible. But we need the metadata for the other files.
I entirely disagree. If I have an empty .json or empty .bin or empty .txt I expect a different behaviour in the software that I am opening the file with, even if in all cases the actual data is zero bytes.
More generally file formats can be polyglots: https://github.com/Polydet/polyglot-database
And, if it happens that the file goes through such a system that doesn't understand mimetypes, you're only back to the status quo of having to sniff the type from hints like extension or content.
Would you say that all the common desktop filesystems, flash memory filesystems, optical filesystems, archive formats, etc. today all aren't robust since they don't do this?
> ... you're only back to the status quo of having to sniff the type from hints like extension or content.
Combining multiple methods does solve that problem, but doesn't it kind of negate the benefits too?
Also, I still think there are some issues. For example, it would need the creation of new mime type changing UI, which might be a hard thing to teach laypeople (they already barely understand extensions). And what do you do if the internal mime type disagrees with the extension? Aside from just being confusing, that could be purposely abused by malware.
I mean, yeah. They all guess at the mimetype, all the time. They all fall on their face fairly often in that regard.
> Combining multiple methods does solve that problem, but doesn't it kind of negate the benefits too?
No? A system that just stores a two-tuple of (type, data) and doesn't have to guess — when it knows the type — is strictly better than a system that always guesses. Where it integrates with other systems that understand how to type the incoming data, it would work flawlessly, every time. Where it integrates with systems that send untyped bytes, there is again no choice but to guess: the data simply isn't there.
> For example, it would need the creation of new mime type changing UI, which might be a hard thing to teach laypeople (they already barely understand extensions).
Yes, I agree. But people are only going to hit trouble where the mimetype isn't known and the sniffing fails, which is the same issue they'd hit today. I'd argue setting a mimetype has a better shot at being a good UX than trying to get them to set the file extension ever will though.
E.g., Right click → Set file type → prompt with different types (use friendly names, if at all possible). E.g.,
This file looks like it is probably one of the following:
JPEG image
(can be opened with GIMP, Photoshop.)
PDF document
(can be opened with Adobe Acrobat.)
> See all options
(accordion dropdown to show all options)
Yes, that still requires the user to know the differences between a JPEG & and PNG and to a layman, that's considerably sub-optimal. But we're at the point where the system didn't record the mimetype, and can't guess it correctly, so there's not much left, really: some human has to make a call.A form of this exists today in most systems, with an "Open with" context menu, and generally with an option to "always open files of this type with this application". (But that's more about the binding between the mimetype — still determined currently through sniffing or extension — and the handler for that mimetype.)
But ideally, if the interfaces were there to allow the process to be deterministic & obvious when the data is known, things could or would grow to adopt those interfaces. (Though it'll likely be a looong time.) Unix screwed up, in that regard, in that it set us down the path of "everything is untyped bytes", vs. "everything is strongly typed bytes". With the latter, the system can start figuring out correct or incorrect actions. With the former, it simply lacks the information to make a decision.
> Aside from just being confusing, that could be purposely abused by malware.
OS X essentially just marks, in the metadata of the file, that it came from the Internet. It could keep doing that. Whatever process "this file could harm your machine" happens with in the browser today could keep on happening that same way.
Granted, there is the possibility of "foo.jpg" being sent with a "application/executable" mimetype. That's a real concern. I think this comes down to the system being clear about what you're dealing with, and the consequences of actions. This problem already exists today: people have crafted ".pdf" files that are valid executables. Having better security controls on apps (not having desktop apps run with the same privilege as the user) would help (limits the damage).
We've learned, repeatedly, that strong typing results in most robust systems. I don't think the answer is any difference with the bag of bytes a file comes with: knowing the type is better than guessing it. We've also learned, I think, that sniffing almost always leads to loopholes…
(I'm not a fan of the "or guess" bit of the proposed idea in my comment; I think a simple "the file carries the mime and that's that" would be better, but I suspect that the roughness of integrating with legacy code that can't communicate what type of data it is reading/writing would hamper that.)
The file type determine which applications could be used to open a file (by drag-and-drop onto the application, or via the application's Open dialog box).
Applications could register to open a type, I think it would default to the creator.
Of course, that wound up with things like resource forks, and that didn't really play well with anyone else's file system.
What was stored in a resource fork (of each application) was the declaration “I am [four byte creator code]; I can open these [four byte file types]. That’s the information that the Finder built its databases from (https://www.folklore.org/StoryView.py?project=Macintosh&stor...)
That easily could have been moved into the file proper (as AppleSingle (https://en.wikipedia.org/wiki/AppleSingle_and_AppleDouble_fo...) did, but another (and IMO better for this purpose) option would have been to move that info into a special code segment in the data fork (in Unix parlor; old-style Mac OS programs stored their code in code segments in the resource fork)
Also, I think Mac OS still uses https://en.wikipedia.org/wiki/Uniform_Type_Identifier, which are an improvement over file types and creator codes (you can, for example, express the fact that every html file is a text file with it)
Never had much luck with xdg-open
Will definitely dig into it to either use it directly or at least try to implement it in mine see.sh.
Also IMHO you should add short 'youtu.be' regex handler :)
Regards.
pdftotext seems reasonable and looks to be part of the Debian lessfilter.
#!/bin/sh
echo -n "$1" | xclip -selection primary; notify-send "URL Copied" "$1"It is for this reason that `echo` is generally frowned upon for programmatic input, and `printf %s "$1"` is recommended; as this is guaranteed to not interpret, and copy dash-leading strings to the output verbātim.
Modern shells also typically have `echo` as a builtin and can have quite different behavior.
https://stackoverflow.com/questions/33784892/bash-vs-dash-be...
Any topics You would like to be covered?
it shows a dialog with all the associated application to the user to choose from
I use BROWSER environment variable for that.
It seems to even override XDG settings.
% xdg-settings set default-web-browser firefox.sh.desktop
xdg-settings: $BROWSER is set and can't be changed with xdg-settings
% echo ${BROWSER}
firefox
Hope that helps.EDIT: You also got me nice idea - to add http(s):// and ftp:// support for mine see.sh opener :)
~/.local/share/applications/mimeapps.list
~/.config/mimeapps.list
In here you find the MIME type pointing at a .desktop file (app launcher), it's a matter of changing that. Mine for example related to using Firefox: [Default Applications]
x-scheme-handler/http=firefox.desktop
x-scheme-handler/https=firefox.desktop
Find the .desktop file you want to launch and just update that bad boy. Most DEs have a GUI tool to manage this for you without having to resort to manual editing...[0] https://www.freedesktop.org/wiki/Specifications/desktop-entr...
On my (Debian) system, firefox and chromium ship their .desktop files.