Wuffs: Wrangling Untrusted File Formats Safely
github.com
github.com
So generally Wuffs is great and you should use it to decode your PNGs. There are some downsides: not all of the obscure bit depths and formats that PNG supports are loaded as-is, some are converted to more standard formats.
Also the Wuffs documentation is a bit hard to understand. It's a litle bit of a mission getting PNG decoding working. You can see my code for that here though: https://github.com/glaretechnologies/glare-core/blob/2c7174c...
Also, it has the funniest testimonials.
The biggest bottleneck in PNG decoding is zlib, which is not part of libpng. There are faster inflate implementations, but nowhere near 5x.
The second slowest thing is unfiltering, but it takes only 10-20% of the decoding time, so even lightspeed implementation would make little difference.
There is possibility of a 10x difference when encoding, but that's not due to libpng being slow, but because it's possible to apply worse compression and there are dedicated crappy-but-veryfast encoders.
Wuffs the Language - https://news.ycombinator.com/item?id=26731305 - April 2021 (75 comments)
Wuffs’ PNG image decoder - https://news.ycombinator.com/item?id=26714831 - April 2021 (138 comments)
Today C makes most sense given the WUFFS language is still in flux.
[Edited to fix a serious typo]
So basically higher performance.
FFI nominally has a runtime and compile time cost - whether that matters for you in particular will depend on your needs, but being able to publish a very simple crate without a build.rs to manage can have an attraction.
It's common for C libraries that do get wrapped today (e.g. openssl) to have a two phase wrapping, a -sys crate which turns the C into Rust C FFI and then another crate to turn the Rust C FFI into something actually palatable to ordinary people.
> that can then be shipped like normal C, so you don't get the ecosystem friction like with ex. Rust.
Emitting Rust doesn’t help with this.
For debugging i believe you can generate your own source maps and use gdb as a backend to talk with your custom debugger.
My understanding is that compiling unsafe C to WASM and back would also guarantee safety with respect to buffer overflows, integer arithmetic overflows and null pointer dereferences.
It’s nice not annotating code to explicitly prove invariants to the compiler like you would in say Wuffs or Rust, but I suppose that’s what limits performance.
What seems nice about wuffs is that it has no side effects and a clear project scope. Deserialization is so riddled with severe issues that it does kind of warrant its own DSL. OTOH, some legacy formats will probably never be ported.
Not fatal, but perhaps annoying.
But the tradeoff is that you need to rewrite your code for Wuffs, while WasmBoxC can sandbox anything that compiles to wasm and prevent it from corrupting the outside, including existing code in C, C++, Zig, unsafe Rust, etc. etc.
I would be glad to have a try but the language specification and documentation is almost non existent.
I understand PDF has a bunch of limbs, but I always assumed the JS stuff was at least separate from the parsing. (I am familiar with the PDF format at a lower level but I never touched any of the weird features.)
I work a lot in OpenSCAD, and had a need to design some custom graph paper. So I found the subset of SVG which was similar to OpenSCAD. :)
It's annoying you can't just "flatten" or "bake" such an svg like yours into one composed entirely of elements (unless one exists?)
Chrome even includes a --dump-dom flag you can use to do this on the command line, although I haven't tested it with an SVG.
Note this does not prevent unscrupulous companies abusing dominant market positions to voluntarily embed machine and serial hash watermarks.
To be clear: formats like pdf, ps, webp, svg, and tiff are so badly implemented in some ecosystems... they can't _ever_ be assumed safe input formats. Thus, at some point people need to spin up an actual VM to transcode a "web" version, and scrub each stage of the rendering pipeline like a virus or header injection is already present.
"I never play where nice things are, and don't break things" (Eliza Mowry Blven, The Humanitarian Review, Volume 3, March, 1905)
Cheers =3
I've yet to find a better method than Honeypots to sustainably mitigate the complex leaky dependency mess on traditional architectures. It has been my experience that "all software is terrible, but some of it is useful".
It may just be my bias, but I see code smell getting worse in recent decades...
Have a nice day, =3
And walking a binary object store to ban problem users is not always necessary... depending what you are doing.
Most other approaches makes the same predictable assumptions:
https://en.wikipedia.org/wiki/List_of_cognitive_biases
Despite popular belief, shitty design does not usually get better in another language. Rather, people just feel more confident it isn't shit anymore.
I have yet to see evidence to the contrary. =3
Have you been imagining that sandboxes are some sort of fairy dust we just stumbled onto one day, supernatural in nature and not, in fact, just software written by people you're hoping are competent and haven't left any holes?
Or put another way, the available attack surface of a bare-minimum fixed environment is much easier to auto-audit, than a pile of daily permuted binaries and self-delusion approach. i.e. if it fails to behave in an expected way, or is modified in any way... the host audit process doesn't have to care why or how it is broken to maintain a service queue as the guest is culled.
Perhaps I am wrong about exchanging 15% of raw performance for reliability, but things can get complicated with licenses and multiple OS specific platforms.
You seem to be getting emotional about this subject, presenting secondary and tertiary straw-man arguments. So I'm going to go eat some Cheese Goldfish crackers... and just agree that your beliefs are interesting.
Have a fantastic weekend... =3
If you're Matt Godbolt the benefits of sandboxing outweigh the cost because Matt is interested in general purpose software. But WUFFS isn't for that, as its name says it's interested in doing one particular task well.
In this deliberately limited domain, WUFFS gets to sidestep Rice's theorem altogether and just prove the software meets the semantic requirements [technically you do the proving, WUFFS just checks your work].
I hope you enjoyed your goldfish crackers but I urge you to use the right tool for the job.
The design in question currently only processes around 1.8M large image files a day, and does not require additional work/re-implementations to support the dozens of questionable user file-formats. i.e. the plain old ImageMagick lib does most of the heavy lifting at the end.
Would I trust such a solution for something like a native client side web-browser etc... absolutely not... but for the core-bound instance overhead, the resource cost was acceptable for almost a decade of uptime on those system instances.
Use-cases are funny like that, as there is no perfect solution... but rather a tradeoff of what features get the system functional and reliable. Part of that is admitting integration of 3rd party dependencies is a long-term liability, and domain specific languages almost always fade into obscurity.
Cheers, =3
Hardly a panacea for fundamentally bad designs that go back decades.
Ever seen a web-server written in postscript? Its worth a look just for the laughs.
Good luck out there =)
The clever idea is to have you the programmer in effect write a proof that your code has the desired semantic properties as part of the programming activity and so then the WUFFS transpiler is merely checking that the proof is correct.
This leverages your understanding of what you were trying to do.
try to ever read any code for PDFs and see all the horrors.
Google gave up and just bought the code from foxit.
They just bought foxit code to save years of development when they wanted to ship PDF reader in Chrome.
Your comment about "the only viewer that is semi-correct" is also wildly off the mark.
Parsing correctly written PDF files is hard but multiple engines can do it correctly.
Parsing real life PDFs is much harder then correctly implementing PDF spec because lots of PDFs are just broken. They generators create invalid PDF files and then PDF readers have to spend heroic efforts to somehow make sense of this brokenness. Adobe does it better than most because... well it would be embarrassing if they didn't. They invented the format, they make money from their tools, they were doing it the longest, they have the largest archive of broken PDFs for testing etc. It's hard to expect that e.g. an open-source project with one or two developers can match that.
I work on SumatraPDF so I know.
edit: it seems Summatra doesn't support PDF forms? Either AcroForms or XFA forms?
[0]: https://github.com/WebAssembly/wabt/tree/44837a7236e85c048de...
[1]: https://github.com/WebAssembly/wabt/issues/2289#issuecomment...
Even if I told you some of the numbers I’ve seen in my experiments and usage, it wouldn’t be wise to trust them or let them taint your opinion.
Please, let the end brackets should be enough.