In return, the kernel side API for sysfs is also a lot cleaner and allows to more-or-less expose individual variables as tuning knobs for a driver.
Of course there are edge cases, and there are e.g. some binary interfaces as well (e.g. for providing direct register access, or implementing a firmware upload interface for a device).
ABI compat issues aside, I think that implementing "a standardized [structured] record format" as suggested in the comments here is a rather bad idea, going into exactly the wrong direction by adding complexity rather than reducing it, which would definitely cause even more parsing related issues in the long run.
I'd rather have structured file than to have open 30k files (for say conntrack)
Hell, just example from the article, /proc/<PID>/stat has 52 parameters. That would be 52 opens and reads with single value per file.
> ABI compat issues aside, I think that implementing "a standardized [structured] record format" as suggested in the comments here is a rather bad idea, going into exactly the wrong direction by adding complexity rather than reducing it, which would definitely cause even more parsing related issues in the long run.
It's literally the opposite. You have to implement it once on kernel side and once in userspace vs every special format that currently needs
If you're on a system with huge numbers of connections, reading from /proc/net/tcp get extremely slow. Modern tools query connection state using netlink instead (ss vs netstat). This was done by necessity: /proc/net/tcp actually doesn't work at scale.
I agree with you, serializing and deserializing files with records is a terrible idea - it cannot be performant. JSON fixes parsing ambiguity but at a cost of being even slower. We already know it won't work.
We have already solved this for specific parts of /proc and it works great. All we have to do is finish the work and provide the rest of proc via netlink as well (or whatever else similar non-text based system for querying structured records)
CBOR could be another option: https://en.wikipedia.org/wiki/CBOR
You can't read /proc/net/tcp within a reasonable amount of time on a system with hundreds of thousands of connections. Even allocating/churning memory to store a textual representation becomes a problematic overhead.
I actually did think about this before posting my original response, and I think this is unrealistic from a practical perspective. To elaborate a bit on that:
First of, a one-size-fits-all structured format is a lot more complex than a directory with ASCII files in it that each store an integer and IMO invites itself to feature creep (i.e. more complexity).
Complexity is IMO the root cause of the issue originally discussed here (if not most bugs). The more code, the more complexity, the more bugs. In my experience, software will always have bugs, complex software more so.
There can never be a "one-true-implementation" for userspace. Because of the complexity, people will write their own ad-hoc versions. They "only need that one thing" and don't want to drag the whole library dependency in. Some people think they know better and write their own "lightweight/suckless/..." versions because the kernel one is "bloated", or "that API sucks". Some will rewrite it in their favorite programming language for whatever reason. NIH syndrome, bike shedding, ...
Then, what if a widely used implementation has a bug? Especially if it's the "one-true-library" itself? You now need to roll out a fix. Across countless Distros, embedded devices that might get maintenance updates every couple years at best, set-top boxes, network appliances, ... You'll have programs floating around that are statically linked against a specific version of the one-true-library. In the end, we have a variety of differently bugged parsers in use, simply because of the spread in versions alone. The original problem that we wanted to solve, remains.
Of course there are issues with the more simplistic approach, but in the case of e.g. sysfs, those are typically corner cases. Adding a one-size-fits-all special, structured format for everything introduces a whole lot of unneeded complexity everywhere else as well. A "one-true-format to solve all problems" that needs special library code for processing IMO introduces a whole lot more problems than it solves.
Then you could just have "load a single value" function that does the unquoting, "load K/V" function for stuff like /proc/meminfo, and "load table" for stuff like /proc/<pid>/stat. Maybe "load records" for stuff like /proc/net/nf_conntrack which is essentially list of KV pairs.
You can't really prevent that. People do funny nonsense in other self-describing data formats like JSON and XML all the time too. There's only so much you can do with a framework.
But /proc is... extremely old, and very heavily used by userspace. In practice it's never going to change.
Sure but you will get more of that if the convention is too simplistic. "one file per value" breaks really fast, just cat /proc/net/nf_conntrack or even just proc/<pid>/stats and see just how many values single entry (file/connection) has.
Doesn't need to be some ASN.1 monstrosity, could be simple conventions like "this is how key/value proc/sys file should look, this is how tabular file should look etc."
Make all escaping use same syntax, make every table separator be \t etc.
> But /proc is... extremely old, and very heavily used by userspace. In practice it's never going to change.
eh, just mount it in /proc2
All sensible ones allow you to just pass an array of parameters to command execution and not worry about spaces in them
run(f”command {arg} -v -p{opt} {target}”)
to run([“command”, arg, “-v”, “-p”, opt, target])It's an array of string parameters.
No semantics! Just an variety of customs about what it means when a parameters begins with a - or a -- or if you have -- by itself preceding some characters, how to break lists of arguments with separators, what happens when you pass the same argument twice, etc etc etc.
To choose the worst possible solution better than that, we could instead be passing in a single string with a JSON dictionary that says things like '{ "recursive": True, "force": True, "files": [ "file1", "file2", "file3" ]}'