Parsing Protobuf Definitions with Tree-sitter
relistan.com
relistan.com
Here's another approach. The AST of a .proto file is itself a protobuf. That's how the codegen plugins work. Protobuf also has a canonical mapping to JSON, so...
What you can do is use protoc to parse the .proto file, spit it out as JSON, and then process that data using your favorite pattern matching language. I wrote a [tool][1] that helps with that. For example, here's some [js code][2] that translates protobuf message definitions into "types" for use in an ORM.
[1]: https://github.com/dgoffredo/protojson
[2]: https://github.com/dgoffredo/okra/blob/master/lib/proto2type...
Also, this reads like they might not have seen the newer proto3 optional keyword, or know about the well-known wrapper types.
https://github.com/metaverse/truss/blame/master/svcdef/svcde...
Wouldn't it have been far more efficient to contribute back to protoc? Certainly posting a patch to a parser takes far fewer work than writing an alternative implementation of the same parser.
In fact, that's the whole premise of GitHub: you fork a repo, you work on a feature branch, you post a PR, anf your fork stays there.
To repeat my point: it's easier to add small features to a parser than it is to reimplement everything, and posting a PR with your small change takes no work at all.
What you said was:
> Wouldn't it have been far more efficient to contribute back to protoc?
My point was about the bureaucracy of contributing back upstream. Nonetheless, your new point, that patching and using your own fork is preferable makes sense on the surface given the ease of forking in go projects. Be interested to hear the author's reasoning.
There is no bureaucracy. Once you have your patch, you can post a PR. Click on a button, write a description, and you're done. That's it. Posting a PR is hardly the hard part of the process.
1. If I have to get back to this project several months in the future, it might be a lot of trouble to bring the fork up-to-date. The fork might not have new features, security updates, compatibility upgrades that I want. The upstream may have rewritten the code, making my fork redundant or refactored it and I have to rework it again
2. I think it is in bad taste to submit PR without intention of getting it merged - it create open ticket that maintainers could perhaps never close and have to sort through.
What's the trouble? What compells you to "bring the fork up-to-date"?
> The fork might not have new features, security updates, compatibility upgrades that I want.
I really don't understand what point you think you're making. The scenario you're discussing is either a) contribute a patch to an existing project, or b) write your own from scratch. Does scenario b) magically introduces feature and maintenance work on your alternative implementation?
> 2. I think it is in bad taste to submit PR without intention of getting it merged (...)
Why are you even considering this a scenario?
Adding a small feature to an existing codebase will never require more work than implementing a minimal working example from scratch.
You're somehow ignoring all the development work you need to invest to go from zero to alternative implementation, not to mention the maintenance work you need to do to come close in stability.
The whole point of FLOSS is that everyone, including you, can build upon all the work invested by everyone else. Otherwise there would be a whole lot of people reinventing the wheel.
- learning a new language
- learning a new codebase
- modifying that codebase
- distributing a that modified fork of a project in c++ alongside what I was building in Go
It was much easier in my case to build it as a small part of the tool I was already building and move on with my life.
Also, if you’re parsing .proto files directly, you have to deal with a bunch of annoying issues like include paths, how you package sets of them to move around, etc. descriptor sets seem like a better solution to me.
The other nice thing about all the descriptors is their schemas are protobufs, so language agnostic and easily serialized/deserialized.
I don't understand the point of using tree-sitter to repeat that work (almost certainly having bugs doing so). Am I missing something?
E.g. just use `string` instead of `StringValue`.
If you absolutely do need it, then in addition to "name" you can also add a field "hasName". If you feel fancy, you can even do that in a code generator.
A better example is Int32Value. By default, protobuf will serialize int32 to 0 if it's not present.
E.g.
- type.specified => “”
- type.unspecified => empty
The same technique can be used to disambiguate between 0 and empty.
Also the wrappers are known to most protobuf code generators, that can then use language features for better ergonomics. E.g. in Rust all wrappers to allow nullability of values will be turned in Option. Like the StringValue will turn into Option<String>.
As for the need, yes, it is needed. This carries intent, which is a good thing in customer facing APIs. An update endpoint will have all fields optional to distinguish between a field needing to be overridden or not, without the need to re-send the value when it does not need to be changed (à la PUT). Having additional boolean fields gets unwieldy quickly.
https://github.com/protocolbuffers/protobuf/blob/main/docs/f...
I'm not as familiar with the Go reflection tools, but getting the information the author wants is trivial in Java reflection.
I'm using it for syntax checking across 30+ languages in Plandex[1], an LLM coding tool. tree-sitter runs in single digit milliseconds on typical files and is highly accurate. When it encounters a syntax problem, it can pinpoint the exact location/expression in the file, and it's fault-tolerant so it can keep going and identify multiple issues rather than stopping on the first one. These results can be sent back to the LLM, which can then often fix its own errors. I was able to reduce syntax issues by roughly 90% with gpt-4o using this approach.
Afaik, there's no other viable option for a use case like this. You'd need a menagerie of language specific linters, compilers, and/or language servers to get anywhere close, and many of those are way too slow to run inline.
This looks good too, from a year ago: https://news.ycombinator.com/item?id=36106421