3,778 karma · joined April 23, 2011
Which link were you comparing FFI in python to json in python? The link I see seems to be about pure JavaScript language Protobuf impls vs json.parse?
Which is not a very generalizable situation.
It looks like the CJK unified space is over 20,000 characters, so that's a real technical magnitude distinction compared to Latin and Cyrillic. "Asian languages be damned" seems like a bad faith read, compared to "Java and Windows char is 16 bits and that will never change realistically" (and in fact they still haven't, even in 2026 things which rely on UTF16 instead of UCS2 are still commonly bugged unfortunately)
Do you have evidence that people would care?
The specific case of this poster it appears they are claiming it was disclosed by the submitter that it was AI and still picked. Maybe that's diffusing the situation after, but it seems most likely these people selected and asked to care about the poster knew it was AI and still selected it. Most people would care less than they did.
Standard Ebooks for high quality public domain books, though they also offer azw.
In practice when you make your new set of hacks the string you can always evaluate whatever cruft in the useragent today, but next browser shows up.
It does make the user agents insane but I don't know if there's any obviously better system for the problem, even with hindsight
Yes, we actually do exactly this (not just protobuf but generally all of the other C code with it for the given application). But there are other approaches that can work.
> But also now that header is deprecated anyway
It's deprecated exactly because the preferred/supported/default runtime is the upb based runtime. The odd gencode design is what made that transition possible as an implementation detail. So the deprecation is just the same topic here, that people successfully setting up the dynamic linking is not really realistic/viable unfortunately, and the effort for open source support is spent making the best behavior on the upb runtime which can't have that feature.
In the bad old days there were so many differences between html, css and js behaviors that if you wanted your site to be nice you had to change it for the browser. The way css padding worked wasn't even the same. Feature detection was rarely viable for any of this.
No user agent would probably have only entrenched IE6 dominance even more by blocking you from deliberately making a site that works at all on other browsers (including IE7 for that matter)
But its kind of a clear example of a surprisingly complicated technical space. The big picture is that it's actually not weird when a simpler thing which has less constraints can be better on some axes users when compared to a complicated thing that delivers on many advanced constraints.
If "readability of the .py gencode instead of the .pyi gencode" was the worst pain point (or pet peeve) here then I actually suspect Buf wouldn't have even bothered, it's just one of these that is easy to see and explain.
It's great to have a healthy ecosystem in the world around Protobuf. Google can't possibly fill all use cases, there's many tools and Buf makes good tooling. The Protobuf team at Google intentionally tries to enable an ecosystem around Protobuf including examples like this. Google Cloud APIs are intentionally usable with any compatible thing that can understand Protobuf encoding including this one.
Kudos to Buf for making something that I'm sure a ton of people will find useful and which takes conformance so seriously.
Just to chime in with some context about Google's own implementations here though (since that's a lot of the discussion otherwise).
Google definitely takes Protobuf seriously including for the long term: you can't really understand how engrained it is within the Google stack without seeing it for yourself. It's not just RPC layer, it's storage, logging, FFI. Html templating is driven off Protobuf messages. Systems which interact with bank XML based systems uses Protobuf schemas. Internally it's widely used for in-memory library api types even without any direct/obvious connection to serialization just because it makes internal details like logging easier. This extremely large surface does create constraints and use-cases to balance. You can see Buf's reported numbers reflect that it is faster for a usecase they expect is typical, but at scale users do fall into the other buckets shown, affecting the performance of preexisting code is a major concern for our implementations that a greenfield implementation doesn't have.
Wide exposure in critical paths alongside long term support directly causes some quirks: for example some of our APIs followed PEP8 when it was created but PEP8 changed. It looks stupid that we have wrong style APIs but also it would be stupider to break compatibility for style reasons. JavaProto as another example still supports Java8 and the runtime is compatible with 2014 gencode which is a pretty major constraint.
Google Py Proto implementation has one extra interesting choice of the same gencode is reused with 3 different implementations (upb, a complete pure python one, and one that uses C++Proto as the in memory representation which libraries like TensorFlow can use to share memory between Python and C++), which is why design the way that it is with runtime created classes, the pyi files are readable but the .py files not.
This definitely has pros and cons, and the direct approach taken by Buf here really makes a ton of sense. It's just that Google's maintained implementation falls into a different spot in a larger technical tradeoff space.
If you see things that appear to make no sense with the official implementations, feel free to file an issue on GitHub and we can look, sometimes there is no reason and we can fix it, and sometimes there's a reason which we can explain.
Kudos again to Buf here, I'm fully sure this will solve some set of real business needs better than Google's (but not because Google isn't maintaining our offerings too).
Are you worried about some foreign state actor like Israel targeting you specifically, hacking your devices to listen to you chatting? Or you're worried about US warrantless mass surveillance wiretapping all citizen's smart speakers, and you're worried the US government may spin up such a program?
In the latter case, the scenario you expect will happen is we'll have ~100 million US households are live wiretapped 24/7 without anyone knowing, you'll be part of the remainder living your life blissfully wiretap-free thanks to not having a smart speaker?
There's still some possibilities here: maybe you came across someone with a broken car and believed that to be representative.
Most other explanations here just involve some form of confusion on your part to be honest: that it was the backup-alert noise that it makes _because_ the car is otherwise too silent, that it wasn't actually a Leaf but an ICE engine that you saw, or something else.
We all also have phone in my pocket 24/7 and my laptop on my desk, both with microphones in them. In the event of the government doing warrantless spying on all devices it seems like that is a strictly higher ROI target for them?
So at least for that one it kind of draws into question your overall conclusions here. You might be hearing an ICE and thinking it is electric, or some specific badly tuned vehicle or something.
As far as I can tell home smart speakers are being used for warrantless mass surveillance, unlike Flock for example. Do you mean the possible future situation where they are?
And then the attack is to trick this recommendation system into putting a link out
I actually the attack is very likely already soft defeated by an interstitial telling you that you are leaving the site though, it would be weird if they didn't do that in general from this surface
I think theres very little chance this particular report made it to any engineer who works on product at all, because if they did they would be completely overwhelmed by reports, the filter which has to handle the many thousands of reports based on a playbook almost definitely filtered it out before it made it that far.
Bots manipulate review scores, posting link spam to other users, crawl your database that isn't open to crawl, etc.
At the upper bound, fraud can always be committed by paying real people with real accounts to perform the desired action in a way that is 100% truly indistinguishable from organic. There's fundamentally actual prevention technique at the limit.
So the entire game is only "increasing the costs until it's not viable ROI", not "holistically prevent", which is why fingerprinting is a relevant technique here.