An Unbelievable Demo
brendangregg.com
brendangregg.com
(Just to be clear, there was no license violation involved in this case; just a lack of awareness of the provenance of the open source software they were using.)
What are some real life use cases for it? When does a developer need such a tool?
The game industry really hasn't used binary diffs and patching since the 90s when they used RTPatch.
Flight Simulator downloads tens of GB every time I start it. :-/
Also, as a dev, you have no idea what version your users are updating _from_. You either need to generate some number of patches for every version you could be updating from, and figure out if you should just download the whole thing again in any of those cases anyway.
Does this happen with more advanced compression algorithms? I've rsynced zip files of different versions of internal software and the diff was always much, much smaller than the entire package.
Heresay, but from what I've heard modern games may ship multiple copies of some assets with different levels or features so they can be loaded as a sequential read off the disk. While a block-oriented compression algorithm might sync up more reliably, if you're packing 200MB of assets for a level and they're all compressed to take advantage of the fact they'll be read sequentially could mean a change 25MB in would still ship ~175MB of changes.
It is just really poor programming, nothing more. And it's everywhere. If find source >/dev/null takes 6 seconds there is no reason for gradle to take 2 minutes on a rebuild. If the dev is used to that, why would they even think about patch optimisation?
The reason games (and software in general) do full downloads instead of binary patches is purely overdefensive and/or stupid. Store software could just check checksums after a patch and re-download only if they fail.
I'd argue that zip is a relatively simple compressed archive format. Its simplicity is its charm and the reason it's so popular. More space-efficient algorithms would be less likely to be "patchable" as there would be less redundancy / structure in the compressed representation to exploit (the best compression seems like it would have similar properties of random data.)
zstd itself also has the (pretty new) ability to use a shared file as shorthand during compression. What that means in practice is that diffs can be REALLY tiny if you have the previous archive download.
Hi dang.
Clarification: .zip (unlike .tar.gz for example, or "solid" .7z) compresses each file separately, that's nothing to do with the compression algorithm used. In addition, DEFLATE, the LZ77-based compression which is by far most commonly used in .zip (and also by gzip) has a window size of 32kB (uncompressed). So yes, even if you used DEFLATE on a solid stream (e.g. zipped a .tar archive) it couldn't remove any cross-file redundancy once it's gone past the first 32kB of each file.
Each file is 1k-20k, of which there are 40,000 or so. But they are catalogued in 3-4 deep directories, so if you just zip them, the metadata takes 30% or so of the zip.
But the metadata does compress very well, so they zip it again.
For zip files each individual file is compressed independently. So unchanged files and prefixes don't need to be resent, even if once a file changes the entire tail end of it needs to be resent.
Some times compression algorithms "reset" periodically. For example the `gzip --rsyncable` patch. This basically resets the compression stream so that a change will only affect part of the compressed file. This does have a cost in terms of compressed size because the compressor can't deduplicate across resets. However if the resets are infrequent you can maintain fairly good delta transfer with little space overhead.
Additionally some delta transfer tools detect common compression and decompress the file "in transfer", performing the delta checks on the original file.
It feels a bit egregious when I have to download a 100MB update just because a few characters were buffed or nerfed. More involved changes end up being over 1GB.
It is a complete image, but phones today have nontrivial state that may be a problem - e.g. your baseband processor might have its own rom with its own update protocol, which changed between image 2 and image 7, so image 10 after image 1 will be unable to update the baseband.
I honestly consider that a pretty reasonable trade-off.
Not all versions are made equal either - one might be a character buff, another might reorder assets in the "big huge binary blob file" for performance improvements. At a certain point, rather than downloading 30MB per update for 25 versions, and applying each incrementally (remember that you have to do them in order too), just download the full 1GB once and overwrite the whole whing.
Most backup software is able to do good binary deltas of arbitrary data for decades. Even dumb checkpointing resolves problem of downloading 25 versions - you download latest checkpoint and deltas from there.
Don't excuse poor design and programming, when you know a file structure, creating a differential update should be short task. With a tiny bit of algorithmic knowledge you could even optimize the process to only download needed assets inside of you big binary blob - if the asset was changed 7 times during your last 25 version you only need to download the last one.
Instead, we just pack the uncompressed files together (frequently using normal zip in a no-compression mode) so that we can avoid needing to ask the OS to open and close files for us or examining the contents of a directory, both of which can be kind of startlingly slow (by video game standards) on some common OSes. Instead, we will generally cache the directory data from the zip file and just use that rather than go to disk.
(of course, the whole download/patch would all be compressed for network transfer, but files would then be decompressed during the installation process)
_You_ mightn't but the last three AAA games I worked on do/did. PS5 expectes compressed files, and does HW decompression (ahem, mostly) on the fly.
Ps4 also did the compressed packages by default thing if I remember right. The upside there being ample cpu for decompression such that no compression was never fastest.
On the Windows/Mac/Linux title I’m working on now, I definitely measure a sizeable improvement to performance when loading from an uncompressed zip rather than from a compressed one. But even that could be down to the particular set of libraries I’m using to handle it.
Did you actually benchmark this? It probably makes sense in your head, but on any vaguely modern hardware it's very unlikely to actually be true because of how exponential the memory hierarchy is.
The simplest one is generate patches for recent versions, where recent can be years in the past. It is a linear operation but you only run it on release so it probably isn't a huge cost. You can also use some heuristics such as if if diff is >20% of the file just stop and force users still on that version to do a full update.
A second option is using zsync[1]. zsync is basically a precomputed rolling checksum. The client can download this manifest and they download just the parts of the file they need. This way you don't care about the source, if there is any similarity they can save resources.
And of course these can be combined. Generate exact deltas for recent versions and a zsync manifest for fallback.
[1] http://zsync.moria.org.uk/
Side note: One nice thing about zsync is that the actual download happens from the original file using range requests. This is nice for caching as a proxy only needs to cache the new data once. Is there a diff tool that generates a similar manifest for exact diffs? So instead of storing the new data in the delta file it just references ranges of the new file.
Games frequently use an override directory or file. The patch contains only the files that have changed and is loaded after the main index and replaces the entries in the index with the updated ones. This is the most common way of doing a patch if it's not just overwriting the original files.
Some games load their file as a virtual filesystem and then the patch just replaces the entries in the virtual store with new ones. Guild Wars 2 works this way. This is only common in MMOs though.
There may be prior art to that, but as a young coder that was the first time I’d seen it
I wouldn't be surprised if the other consoles also do things this way. It's a very sensible way to manage updates — especially when a game is running off of physical media but the updates are held in local storage. It also means there's no point where the update gets "merged in" to the base image, which means updates can be an atomic thing — either you have the whole update file downloaded + sig-checked (and thus it gets added to the overlay-list at boot) or you don't.
And, if all the consoles are doing it, I wouldn't be surprised if studios that do a lot of work on console don't just use that update strategy even on PC, for uniformity of QA, rather than for "obfuscation."
Games are directories/packfiles containing many individual files, mostly binary art assets, plus one executable that takes up a negligible proportion of the total size. When binary art assets in the directory/packfile are updated between versions, they don't really "change" in the sense that a source-code file might be changed a git commit; instead, they get replaced. (I.e. every file change is essentially a 100% change.)
The "binary diff patching" you're talking about the game industry using, was just the result of xor-ing the old and new packfiles, and then RLE-encoding the result (so areas that were "the same" were then represented by an RLE symbol saying "run of zeros, length N"). For the particular choices being made, this is indeed much less bandwidth-efficient than just sending a new packfile containing the new assets, and then overlay-mounting the new packfile over the old packfile.
bsdiff isn't for directories full of files that get 100% rewritten on update. (There's already a pretty good solution to that — tar's differential archives, esp. as automated by a program like http://tardiff.sourceforge.net/tardiff-help.html .)
Instead, bsdiff is for updates to executable binaries themselves (think Chrome updates), or to disk images containing mostly executable binaries + library code (think OS sealed-base-image updates — like CoreOS; or, as mentioned above, macOS as of Catalina + APFS.)
In these cases, almost all the files that change, change partially rather than fully. Often with very small changes. The patches can be much smaller, if they're done on the level of e.g. individual compiled function that have changed within a library, rather than on the level of the entire library. (Also, more modern algorithms than xor+RLE can be used — and bsdiff does — but even xor + RLE would be a win here, given the shape of the data.)
There's also Google's Courgette (https://www.chromium.org/developers/design-documents/softwar...), which goes further in optimizing for this specific problem domain (diffing executable binaries), by having the diff tool understand the structure/format of executables well-enough to be able to create efficient patches for when functions are inserted, deleted, moved around, or updated such that their emitted code changes size — in other words, at times when the object code gets rearranged and jumps/pointers must be updated.
The goal of tools like bsdiff or Courgette isn't to reduce an update from 1GB to 200MB for ~10k customers. The goal is to reduce an update from 10MB to 50KB for 100 million customers. At those scales, you really don't want to be sending even a 10MB file if you can at-all help it. The server time required to crunch of the patch is more than paid off by your peering-bandwidth savings.
In fact, if you use smarter compression than RLE, I wouldn't be surprised if the update was larger than the original binaries after the xor, as an offset xor will likely increase chaos (entropy) in the file, making it compresss worse than the original.
bsdiff was specifically designed to intelligently handle these situations, which is why it works.
Just tested it on Chromium from my package server (90.0.4430.72 vs 90.0.4430.212):
- Original binary: 266MB
- Gzipped binary: 106MB
- Gzipped XOR "patch": 228MB
- bsdiff patch: 47MB
XOR-and-RLE works well for binaries from non-HLL languages (assembler, mostly) where — due mostly to early assemblers' lack of support for forward-referencing subroutine labels from the data section — subroutines tend to ossify into having defined address-space positions.
You can observe this by the fact that IPS-patchfile representations (which, while a different algorithm, is basically equivalent to XOR-and-RLE in its results) of the deltas between different versions/releases of old game ROMs written in assembly, are actually rather small relative to the sizes of the ROM images themsleves. The v1.1 ROMs are almost always byte-for-byte identical in ROM-image layout to the v1.0 versions, except for where (presumably) explicit changes were made in the assembler source code. Translated releases are the same (sometimes, but not always, because they were actually done by the localization team bit-twiddling the original ROM, because they didn't have access to the original team's assembly code.)
(This is also why archives that contain all the various versions/releases of a given game ROM, are highly compressible using generic compressors like LZMA.)
That said, patches aren't really downloaded as standalone patches anymore because of Steam distribution. The way Steam handles it is documented, and if you're interested, it's available here: https://partner.steamgames.com/doc/sdk/uploading#Building_Ef...
But as an overview, Steam splits files into 1MB chunks and only downloads the 1MB chunks that have changed. The 1MB chunks are compressed in transit. Steam also dedups the 1MB chunks. I would assume that this works fine to manage the tradeoffs between size and efficiency.
Another reason is that certain operating systems originating in the state of Washington have performance problems when you access small files or directories containing many files.
I actually could really see using it, now that I understand what it does.
I've worked on some firmware projects where we did OTA updates and were guilty of shipping the entire binary. Luckily, even the entire binary was rather small, but still it would have been very cool to be able to create a diff and ship only the diff!
/s end open source entitlement section
Thanks for your great work
https://www.chromium.org/developers/design-documents/softwar...
https://blog.chromium.org/2009/07/smaller-is-faster-and-safe...
The explanation here is pretty fascinating
https://www.chromium.org/developers/design-documents/softwar...
It‘s a binary diff/patch utility.
> What are some real life use cases for it? When does a developer need such a tool?
Incrementally update binary files e.g. assets or runnable binaries instead if having to re-send the entire thing on every update so e.g. games, browsers, package managers, …
The standard diff/patch utilities are text-based and even when they do extend to binary data their algorithms and heuristics tend to be biased towards textual contents.
Bsdiff was built specifically with an eye towards executables.
There's a collection of articles they wrote talking about regexes and various pitfalls: https://swtch.com/~rsc/regexp/
1#token
(this is in the HTTP spec's notation: it means 1 or more `token`s, comma separated with optional whitespace around the comma) and hit this, by trying to split the incoming values with, / *, */
I was shocked that this wasn't compiled to a DFA. (I checked, too: my JS exhibited the behavior in the bug report.)This is, I also think, another reason why "simple" text protocols are not really so simple. The grammar above is "trivial", and yet, this is the end result. I don't feel like the library is particularly at fault: I doubt I would have caught this in code review.
That's an odd thing about the tech world, it's accessible. As you get better in different areas you are actually more and more likely to make contact with important people (big names? people who did important stuff?). This can creep up on you if you're not aware what level you're operating at. It can be a small world.
about 15 minutes later Guido van Rossum answered my question.
It's amazing how often doing this completely bypasses any corporate first-line-support structure in the way, and just puts the email right into the inbox of the line engineers working directly on the product. It's also amazing how quickly those line engineers reply. (It's as if they treat "replying to random messages on the product mailing list" as their highest-priority job. Or maybe it's just that they're technical people, and my questions are usually very nerd-snipe-y, and get them hooked.)
I've had similar experiences in other language/framework communities. It's amazing how helpful some of these very productive people can be to random chat visitors :)!
I do feel a little bit embarrassed if it turns out they're reading the docs to me, but I feel embarrassed about that whether it's an expert or a fellow n00b ;)
I almost always solve problems with the products/services we use myself — up to and including forking the vendor's codebase to fix their shit for them — because it's almost always the fastest way to do things. I've already been working with their product for a while, and I already know exactly what my own problem is. Provided I also know the language their code is written in, that translates to being able to code a patch myself, faster than I can get someone on their end to comprehend the problem I'm having.
That applies up until the point where there's a problem surface that's just plain inaccessible to me (i.e. the inside of a proprietary mobile app or SaaS service), at which point I have to reach out to tell them that it's broken / missing something on their end. (And even then, if I have a spare hour and access to the offending binary, I'll reverse-engineer it a bit to see if I can hotpatch it while waiting for them to get back to me.)
I suppose, for people who don't think this way, there can be value in "support." But IMHO there's more value in just hiring some DevOps engineers who do think that way. Then all the easy "support" requests get handled in-house, and so you'll only ever need the kind of "support" that involves direct bug reports to the engineers from the vendor who built the thing.
How well does that work for a hosted cloud service?
But otherwise, like I said, that's when "the problem surface is inaccessible."
if your product is targeted to everyone but only power users can figure it out when there is an issue... well you have a problem.
also, being able to figure something out != you should figure it out. your time is limited and the complexity of remembering all those things that you figured out (even if you have the time) will quickly overwhelm you. unless it’s literally your job to support the product you should care about the interface of the product and what guarantees it makes
re: hiring devops engineers. i’m sorry, what? if my email suddenly does not work I’m supposed to hire a devops engineer now?
> Maybe it's just that they're technical people, and my questions are usually very nerd-snipe-y, and get them hooked.
Integrating sales or classic customer support is boring.
I mean, I get that it pays the bills, but when I've got a million priorities, boring work that I don't really get credit for goes to the bottom of the pile.
My point was that there are often public mailing lists, where engineers with real engineering problems could discuss those problems with the engineers responsible for the product/service; and yet the engineer with the problem nevertheless doesn't even think of using the mailing list to reach out, but instead decides to go through regular customer-service support channels to get their problem solved.
Imo, this is not scalable or sustainable, and mailing lists are not a replacement for adequate customer support.
The only reason sending emails directly to mailing lists for specific Google products works is precisely because those mailing lists are not public and not flooded with bajillions of emails from the general public. So those who send the emails are already somewhat pre-screened in a way, because if you know that mailing list email address in the first place, you are very unlikely to send something like "my cousin couldn't remember password to their google photos account, can you fix this please". That's why everything there ends up being read and addressed. If those mailing lists were public, then they would be just as useless and ineffective as the current customer support routes currently are for Google.
Tl;dr: mailing lists for specific products are a nifty workaround for the time being, but they aren't a good sustainable solution for shitty customer support. Making those mailing lists public will not only not help solving the problem, it will just make those mailing lists as ineffective as the current customer support. There is no "one weird trick" to solve the customer support adequacy issues with Google,it has to be an actual customer solution that won't be easy and will take time.
I like Discord, especially in the early days, you can reach out to the principal dev, etc. But it soon seems like they either disappear to get work done (good) or spend all thier time on it (bad). Either way you end up with chaos.
1. they're encouraged to do so in public, so that other community members can help if possible, and/or so bots can reply with suggested FAQ answers;
2. the Community Manager will answer with the company line for questions the company has set answers to (e.g. "when are you releasing X?" or "why is [abusive DoS-like pattern of requests to your service] not working?");
3. otherwise, if the Community Manager knows the answer for sure off the top of their head, they'll give the answer;
4. and if not, the Community Manager relays the question to an engineer in our Slack, where we either have an answer off the top of our heads, or we file it as an issue.
Seems to work just fine for us so far.
Some of the engineers are also sometimes in the Discord (and we're all registered to it), but other than the Community Manager, it's not our job to be in there.
By my thinking, that's a "public" mailing list. They're not hiding it from you. The opposite, really — they're trying to get everyone to know and use it, by making it free to any GCP customer, while the actual CSR kind of support requires paying for a subscription to a higher support tier. The mailing list, presented in Google Groups format, is literally what GCP calls their "support forum." It's supposed to take on all comers, including dumb customer asks.
The point is that I shouldn't have to bypass the official channels that way. These organisations are operating at the level of ad-hoc individual heroics, which is the lowest tier in terms of organisational maturity. In a start-up where everyone has to do everything and no-one has worked anything out yet, that's completely understandable. In a many-billions-of-dollars business with enough influence that someone's quality of life or the viability of some other business could be profoundly affected if the giant screwed up, we should be demanding better by now.
In one sense, it's hard to blame them. After all, if no-one who matters to their revenue stream is actually going to change behaviour because of that dismissive policy, it saves them all the overheads of providing useful support and costs them practically nothing. It's just good business, right?
What is strange is how they've got away with it for so long and most people still don't seem to be switching to alternatives, even as the tech giants casually squash them without even noticing. At some point around here, the words "competition" and "regulation" enter the room.
Apparently the guy answers about 1000 emails per day.
I didn't realise the author of that tool was in the channel - kinda neat that we have such a flat open structure at times.
Reportedly, this was also a major factor in Google's strategy shift to open-source a lot of their infrastructure (GRPC, Bazel, TensorFlow, LevelDB, etc.)
In 2007, an academic argues with PG and others. After long exchange, academic defensively states their own achievements.
Another user challenges academic with GP question, to which academic provides unexpected affirmative response.
More often, we refer to ourselves in the first person plural.
Tarsnap's still there though, so I'd say that's a good defense.
But I was having a bad day, and I said things in a different and more abrasive way than I normally would.
Pretty good bad idea
10-ish years later, I'm doing other stuff but still working with Blizzard-related tooling sometimes. I was talking to a Battle.net engineer about their latest-and-greatest game update protocols. He tells me they're thinking of adopting this great thing called bsdiff for the next version. I giggled a bit.
Anyway, hi! I implemented a version of your algorithm when I still barely knew wtf I was doing. https://github.com/jleclanche/python-ptch
The team he'd worked on had produced a tool that was only ever intended to be used by the team to solve a particular problem they had. It contained proprietary code.
Unknown to the team, word had spread about the tool, and others had started to use it, including solutions architects. Who started shipping it to customers to use, who absolutely loved it.
That'd be fine except one of the core libraries it used was GPLv3 licensed, and there was non-open source proprietary code used in the tool.
The nightmare scenario he found himself in was having to rapidly re-architect the tool around a non-GPLv3 licensed library, without breaking any functionality, all the while having to have regular sync up meetings with a furious CEO and Legal department (who, to be clear, were mad about the situation, not this particular developer or his team, who weren't to blame)
... or just go with it and have it be open source? The old version is already open and free for anyone to request the code of. No rush at that point, you can withhold updates for a little while while you rearchitect this or take the situation as it is and have the next few bugfix releases also fall under GPL until you get around to replacing the core component (iff one insists that the future additions must absolutely be proprietary).
Quickly removing the code doesn't change the previously released versions' license.
But would $BIG_CORP publish source on request for a proprietary product just because they built one version with a GPL library by mistake and later fixed it? Has this been successful, ever?
The built-in conflict resolution in the GPL is no-distribution.
They didn't have the legal right to do so.
Company B uses that library and a GPL library in a product. They distribute the product.
Company B has no right to relive se Company A’s commercially licensed library under the GPL. Hence, stop distribution and replace GPL library.
> ... since there is no mention of other claims or parties to the mix [like company A]. That means the company [B] owns the copyright to the added code and is free to comply with the contract (license).
I'm also not sure whether the confidentiality clause would weigh heavier than a 'must provide source on request' clause, perhaps it could be resolved by not distributing the part that's covered by the confidentiality clause since another standalone library is clearly not a derivative work of the GPL-licensed library. Then only B has to distribute what they made for everyone's benefit.
Next thing I knew someone had copied it and started running as a rule on every data set / customer they could get their hands on, and of course it was false positive city.
Finally after lots of emails where I would just type "Don't use that script, it doesn't work." some engineer wandered up the stairs to support to talk to me.
They were fielding escalation after escalation for these false positives. Support would run the script, see some flags, and turn their brain off and escalate. Management was so scared of this bug / issue that they would do the same.
So he tells me to use a trick he used with a similar situation.
I announced a new and improved script and management militantly demanded everyone use it, and that all escalation using the 'outdated ' script would be rejected.
The new script just identified if it was the right customer for that script and set a bit if it wasn't. If that bit was set the engineer knew immediately they could ignore the output and would say that their analysis didn't find the problem in question and advised some basic troubleshooting next steps (copy and pasted mostly).
The support soon realized that just running that script wasn't getting them much more than a few minutes of breathing room away from the case, management realized this too and saw all these next steps coming back and the focus switched to 'hey we should do these next steps all the time too'.
That was also one of the ways I started to understand how the engineering team worked and really helped start a good relationship with them.
it's still untested in the real world if gplv3 "taints" derivative work that broadly.
Would they had to open source the entire solution? or just changes made to the library? nobody knows, and there's a lot of FUD to promote BSD licenses instead of sitting down and properly defining the limits in a practical way.
I'd get a lot of other students coming to me for coding help. Most just wanted me to do their job for them and I was too naive to say no. One wanted to count cells in a microfluidic device using image processing. I sat down with them for a couple of hours and walked them through a few methods they could look into to get started collecting all the examples in a script. Basic stuff so they wouldn't feel overwhelmed. A few months later I see he published my simple introduction as a paper with zero modifications. He had the good grace to at least thank me in the acknowledgments.
Several years later while working a $BIG_TECH lab we interviewed a candidate from my old lab. They presented their work and had performed some data analysis of thermal camera images. Turns out they were using a script I'd written there and it was still actively used to work with the thermal camera. Nobody ever modified or improved the code — many engineering students are terrified of code. I was annoyed because there was no interest in further development the code, I don't think they even read it or understood it.
While at the spinout company I developed a software tool for the PCR optofluidics platform that was being developed. It was considerably faster and more robust than the hacked together script they were using before and had a user friendly UI that I build with feedback from the biologists on the team. A few years later the founder and one of their new students published a paper documenting their amazing tool without any reference or acknowledgment whatsoever. That one pissed me off.
There is a lot of ignorance around code authorship and respect for the developer in physical sciences research. Like I said, many are terrified by code but don't value the time and expertise it requires; once they have it they no longer think about its maintenance or acknowledging the author.
That can't be the whole story, surely? You verbally suggested a couple of possibilities for what might work and wrote down a couple of lines of code, and then I imagine the student tried out all of those possibilities and reported what did and didn't work? I mean, what academic journal would want to publish half-working examples with unstudied properties?
Mind you, I agree about your broader point that (especially in academia) a lot of people don't really understand and respect code authorship.
The overall problem is that people tend to think only in terms of first order consequences. A little copying here and there might have minimal financial or reputational risk. But the second order consequences of that becoming the example and the norm for the junior ranks and the next generation causes larger organizational risk. So got to consider the bigger picture and nip the ethical lapses immediately.
He had some sample images he had taken and I used them to demonstrate the basic code. The figures in the paper are the ones I generated in my sample code.
I'm sorry to hear about you not getting credit, that's inexcusable, in the same way that not accrediting the researchers who did the work in a paper or book (I have heard this story too often) is inexcusable.
There were a couple of downsides; he got calls for years after asking for help with the tool, and one boss seemed envious which led to other issues.
A long time ago (early 2000s) I wrote a handful of tools to ease physical to virtual (P2V) migration for Windows Server. I was a systems administrator in the UK working for a large American firm.
I wrote the tools in my own time and unlike in other countries my employer had no ownership of them, just to clear that up at the start. I did use these tools at work but never developed them on company time or resources. They were released as open source.
Fast forward 6 months and we had a meeting with a virtualisation consultancy trying to sell us some tools to assist in a wider P2V programme. We had a sales guy and a tech guy visit to show us their stuff. After half an hour of them talking up all they can offer the tech guy fired up a tool and I instantly recognised it was my tool but with some rebranding.
My manager looked at me slightly confused as he recognised it too. I let them continue for a few minutes to properly confirm my suspicions then mentioned that this tool was in fact a tool we already use. They were very confused until I loaded up my version on a remote system to show them.
Needless to say the rest of the presentation was extremely awkward. I believe these two gentlemen thought they were in fact their own tools developed in house. It turned out several of "their" tools were in fact rebranded versions of mine.
I would love to say there was some kind of exciting conclusion but in reality all that happened was they were clearly spooked by this as their/my tools were removed and never included again in their P2V toolkit.
I suspect the moment they left the meeting with us they called to report what happened and rather than risk me following up (not sure how I could do that to be perfectly honest, it was FOSS after all just not used the proper way) they decided it would be safer to just pull these non-critical tools. They were just "nice to haves" anyway.
My boss excitedly share the experience with the team and we had a good laugh about it and how the sales guy went from Mr Confident to stammering and stressed in a matter of seconds.
We never did buy their toolkit.
This meeting does not sound at all awkward to me if I was the sales person. To me this sounds like how lots of companies work with open source to make a business.
They take some open source tools (maybe they built it themselves, maybe not), package then and maybe sell some cloud service around it or simply just support.
If I was the sales person I would be delighted to meet the person who wrote a part of the package.
The only reason to be embarrassed is if it did not contain the correct attribution and recognition.
Edit: the fact that it happened 2000 could have made it embarrassing though. Many things have changed since then...
The reason behind it being awkward was for half an hour the sales guy had talked up how they were the only company developing tools like this, etc, etc. How the 'big players like VMware don't care about these pain points admins have to deal with' or words to that effect (which was true and why I made the tools in the first place).
Then the moment he finished the sales pitch and I see these 'one of a kind tools' they have been developing I respond with "That's my tool, see..."
I believe the sales guy, at least, believed all the tools he was talking about were made in house. Open source wasn't a widely understood thing back then with Microsoft talking about Linux and open source being a "cancer" and such. Really knocked him off course as I guess he had never been in that situation before (I doubt many have?).
As for following up with the company. I did nothing. I was young and while I knew what they did was wrong (they literally removed my name, link to my website, etc. and put their company name but changed no functionality of the tools as far as I could see and certainly no source code was available!) I didn't have the confidence (or desire tbh) to chase up on some little tools I made to learn and make my life a little easier at work. The company disappeared (I don't know why) sometime around 2010 iirc.
I mean it was GPL2 so they could have just used it as is.
Of course this was the early 2000s where many people saw "open source" and felt like it meant they could just do whatever they want. I bet they never entertained the idea I would ever find out let alone be sitting in a sales pitch :)
Shows Brendan's maturity.
I am not sure what should be the appropriate reaction or corrective measure in these situations. We should talk more about handling these unfair situations.
Someone else can become more successful building on top of one's open source project. On a resume, a top contributor and a minor contributor to open source project might have same weightage depending on how you present it - making the situation unfair for a person dedicatedly working on a single project (quality) vs minor contributor to multiple projects (quantity).
But deleting name and credits is wrong. An acknowledgement from the benefitting person (if not the recognition/reward) has far more positive impact on career than justifying to other's that your work was stolen.
It was a bit strange to read some of the initial negative comments. I see Brendan being a sport. I would argue that reading the story as a report against unknown persons at Sun makes more sense. I don't see much sense in blaming victim. And, in my opinion, the VIP had a good run but he isn't the bad guy here.
Start recording; have them, a big multinational with a massive legal department, admit to violating and stripping a license from source code. Then sue them. They should know better, and they're making billions off of other people's work. That in itself is fair enough, if the license permits it, but removing the license is crossing the line.
Better make detailed notes, who said what with time and date.
Some places...nord Korea? You the recorder have to consent?? I consent to myself that i record others without their knowledge?
https://www.sydneycriminallawyers.com.au/blog/is-it-legal-to...
Alabama, Alaska, Arizona, Arkansas, Colorado, District of Columbia, Georgia, Hawaii, Idaho, Indiana, Iowa, Kansas, Kentucky, Louisiana, Maine, Michigan, Minnesota, Mississippi, Missouri, Montana*, Nebraska, New Jersey, New Mexico, New York, North Carolina, North Dakota, Ohio, Oklahoma, Oregon, Rhode Island, South Carolina, South Dakota, Tennessee, Texas, Utah, Virginia, West Virginia, Wisconsin, Wyoming
[1] https://recordinglaw.com/united-states-recording-laws/one-pa...
Note that I believe (and, IANAL) that if at least one party to the conversation resides in a "two-party consent" jurisdiction, you will need the consent from all such parties.
>But the reality is that it is normally against the law to record a phone call without the other person’s consent.
>In fact, ‘covertly’ (secretly) using a listening device such as a mobile phone or digital recorder and publishing or otherwise distributing that material can amount to a criminal offence.
Recording private conversations:
>The laws only apply to ‘private conversations’, which is one where the parties may reasonably assume that they don’t want to be overheard by others.
>One of the exceptions to the prohibition against recording and/or publishing or distributing records of private conversations is where police officers have obtained what’s known as a ‘surveillance device warrant’ – also known as a ‘wire tap’ – which allows for the recorded material to be used for investigations and tendered in court provided, of course, that the material is relevant to the proceedings at hand.
Between jurisdictions:
>It is legal in all jurisdictions to record a phone call if ALL PARTIES to the phone call consent.
https://www.sydneycriminallawyers.com.au/blog/is-it-legal-to...
But hey if your a Police Officer working on a case your are correct, you don't need the Consent of the other person ;)
Vic, Qld, NSW, SA, Tas, all OK. That's most of Aus.
That's whats you said, and it's not true, without consent from at least one recorded or being in a situation where recording is normal (tv etc) it's pretty much everywhere illegal.
And i wrote 'Most Country's' which is a hint that there is some 'Some others'. There is even a Country where singing under the shower is forbidden, in MOST others...it's not.
I think it's a small world, and everything is software, so the chance you'll bump into someone who wrote software you are using I think is pretty high. I was once trying to get my head around Andi Kleen's pmu-tools, and I had the github repo open in my browser on my laptop I was carrying, when the guy sitting next to me on a bus says he's Andi Kleen. (Ok, it was a bus taking Linux conference attendees to an event, not a random bus, but I still found it remarkable timing -- I was studying pmu-tools at that exact time!)
He leaned over, asked if I liked the blog, and (slightly proudly if I remember correctly) mentioned that deep mind had hired Andrej for an internship starting soon.
I don't understand. What kind of funny looks were they? Disbelief? Distrust? Fear of your mental health? Realization of having been lied to by their bosses (oops it wasn't really an internal tool)?
Also, what were the impact of those funny looks? How did they make you feel? Was there any longer term consequences of telling them you wrote the thing?
Maybe I just don't look or dress or sound like what one would expect. But there's context here too: At the time it's when these things are flagship features and on the booth monitors, and the booth staff are explaining the virtues of these features to everyone they meet. They are making it a big deal of it at the time, so maybe that makes it even more unbelievable that the inventor would wander by at that moment.
Now imagine what would happen if companies had a thanks page along with the other boilerplate pages (contact us, about us) on their website. If you're making millions from a thing, thank the original person for that thing. (I put thanks pages at the end of my slide decks, it's not hard.) These interactions would go a lot better -- "my name is on your company website" -- and could lead to fruitful discussions and collaboration instead of weird looks.
When it happens internally, i.e. I catch someone doing it, then either it is a "first time offense" of the clueless, or it is the act of an unethical person who will be unrepentant. For the clueless, it might be that they undervalue themselves, and therefore undervalue assigning credit. The unethical person, however, understands what they are doing and is simply untrustworthy. They will also likely have a lawyer, because they've done this before. So it can be pricey to get rid of them, but get rid of them you must. They are poison to your team.
She couldn't believe what she saw. A major govt. backed logging company (which does do a lot of dev work themselves) were showing off one of our projects as theirs! We were using depth cameras to estimate the volume of wood loaded on a truck as it drives trough a gate. They even used screenshots that I had made!
Now, they were involved in the project. But they were basically clients of our client. They provided us with a place to test the system as their trucks ran trough. They didn't own the software let alone did any work on it. Why they would present it as something interns could potentially work on is beyond me.
People are weird.
However one of my favorite questions to ask is, "What don't you like about what you invented?" A true creator of something is always acutely aware of its flaws. They're unsettled about its shortcomings and wants to do better. In describing the flaws they demonstrate deep insight into the problem space, and they explain expertly how things don't go perfectly in certain corner cases or how the code could have been better organized.
The poser will almost always struggle to come up with criticisms of what they "invented" and try to pass it off as a spectacular feat of engineering brilliance.
I can't speak for the US but, in the UK, don't misrepresent your work in a job interview. I can't say you'll never miss out on an offer by being honest, though I don't believe I ever have, but would you really want to work for people who'd prefer you to lie or misrepresent mistakes you've made than be open and truthful about them?
To me that's something of a red flag: it's at least indicative of a culture where mistakes are likely to be covered up, leading to a lack of reflection, learning and improvement... and also quite possibly storing up bigger problems for later.
(FWIW, I started as a developer and am now the CTO of a mid-sized multinational market research and insight company. This is nowhere near as grand as it might sound, and isn't meant to be a boast, but hopefully illustrates that being honest doesn't appear to have done my career any long-term harm. Some things that have, if not derailed my career, caused me to take some fairly substantial detours: (i) taking things too personally, (ii) placing too much weight on others' assessments of me, (iii) and I say this as somebody who is wary of people who change jobs too often, but... staying in a job way past the point where there was anything else I could learn/give/progress. I am, of course, but a single data point.)
I wouldn't paint all tech companies in the U.S. with such broad strokes. In the interview loop for the job I have now in the U.S., every single interviewer asked me a question about "how things could have gone better." I talked about mistakes I made, lessons learned, and how I could do better next time.
I am told the feedback from that loop was across-the-board "outstanding."
"Leaders listen attentively, speak candidly, and treat others respectfully. They are vocally self-critical, even when doing so is awkward or embarrassing. Leaders do not believe their or their team’s body odor smells of perfume. They benchmark themselves and their teams against the best." [1]
OH YES.
I've done a lot of things for both work and side, and this is the one thing I can go on and on.
Things that are a choice of the lesser evil, things that are unfortunate, things that we just don't have time for.
(No credit of course, and the marketing copy around it made it sound like it was all their own code. Welcome to open source I guess!)
Could you please link the voxel engine? Which version of Minecraft was this?
The game is still live though: https://classic.minecraft.net - Originally it supported multiplayer, but that stopped working the day after the game launched.
The voxel engine is here: https://github.com/andyhall/noa/tree/develop
lol. Microsoft should just release all the source for this since Minecraft isn't exactly state of the art these days. The brand is nearly all of the value.
So now they see that their code has been stolen from the previous company, and they're being asked to work on it, what will they do? Refuse to use the code and risk being fired for being too slow? Report the crime and risk being accused of attempted to cover their own tracks? Talk to HR and being hated by one's closest bosses?
I see no way to really win here. At least, not reliably.
Just in case what happened? I'm not understanding the work environment that would lead to this practice
As if it had been Bryan, Brenden would have been more sure.
While he and I had corresponded a bunch, ironically the first time I met Brendan was in 2005, it was in Sydney, and I was on a tour of Australia talking about Solaris 10 -- but it followed this incident by several months (if I recall correctly). I was excited to meet Brendan, and he was excited to meet me -- especially so after his poor experience several months prior!
My impression is that reason for the stealing usually makes sense. For example, a key library that's hard to write that's just copied into the source tree, ignoring licensing. Or an appliance developer didn't want to deal with licensing for Linux or BusyBox. Or an individual developer in over their head quietly copies code.
The time I heard an explainable incident happened with my code, was in mid/late-'90s. An acquaintance, who'd offered to be one of the testers for an unreleased Java desktop application I wrote, then reportedly ran it through a decompiler, and passed it off as his own code, in a demo to investors. He later acknowledged doing this, and said he'd send me a Sun (ha) workstation as compensation. I declined.
Then there are incidents for which the reason isn't obvious, like the one from the article. I speculate that sometimes the explanation might simply be that the perpetrator wasn't quite right in the head at the time, like in some famous cases of journalism fabrication.
An inexplicable one involving my code was when an open source developer took a substantial and novel package that I wrote, stripped out my name and license notices, including out of the main file, and posted the package with themself identified as the author. There were also a couple other incidents with that person that seemed like that hadn't yet learned how to play well with others, in engineering or open source. I asked a mutual acquaintance, in confidence, what was going on with that person. The acquaintance checked, and was also baffled. In that case, I suppose that maybe the perpetrator was going through a difficult time, and not thinking clearly. Or maybe it was a combination of unlikely accidents that looked worse than the intent was (which happens).
In the article's story of the Sun incident, I'm a little surprised that (speculating) an engineer could do this despite all the other people in engineering who might be in a position to notice something funny going on. And Sun had been the dot in some dotcoms by 2005, so presumably they had some strong engineering processes around what goes into product.
Maybe the demo was something put together by a systems engineer, working as part of a small marketing/sales team, rather than under an engineering organization, so a lot fewer engineers were aware of it?
I wrote a few silly scripts in my lifetime, all small-time stuff (so trivial it never got me hired as a Dev anywhere). No matter whether I marked things as GPL or BSS, I often found them copied on GitHub with my copyright notices stripped. In the case of the BSD license, that's literally the only thing you can't do!
In some cases it was utterly blatant: cloning from my own GitHub repo and then replacing authorship notices in the very first commit. I mean, come on guys, at least try to be smart; you can add your name just fine... When I politely asked to, y'know, respect license terms and reinstate the notices, some people just took down their repo rather than doing that. I bet they then re-uploaded it somewhere else.
When I contacted them to ask how they thought it was ok to republish my work like this (I was giving them a chance to explain), they just took it all down.
EDIT: For some reason I assumed OP was the author. Someone with a twitter account should ping him and get him over here to participate in discussion.
"You might wonder why I didn't talk about this first case publicly at the time. I'd already informed them of the problem privately and how to fix it, so there wasn't more to say. Also, Sun was the number one employer in town, and to be publicly critical could be career ending."
I don't know what happened to the US developer. But I didn't think it made sense to burn bridges with Sun about it. I was more worried about giving people a bad experience with DTrace by running older versions of my software.
Perhaps largest in Sydney...Great article btw!
I thought Solaris is being completely eclipsed by Linux.
Even windows are surprising because hardly anything except gaming and limited local clients are windows. Granted some people choose to develop on Windows.
If you can prove it's your invention, and they're distributing it without your permission, they won't want to waste money defending a suit.
You may be a little guy, without enough dosh to launch a suit; but you could sell your interest to $COPYRIGHT_TROLL, then they'd be in trouble. So just ask them nicely to propose a settlement.
It's not like war, or anything; tech-company legal departments are there for just this kind of thing. Unless there are arseholes involved, it can all be friendly and business-like.
No, GPL FAQ specifically addresses internal distribution: https://www.gnu.org/licenses/gpl-faq.html#InternalDistributi...
> Is making and using multiple copies within one organization or company “distribution”?
> No, in that case the organization is just making the copies for itself. As a consequence, a company or other organization can develop a modified version and install that version through its own facilities, without giving the staff permission to release that modified version to outsiders.
> However, when the organization transfers copies to other organizations or individuals, that is distribution. In particular, providing copies to contractors for use off-site is distribution.
Please consult sources before making a strong statement.
Which couldn't be more irrelevant, because
a) This software doesn't appear to be licensed under the GPL.
b) The question here is not whether or not it's distribution, it's if it's copyright infringement. Making copies of copyrighted material is illegal (up to fair use and licenses) regardless of whether or not you distribute them.
c) The GPL also requires you maintain the license header when you are making copies (GPLv2 term 1, "You may copy [...] the Program's source code as you receive it [...] provided that you [...] keep intact all the notices that refer to this License [...]" (other terms and conditions do apply, hence the ...'s).
d) The license (and the law, and court rulings) is the source material, not GNU's not legally binding faq.
> Does the GPL require that source code of modified versions be posted to the public?
> The GPL does not require you to release your modified version, or any part of it. You are free to make modifications and use them privately, without ever releasing them. This applies to organizations (including companies), too; an organization can make a modified version and use it internally without ever releasing it outside the organization.
You're not free to make copies without the license. You are free to make copies with the license, because that's what the license granted you permission to do. You aren't free to do so without preserving the license, because you don't have a license to do that and that's copyright infringement.
You aren't required to publicly release derivative works or their source, simply because the license grants you the ability to produce derivative works (provided you keep the copyright information intact) without publicly releasing them. You are required to keep the copyright information intact, because the license does not grant you permission to produce derivative works if you do not.
The requirement to keep the copyright information attached is simply not connected to your choice to distribute it or not. It's a requirement any time you make a copy or make a derivative work, even if you're only making that copy or derivative work for yourself.
I still control all 1000 computers, but I'm doing something I fundamentally cannot do with a single copy.
The question is a bit murky, not because it's unclear I'm making copies, but because some copying is fair use. For example to execute a program I need to copy it from the hard drive to ram, and (pieces of it) from the ram to the CPU. That's generally considered to be legal, despite being copying, approximately because it's necessary to use the program. The exact line here is not well defined.
(Also not a lawyer, just a nerd)
So, shout out to Mr. Pearson for a great presentation and a really helpful SO answer!
A blog post on my blog.
In fairness to him, it was still on a domain that didn’t identify me at all, though the “About” page had my name on it.
He then proceeded to launch up the chain of command to get me in trouble for publishing company secrets on said blog.
I promised to only write it outside of work hours and not to directly copy any code (scripts, one-liners) I wrote at work, and that was the end of that.
There were other less-antagonistic instances of others in this large company finding solutions on my blog, which says something about the value to the company of posting generic stuff like this on publicly-searchable platforms instead of an internal-only SharePoint or Confluence instance, and definitely instead of the PDFs our senior management insisted on for all documentation.
AOL's docs were the same ones from the open source project, with the GPL message intact.
I once received a library under a modified Apache 2.0 license, where all the conditions under section 4 (attribution, etc.) were just inverted for a given number of years (effectively an NDA), after which the normal Apache 2.0 would apply. Which worked because it came straight from the original copyright owners, but here I assume AOL weren't.
Fast forward 10 years or so, I'm consulting at a company. When I see one of their support engineers using a tool that looked vaguely familiar. Turns out the company had acquired the one-man ISP and had continued developing the little helpdesk application I'd made for WAP all those years ago.
I'm a generalist, but I've been thinking lately if performance engineering is something I should be specialising in. I'd love to hear any advice from those in the field.
The other plus is, if you frame yourself as a specialist in X, it might be easier to explain not knowing Y. You can't simply know everything, especially if you invest heavily in specializing in other things.
Personally I've been a web performance (i.e. mostly JavaScript/HTML) guy lately and it's been fun. There's a million of React devs in the world, but only a few dozen web perf guys with web presence (Twitter / blogs), and I almost know all of them by name at this point.
Of course it depends what you want to do. Typically it's big corps who look for specialists. Small companies prefer generalists.
My problem with picking a niche is that ones that are easy to learn from a book are too crowded to be considered niches, and the rest can only be learned by doing in real production environments. Now there’s one where I have a foot in, and it’s one that I enjoy.
As someone who has spent a lot of time in open source, forks are not a problem, they are indicative of a problem.
Don’t cry when people fork and go a different direction, try to figure out why and see if you’re willing to change the project to accommodate them. Dropping chastising “more wood behind fewer arrows” platitudes is pointless when half of the wood wants to break off in a different direction anyway.
Forks because the main project is not responding to a need are not a problem. Forks just because some project or product believes it gives them more control and they don't interface with the original project usefully are a sort of a problem. Those aren't necessarily started because of a community need, but some business need or perceived business need which may have nothing to do with the reality.
I think that's the situation he's asking to avoid here, as he mentions observability products. He's saying please don't fork just to have it under your own repo and for no other reason, when it can be developed jointly.
It’s worse to have people trying to jam shit into the main project if they don’t actually care about the main project. That’s how you get contributions that are huge hacks and require more work to review and iterate on than is worth it to the community.
If some company forks for an observability product and doesn’t contribute back, they clearly don’t want to. Don’t try to force them.
git is an incredible tool for managing forks and this anti-fork mindset is right out of old school open source culture from the early 2000s.
Demanding people contribute to your project instead of forking just sounds like demands to pay homage more than anything.
And yet the author here is saying please work together. Presumably he or the community is okay with picking apart those huge hacks. I'm not sure why you should care what that community is willing to accept, unless you're part of it.
> If some company forks for an observability product and doesn’t contribute back, they clearly don’t want to. Don’t try to force them.
Who's forcing anyone? Did you not notice the "please" he starts that statement with?
> Demanding people contribute to your project instead of forking just sounds like demands to pay homage more than anything.
That's just a straw man, nobody demanded anything, and I'm not sure why you'd even try to insinuate he was when he literally says please.
If I had to paraphrase the part of the post you're critiquing, I would do so as "If you're using these tools, don't feel like you have keep any stuff you do separate. We'd like to build something better for everyone, so feel free to contribute to the project rather than keeping it separate and we can all build on each other's work." That's pretty standard open source ideals IMO.
Sure, that’s just a completely different writing of what’s actually there. The original post is an instruction not to do something. Your interpretation is a much more open “pull requests accepted”.
His reply clarifies that he doesn’t want broken shit out there so your reading is incorrect. He does want to discourage forks because they are subtle to get correct and doesn’t want shit out their sullying bpf’s reputation.
- Someone sells you BPF observability that's buggy and dissapointing, and you avoid BPF in the future. So in that case it's hurt the BPF community.
- Funding goes to a project that gives nothing back, instead of funding a project that does. Again, it hurts the community as funding isn't infinite.
Ok, so someone sells me closed source tooling that uses BPF and it’s buggy and disappointing, so I avoid BPF in the future. This association problem exists regardless of closed source vs a fork vs an old release.
> Funding goes to a project that gives nothing back, instead of funding a project that does. Again, it hurts the community as funding isn't infinite.
This only hurts if it’s a closed/hidden fork. This also presumes that the fork isn’t just stripping out tons of upstream features (I’ve had to do this for clients because “security”).
My view is that forks of open source (GPL in particular) should be encouraged and done in public. Iterating on central projects only is just so slow and stifling.
Maybe your projects have very little churn so lots of people experimenting in parallel doesn’t make sense, but this has been my experience with larger infrastructure projects. Let a thousand flowers bloom.
Somewhat like when powershell on windows had aliases for wget and curl that was just shell aliases for a simple downloader, making some scripts work without those two installed, but also confusing anyone that wanted to do anything outside of "get one single file from url".
But also yes: There are two particular issues with bpf tooling forks. 1) They look deceptively simple, but those that are kprobes-based are really kernel-specific and brittle, and need ongoing maintenance to match the latest changes in the kernel. One ftrace(/kprobe) tool I wrote has already been ported a bunch of times, and I know it doesn't always work and one day I'll go fix it -- but how do I get all the ports updated? No one porting it has noticed it has a problem, and so the same problem is just getting duplicated and duplicated. Which is also issue 2) Unlike lots of other software, when observability tools become broken it may not be obvious at all! Imagine a tool prints a throughput that captures 90% of activity and no longer 100% (because there's now a fast-path taking 10%). So the numbers for some deep kernel activity are now off by 10%. It's hard to spot, and that increases the risk people keep deploying their old broken ports without realizing there's a problem.
It seems like there is a missing formal interface here if this is so brittle, no? If it’s hitting a bunch of internal kernel stuff shouldn’t this stuff just live with the kernel itself?
But kprobes is basically exposing raw kernel code that the kernel engineers bashed out with no idea that anyone might trace it. And they can change it from one minor release to another. And change it in unobvious ways: Add a new codepath somewhere that takes some of the traffic, so gee, seems like my tool still works but the numbers are a bit lower. Or maybe I measured queue latency and now there's two queues but the tool is only tracing the first one, or now there's no queues so my tools blows up as it can't find the functions to trace (that's actually preferable, since it's obvious that something needs fixing!).
I really don't like using kprobes if it can be avoided (instead use tracepoints, /proc, netlink, etc). But sometimes it's solve the problem with kprobes or not at all.
Now, normally such code-specific-brittle things should indeed live with the code like you say, so normally I'd think about putting the tools in the kernel code. But we don't want to add so much user space to the kernel, and, it also opens the door as to whether these should actually be tracepoints instead (which begins long discussions: Maintainers don't want to be on the hook to maintain stable tracepoints if they aren't totally needed).
Another scenario where the tools should ship with the code base would be user space applications. E.g., if someone wrote a bunch of low-level tracing tools for the Cassandra database that used uprobes and were code specific, then they would be too niche for bcc, and would probably be best living in the Cassandra code base itself.
What does it mean, for example in the JVM arguments start with "D" as well.
Any history to this?
In order to trace arbitrary code, it has to inject these into the code site that you want inspected.
If you could just inject arbitrary C, you'd get the issue of potentially adding probes which change the behaviour of your code under test to a degree where new bugs/behaviours are introduced or old bugs/behaviours are masked.
DTrace solves this by using a C subset which helps you avoid such unwitting changes, by not including loops or other operations which could change the memory or timing behaviour of the existing system.
Which is completely unrelated to the other language called D, also known as DLang. (which was released 4 years earlier than DTrace by the way)
To an Australian, introductions in the US can sound
boastful, but they can also be useful as a quick
way to share one's specialties.
But this one puzzles me a bit. Typically, we talk up the people we're introducing - I'm not sure I've ever heard anybody talk themselves up during an introduction!"Sally, this is Bob. Bob's been doing some really cool stuff with XYZ lately. Sally, I know you have too!"
(At which point Sally and Bob often politely insist that no, they're nothing special at XYZ)
I've always thought of this as gracious and not boastful. I definitely agree it would be obnoxious to talk one's self up!
I work quite a few years in IT and never, during any interview or meeting, I've been introduced as anything more than just an engineer. This must be cultural gap, no doubt, but I'd feel weird if someone would detailed my career in front of other participants. Of course I have nothing against filling some details in by myself but only if applicable in given situation. Truth to be told I've never worked in Australia or US, but I did some job in two EU countries and in Japan and, as said, never encountered detailed introduction.
Is this still a thing?
Management, when alerted, made it right, but I think the point of my story is that this is perhaps more common than anyone realizes.
What's the point of publishing with a copyleft license if you aren't going to do anything when someone literally walks into your office and says "we at Big Corp are selling your work without any attribution?"
I was hosting a research talk given by quasi-famous professor at a biotech startup that I worked at.
Quasi-famous prof was describing a gene (gene "xyz") being used as a tool in his lab, "but the specifics of what gene xyz is and what it does are not important. It's just a gene we use in these assays...."
Me: Do you know who has 2 thumbs and discovered xyz? This guy.
Any non-Americans want to chime in on what is a heavy American accent? I’m imagining heavy southern accent, but maybe this is something that can only be heard by non-Americans?
Specifically in my limited experience stuff like pronouncing t’s like d’s and soft back-in-the-throat r’s. “Budder” vs. “buttah”, for example. American vowels tend to sound larger as well in my experience.
> ...which is like a mouth full of down feather.
This means speaking soft, muffled words - roughly the way you'd sound if you had your mouth full of something soft. Like feather. Or food.
Not a perfect analogy, but hey, works as a hash for an accent, at least among the people I know locally.
It's a bit like saying that someone from rural Bavaria who speaks with a strong Bavarian accent has a heavy German accent while speaking German.
So I'd say a strong American accent to an Australian is the most divergent one, not one that somebody in US might consider strong.
Living in upstate New York I also frequently heard "oh, you have an accent" to which I always tried to explain "so do you", and several times I got the response that their accent was either neutral or closer to generic English.
It is amazing the subjective differences in how people experience accents, and how they feel about their own.
But one time I was at a wedding in Karlskoga and talked to someone I hadn't met before. I opened my mouth and managed to get half a sentence out before he interrupted with "Oh! You're from Göteborg!"
Peter Sellers called Americans "The Herns" [1]:
> Various American characters with the surname Hern or Hearn, often used for narration, outrageous announcements or parody sales pitches. The Goons referred to Americans as "herns", possibly because saying "hern hern hern...." sounded American to them, possibly because Sellers once said that a decent American accent could be developed simply by saying it in between sentences.
[1] https://en.wikipedia.org/wiki/List_of_The_Goon_Show_cast_mem...
Personally I'd rather have Americans "lean heavily on the R" than act like the letter doesn't exist (rhotic vs non-rhotic). I think it's another factor why American English is more popular than British English (besides the huge economic factor, the US economy being 5x the UK one), since their pronunciation is clearer and more explicit.
Plus... American multinationals are somewhat close behind. If you want to get a good, well paying job at an American multinational, you have to speak English at least a bit, and you have to know it well if you want to move up the ladder.
https://en.wikipedia.org/wiki/Rhoticity_in_English
>Rhoticity in English is the pronunciation of the historical rhotic consonant /r/ in all contexts by speakers of certain varieties of English. The presence or absence of rhoticity is one of the most prominent distinctions by which varieties of English can be classified. In rhotic varieties, the historical English /r/ sound is preserved in all pronunciation contexts. In non-rhotic varieties, speakers no longer pronounce /r/ in postvocalic environments—that is, when it is immediately after a vowel and not followed by another vowel. For example, in isolation, a rhotic English speaker pronounces the words hard and butter as /ˈhɑːrd/ and /ˈbʌtər/, whereas a non-rhotic speaker "drops" or "deletes" the /r/ sound, pronouncing them as /ˈhɑːd/ and /ˈbʌtə/. When an r is at the end of a word but the next word begins with a vowel, as in the phrase "better apples", most non-rhotic speakers will pronounce the /r/ in that position (the linking R), since it is followed by a vowel in this case. (Not all non-rhotic varieties use the linking R; for example, it is absent in non-rhotic varieties of Southern American English.)
>The rhotic varieties of English include the dialects of South West England, Scotland, Ireland, and most of the United States and Canada. The non-rhotic varieties include most of the dialects of modern England, Wales, Australia, New Zealand, and South Africa. In some varieties, such as those of some parts of the southern and northeastern United States, rhoticity is a sociolinguistic variable: postvocalic r is deleted depending on an array of social factors such as the speaker's age, social class, ethnicity, or the degree of formality of the speech event.
> (rhotic vs non-rhotic).
:-)
Regular Brits use their regional access, yes. And those can be much, much harder to understand than your average American accent. For precisely the same reason rhotic accents can cause issues, those regional accents tend to eat up sounds and sometimes entire syllables.
It's not just the r's (in fact, some British accents are rhotic - around their South West, if I'm not mistaken?), there are all these glottal stops and whatnot.
But, from my observations at least, there are also big discrepancies related to social class. When I moved to the UK, I had no problem whatsoever talking to, say, a local librarian - but a plumber would be nearly impossible to understand for me, in the first months at least. I didn't really experience it in the US, certainly not to such an extent.
American accents are more familiar because of Hollywood, so they tend to be less surprising; and likely because a lot of them were actually developed by people who learned English as a second language, they are often exaggerate in effect, very clear, and actually more regular (particularly on names, where UK "rules" are anything but).
This said, "deep south" US accents, when pushed hard, can become as inscrutable as certain UK dialects.
> than act like the letter doesn't exist (rhotic vs non-rhotic).
Every language changes. Nobody is acting like "the letter doesn't exist", just occasionally that phoneme has changed or dropped in their dialect. Even in non-rhotic accents, an r in the orthography can indicate a change in vowel quality.
> I think it's another factor why American English is more popular than British English (besides the huge economic factor, the US economy being 5x the UK one),
I think the greater population, and the fact that the Hollywood content has embedded itself globally, has resulted in more exposure to American content (of which there is more of). Any dialect will sound clearer to you if you're exposed to it more often than others.
Edit: another example: "door" => "DOH-ur"
I also perceive Canadians to have "heavier" North American accents than Americans do. Some Canadians speak as if from the backs of their throats, with leaden vowels and really round r's. They also often overcorrect /a/ to /æ/ so "drama" becomes "dramma" (like the first two syllables of "Dramamine". And of course there was William Shatner's famous "sabotadge"...
I’d say something to a clerk, and they’d reply with “you’re an American!”
Yes, how could you tell?
You’re so loud!
Thanks?
[0]: https://youtube.com/playlist?list=PLOgT48pM4GctCy-88GTDyyqyy...
Thankfully, my employer had some licenses for the library without my knowledge, but it ain't fun to break licenses at work, especially when you don't notice until months later.
there were many people inside Sun that were not Bryan Cantrill. In fact, almost all of them were not Bryan Cantrill. But there can only be one VIP and it can only be Bryan?
I'm not sure if this is just my personal feeling but I would say that stealing intellectual property was sort of common at that time. Open source was not widely known, knowledge was scarce, communities were just ramping up and really anyone with lack of principles could pretty much steal anything and get away with it.
It happened to me a few times with online content I wrote. Essentially tutorials, articles, etc. around Open Source. Once my own company sent me a newsletter which contained one of my articles signed by another employee from a different place. It felt pretty weird.
I have to say that my limited experience in dealing with Sun as a customer mirrors Brendan's comments around a remarkable arrogance, and it probably played no small part in their downfall.
Yeah, I remember that. I kept thinking "this tool might be nice for C wizards but it does nothing for my day-to-day experience as smalltime Linux user / admin". The other big thing was ZFS, which was interesting, but they were extremely uncooperative with the license, basically ensuring it would never make it big.
This is a made-up story, so you are free to make up your own ending as to whether this was the world-travelling VIP, and if so, whether he had some partial or complete flashback on hearing Brendan's name. One thing we can be sure of: whoever 'carelessly' stripped the copyright notice out of Brendan's code had seen his name before (and it was, perhaps, the part of the code he was most familiar with!)
Can you get more money if they violate the OSS license if you offer the software under a commercial license as well?
I'd much prefer to see the people who build Valuable Things show more interest in capturing some of that value.
There's this overwhelming narrative revolving around Open Source that makes it seem shameful to profit from your work. It's maddening to watch. There's no reason we as developers need to be the low man on the totem pole getting tread on by business people. We just set ourselves up that way and socially punish anybody who doesn't.
The problem he should have noticed was that the Sun was selling the code he wrote for hundreds of thousands of dollars and not passing any of that on to him.
Step one shouldn't have been to worry about putting his header comment back in place and getting them the latest version of his code to sell to their customers. It should have been negotiating a redistribution license for his code if they wanted to continue selling it.
On this topic, I work for Red Hat where we made $3.4 billion in revenues in the last published year (before being acquired), making exclusively open source software which you can download yourself for no cost.
Semi-related... I TA'ed a Comp Sci class a long time ago (back when they had to submit their code as print-outs). I read through the print-outs and noticed that the "look" of one of them seemed oddly familiar (blocks, line lengths, indenting etc). I went back through the others and found another one that was almost exactly the same.
Took it to the Prof and we agreed the ones that copied got 0 on the project, and the ones who allowed the copying got 50% of their mark. Honestly, they got off quite easy.
To be fair, yes those articles were on the same topic, but I'm actually trying to take the next step here. :-)
The best part is when my own blog posts actually have the answer I was looking for - at which point I start feeling old and forgetful.
I feel it would have been ratting to the original creator to do nothing.
Yep. See Sun's response to the Linux SPARC maintainers technical critique of Solaris for SPARC. Meanwhile at Red Hat we started hiring all the ex-Sun people that had been pushing for x86 internally at Sun and gotten frustrated with the flip flops (RHEL 3 on Xeon already demolished Solaris/SPARC) .
Vendor lock-in is not great for the customer, but when you try to milk the customers, you end up risking taking them to the FYO point [0], and that is catastrophic for the vendor, and the vendor never sees it coming and can't help themselves.
Sitting on your laurels is not good. Don't do it. Innovate. Then innovate some more. Then never stop innovating.
[0] Let me google that for ya: https://www.google.com/search?q=fyo+pointAs a former Sun employee, I can tell you that's true: people at Sun wouldn't even look at competing technologies because they were sure there was nothing useful to be learned from them.
That's why Bill Gates and Anders Hejlsberg were able to screw Sun so badly by examining and copying the good things about Java, and making something much better: C# and CLR.
While Sun totally failed to learn from any of the good things that Microsoft or Apple or anyone else did.
>We are better off with all the wood behind one arrow.
Nice reference to the old Sun slogan and 1999 April Fools Day prank, in which Sun employees put an enormous arrow through Scott McNealy's office.
https://findery.com/johnfox/notes/all-the-wood-behind-one-ar...
All your wood notwithstanding, it also helps to chose an arrow that isn't flawed and doesn't totally miss the target. For what it's worth, Scott McNealy also put all his wood behind another arrow named Donald Trump.
Scott McNealy has long been one of Trump's few friends in Silicon Valley
https://www.sfchronicle.com/politics/article/Scott-McNealy-h...
Trump held fundraiser at former Sun CEO Scott McNealy’s Silicon Valley house on Tuesday
https://www.cnbc.com/2019/09/17/trump-silicon-valley-fundrai...
Scott McNealy gets touchy feely with Trump: Sun cofounder hosts hush-hush reelection fundraiser for President
https://www.theregister.com/2019/09/17/mcnealy_trump_fundrai...
I mean in practice it's tiny utilities that I'm too lazy to reimplement in the exact same way, but still, it's something to keep in mind.
Someone downloaded the code for GANs (generative adversarial networks) from Github, generated (sampled) a painting using it, slapped the GAN objective function as a signature, and sold the painting for over $400k!
While I don't believe in karma, I do believe there are consequences to misbehavior, even if one is not caught:
The temerity and lack of ethics to appropriate a project that someone else has written and claim it as one's own leaks into other areas of one's life. That kind of behavior is not isolated to just open-source software. Perhaps the world-weary thief of the OP took his experience as a lesson and changed his ways, but he likely continued bumbling around, behaving dishonestly, losing the respect of his peers along the way. Perhaps even that of family and friends. Perhaps he miscounts the points in a boardgame, or cheats on his spouse, but it won't be limited to this.
Contrast that to the life and career arc of the OP himself.
So, no, the fools are those who steal, caught or not.
That ghost user was an older account of mine. Their interpretation of the post was wrong though.
I reacted to the situation by laughing slightly but didn't even bother explaining why. Because of the poor tone, I decided to let that person enjoy being wrong.
MIT/BSD-style licenses are practically begging large billion-dollar corporations to rip off your work wholesale and use it to generate profits while contributing only the occasional patch or two. I used to see this phenomenon on HN, where every time some distro had a new release the BSD folks would be in here reminding us that BSD runs Netflix and routers and Playstations and won't we please just donate? As if Sony and Netflix value these projects enough to use them for critical infrastructure but not enough to keep them financially solvent.
(The GPL is of course not a panacea; as TFA demonstrates, Sun would have got away with this, possibly forever, had the author not made his serendipitous discovery)
GPLv3 is not really workable in terms of preserving developer freedom to do what they want with code (as long as they share their code back).
Obviously the powers that be disagreed and the GPL has been "upgraded" to v3, but was never impressed, and now much happier to contribute to MIT / BSD licensed products (which do allow you as the developer to do what you want with the code).
Would be curious if anyone has analyzed this.
With that said you’re right, with GPLv3 they moved the fight into new domains that aren’t as obvious or really “as big of a deal” to a majority of people. Also with the rise of things like JavaScript GNU, under Stallman, became the old man crying at the children. Stallman hated the rise of “non-trivial” JavaScript and refuses to work with proprietary code so GNU could never have developed e.g. React or Tensorflow. We now have a new generation of open source tooling that was developed more in spite of the FSF than by it.
The current environment will in turn spark a new generation of tooling down the line that is even more removed from the likes of Stallman; who once responded to a request I sent to work on a JS library advertised on the GNU website with the fact that I should call it F/LOSS instead of open source (and nothing else.)
So while I respect the FSF for the work they did in creating what we have today, the time where they were leading the fight to save software is long passed. They won in some ways and the world is better for it, but they lost in others and I’m not sure the world is that worse because of it. Having trade secrets isn’t completely a bad thing as it allows competition and different implementations, though that part is simply my opinion I suppose.
I suppose you're right though in as much as it splits mindshare.
The sample license header here [1] says:
> This program is free software; you can redistribute it and/or modify it under the terms of the GNU General Public License as published by the Free Software Foundation; either version 2 of the License, or (at your option) any later version.
That is only a suggested header.
From Linux kernel here [2], you can see:
> under the terms of the GNU General Public License version 2 only
[1] https://www.gnu.org/licenses/old-licenses/gpl-2.0.en.html
Remember that the idea of the GPL is the authors preserve and propagate the rights for the USERS.
He also once said "proprietary software subjugates people", which I thought was sort of over-the-top to say, but over time I think software in this era of dark patterns and privacy has unfortunately become very obvious.
Sadly a lot of people don't even respect that.
A company using open source software to make money is making money off commons. Makes sense they should feel obliged to contribute something back, and since they have the surplus of the best form of contribution - money - it's reasonable to expect them to donate some of it.
> it's reasonable to expect them to donate some of it.
No, that’s actually quite ridiculous. “Reasonable” implies some level of reasoning behind it. There is no “reasonable” proposal of how much money should be given when it’s against the very spirit of the license to expect payment based on usage.
What is the percentage amount of profit an individual or corporation should contribute? Give me a concrete calculation of software usage and how much should go back to the project. Is it measured in percentage of clock cycles spent executing that code across all of an entity’s compute?
Presumably the IRS should also give a cut of all tax revenue collected by the US to the open source projects it uses too, right? If not, your beef seems to be purely with private enterprise being successful more than any fairness based billing.
And no wonder why Sun failed to compete. A Cultural and management failure.
Aussie and NZ compared to the US are completely different worlds in my eyes.
(I'm a Kiwi btw)
I'm a Kiwi too - but living in Aus.
The obvious examples that come to mind would be sports. NZ excels in Sailing, Cricket and Rugby against much larger countries.
Though it could be that in general, Aus/NZ have the same skill distribution as other countries, but they just stand out more because there are lower expectations or highly skilled people are rarer overall due to lower populations.
Those are niche sports, very popular in that part of the world. They're not major sports, though.
Cricket is sort of a major sport, but it's also very culturally concentrated. Outside of UK & some former UK colonies, almost nobody plays it/watches it.
Does New Zealand have any famous footballers, basketball players, athletes, etc.?
I am French and we are a country with a great past - but not looking ahead. We still assume that we radiate across the world and that our voice counts.
It does not, and this has to be clear. There are countries that can influence events, but this is not ours.
A typical example is how angry we were after the recent pirate act of Lukashenko (Belarus forced a EU plane flying form one EU country to another to land when above their territory to arrest an opponent of theirs). Our president said that there would be consequences and there are no consequences.
This is not to spit on my contry but sometimes egos are much bigger than the reality. We are in good company there.
For Roman there is no way back after that. It's an atrocity, but ultimately Belarus authorities are responsible for those and the torturing.
Europe needs to ensure its planes are not hijacked mid-air like this in the future, to protect all of us and foremost journalists and activists.
If Brexit showed us anything it is that the UK is no way near as important as they like to think they are. Watching the UK and [redacted] at loggerheads has been a mix of frustration and amusement for someone like myself clearly caught in the crossfire.
As an "outsider" in [redacted] my biggest complaint is the staunch opposition to change. Any change. As you say they simply cannot look ahead. It is as if they only know how to live in their past glories rather than working towards future ones.
The UK is sort of the opposite but in a terribly executed manner. They have dreams of the future but do everything possible to make those dreams harder to achieve due to arrogance they can "do it alone". Harping back to "the good old days of the Empire" and "Blitz spirit!" as if the Blitz was some wonderful time (wtf?).
The sad thing is the [redacted] could learn a lot from each other but seems both sides are too myopic to do so.
I am French and we are a country with a great past - but not looking ahead.
I feel the same as you about both the British and the French.
Add India to the list.The right-wing in India are obsessed about our past, and worse, desperate to associate everything about it with "Hindu religion" or "Hindu culture" despite the huge influence of Buddhism, Islam and the imperialists (largely the British who finally got the upper hand on our sub-continent).
"We were the richest country in the world till we were looted by Muslims and Christians."
"We had brilliant Hindu brahmin scientists who excelled in mathematics, medicine and astronomy / astrology.
"Look at these huge ancient temples built with extraordinary artistry that have survived for centuries."
And so on, are proof of our "great past" that they believe automatically should earn us the respect of the world.
In their obsession with the past, they totally disregard the achievements of modern independent India just because our freedom movement were lead by people like Gandhi and Nehru who opposed their idea of a theocratic-fascist state, and chose to create a secular state that treated everyone as an equal and gave every citizen equal rights.
For a country that won its independence in a non-violent manner from the most powerful empire of the world, and a country that has lifted millions of its citizens from poverty and is today self sufficient in agriculture (one of the largest in the world), and one of the few countries with an active and self-sufficient space, nuclear and defence program we have really made a lot of strides.
But for the right, India is not "respected" by the world because anti-Hindus chose secularism and thus since "we Hindus" don't respect our "Hinduness", the world also ignores Hindu cultural achievement and denies us our true place.
France is a major actor in African and regional conflicts.
It is the third largest exporter of weapons ( After the USA and Russia ) to conflict zones.
It was the case up to the 2000's, I think, but I can't even remember the name of a new French director since then. The last one I remember is Luc Besson. Kind of similar story for actors/actresses, are there some major French stars popular across the world?
I feel most of the influences are leftovers from a different era, folks like Depardieu.
TV/film - is there something new even remotely popular across the world that's French?
Indian here who doesn't know French - just finished watching all four seasons of Dix Pour Cent (Call My Agent) and really enjoyed it. I also watched Lupin (after learning that it stars Omar Sy who I loved in the movie The Intouchables).Zone Blanche Lupin L'ascension Le grand bain
I'd group countries in 4 categories:
1. Continent sized (Russia, Canada, China, US, Brazil, Australia).
2. Large countries: anything below 1 up to probably about 500k sqkm.
3. Middle of the pack countries: everything from 500k down to about 100-200k sqkm.
4. Small countries. Everything below 100k sqkm (200k sqkm if you want to stretch it out).
The categories are somewhat fluid since for example Indonesia is not continent sized in landmass, but it's an archipelago that does stretch over the area of an entire continent, when you consider it end-to-end and include its territorial waters. Plus having a very high population for your group also moves you up. Germany is average in size but in population it's a large country. Same for Japan.
Swings and roundabouts.
P.S: if you constantly boast about how humble your country is, you're probably not that humble.
He mentioned this "VIP" is a "Developer and dtrace expert". But reading that and the other details, I think this is probably not the reality and maybe was communicated incorrectly to him. I really doubt this guy was a "VIP" as he says.
My guess is this "VIP" was actually a pretty normal member on the dtrace project, could be a little senior and got the opportunity to go around and talk about it. I am sure they had a team somewhere who put together most of the software, maybe he was involved a little bit, but probably he was just as confused as everyone else about using that open source software - he probably knew enough to teach it, and how it worked, but so many people work on these type of projects, unless they sent the lead engineer he probably didn't know it deeply except enough to evangelize and teach how it works.
He mentions about being slighted by this guy a lot, saying things like "He wasn't impressed", "gave me a look like he didn't really believe me" etc. This might be true, but i suspect it's coming from his negative interpretation of the situation. This guy just traveled all the way around the world, was super exhausted, was possibly honestly confused what's going on - i certainly have been in that situation before.
The author also mentions he felt it odd that he (the author) was producing more dtrace tools than Sun was. This almost sounds a bit like indirect boasting. Large companies are slow. A dedicated passionate developer who is working alone or with a small team will always run laps around huge companies. This isn't odd at all. Companies often get distracted, can't focus on what's important, or decide not to do what is important for a product due to other business reasons.
In fact, as he found out, some engineer somewhere just ripped his stuff cause it was faster and easier for them to do it. Sun's team was not professional at all, even possibly breaking the law, which I think is the point of the article but the descriptions of the Dtrace guy who's job was to show Dtrace around the world lessened my enjoyment of the article.
- DTrace is the new hotness, we need it in our UI.
- Everyone's using Brendan's tools, let's add them (so far, so good).
- Oh, why do they say copyright Brendan? He made a mistake: Sun employees should be putting copyright Sun on them. (THIS is the mistake, as I wasn't a Sun employee).
- I'll just delete his name and stick copyright Sun on them all.
- Developer gets picked to go do a world tour (and may genuinely not know what happened).
As for how I was treated: I guessed why in the article as well, the low-key introduction as is the norm in Australia.That's really interesting, since you were so close to Sun they actually thought you were a Sun employee!
Large companies are slow, indeed. In the late 90s I wrote a few operating system plugins (nss_ldap, pam_ldap, GSS SASL plugin for the Netscape directory server) which were eventually obsoleted by native Solaris equivalents. The Sun versions were on the whole better engineered, if less flexible, because their OS team had a depth of experience that I didn't have at the time.
[0] https://www.usenix.org/legacy/publications/library/proceedin...
Brendan was the most amazing and prolific user of DTrace, from very early on. Brendan did not create DTrace, but in a sense he "made DTrace" what it is. And not just DTrace, but eBPF.