I'm always surprised at the fact that almost every product has more lines of code than the entire Linux repo. The scale of these products is astounding.
I'm always surprised at the fact that almost every product has more lines of code than the entire Linux repo. The scale of these products is astounding.
For these translation files, I’d imagine there may be occasional work to modify them even after they are initially generated.
Or, in more words: The format of the files is just the representation on disk - it’s not directly connected to how the files are generated or edited. XML files can be written by hand with suitable editor support.
I think that this is not a sensible definition of a generated file. A more sensible definition is that a generated file is created automatically from some source, which is not user input (i.e. an other file). This means generated files do not need to be kept under git, as long as their source is checked in.
Translations files, even if they are not created with a plain text editor but with some other tool that handles the XML layer, are clearly not generated, as long as the translation is done by a human.
This is a very typical workflow. Most people are not out there modifying xlf files by opening them in a text editor. For a start, translations usually aren't done by developers.
(Huge shoutout to Lokalise btw. I can highly recommend it. It makes building a multi-lingual app across different platforms so much easier.)
You opted to keep translation files out of version control; you could also keep images there, or source files. All this stuff is the (pretty direct) output of non-deterministic human intervention.
(BTW, how do you build an old version of your application? Is lokalise able to give you the appropriate translations for a specific git commit / app version?)
There is also '-merge', which will cause git to not attempt to merge the contents, but just ask you to pick a side.
The challenge however is then verifying the contents of these files in things like merge requests.
Is there a reason why that type of file couldn't be better place into an artifact repository, or just generated and consumed in CI as part of generating a final build output?
This adds yet another moving part to the system, and another place things can go wrong.
> generated and consumed in CI as part of generating a final build output
This can get quite slow, and on larger projects you have to expend a lot of effort to keep build times reasonable.
Also, if you're serving a library for public consumption, you generally don't want to add the burden of extra build steps for the user to follow before they can use it. If it can all be automated to the point of invisibility to the user that's fine, but often it can't.
This is not surprising at all. In fact, it's quite standard to commit string translations. Just because you can run the code generation/string replacement step as part of the build that does not mean it's a good idea to generate everything from scratch at every single build.
String translations hardly change once they are introduced, running the build step takes significant amounts of time, and if anything fails then your product can break in critical and hard to notice ways.
Additionally, to respond to your comment, if string translations don't change much then it may be possible to push them out as an internal 3rd-party library, and then they're even quicker to build.
You're missing the point. Storing translated files is caching things in git, and it is not a bad idea. It's a standard practice that saves your neck.
You either place faith on a build step working deterministically when it was not designed to work like that, or you track your generated files in your version control system.
If you decide to put faith on your ability to run deterministic builds with a potentially non-deterministic system, you waste minutes with each build regenerating files that you could very well have checked out and in the process risk sneaking in hard to track bugs. Then you need to have internationalization test steps for each localization running as part of your integration tests to verify if your build worked, which consume even more resources.
Or... you stash them in git?
You use git to track changes, regardless of where they came from. Just because you place faith in some build step to always work deterministically that does not mean you are following a good practice and everyone else around you is wrong.
If one has to be more granular than that, and have versioning and verification against the repository, they can still store the multiple versions on another service and store the hashes on git. Even though I'm not a fan of this for translation (especially if you have lots of languages/lots of strings), since there's an advantage of decoupling the translation process from the development process.
The problem with storing those files on git is that it can cause more problems, including developer experience issues.
It depends on how much you're storing on git. Some CSS files? Fine. 70% of files of the project, like in this case, slowing down everyone's workflow? Definitely not.
I'm sorry, what? Why would a build not work deterministically?
> If you decide to put faith on your ability to run deterministic builds with a potentially non-deterministic system
If your build is non-deterministic, how can you have any faith in the binaries it produces? You would have much larger problems in that case.
> You use git to track changes, regardless of where they came from
You probably don't want to do that if it is 70% of your codebase and slows down all your developer's git.
> Then you need to have internationalization test steps for each localization running as part of your integration tests to verify if your build worked
I'm convinced you've never used a build system before. Your build should fail if required files are missing. Downloading translation files at build time from some artefact repository vs storing them in git is how a lot of companies do it.
Because they don't and never did?
Do you understand build systems and individual tools were not designed to ensure deterministic behavior?
https://reproducible-builds.org/docs/deterministic-build-sys...
Anyone with any professional experience developing software can tell you countless war stories involving bugs that popped up when building the exact same project separate times. What leads you to believe that translations are any different? In fact, more often than not we see unexpected changes during translation update steps.
> If your build is non-deterministic, how can you have any faith in the binaries it produces?
First of all, all builds are not deterministic by default.
To start to come close to get a deterministic build, you need to do all your own legwork after doing all your homework.
Did you ever did any sort of this work? You didn't, didn't you? You're not looking and are instead just placing blind faith on stuff continuing to work by coincidence, aren't you?
> You probably don't want to do that (...)
Yes, I do. Anyone with their head on their shoulders wants to do that. It's either that or waste time tracking bugs that you allowed to go to production. Do you want to waste your time hunting down easily avoidable and hard to track bugs? Most of the professional world doesn't.
You're also doing that everywhere else. How do you think anything works? Why do you think Git is deterministic somehow? Why more so than including some files in a build?
No reason at all, but when you need the files during development, and testing, and CI, and in production, and you don't want those things to fail when your artefact repo or source of data is down, then putting the latest versions in git makes sense.
The cost of having them in the repo is a tiny bit more complexity in your git workflow and config. The benefit is being able to access those files everywhere you access the code. It seems like a no-brainer to me.
You're using git as a cache. You don't need to version a cache.
If it was a physical product, you couldn't keep making it bigger and more complex ad infinitum, because making a physical thing bigger takes more material, and bounded physical resources would be consumed. With software, it's all just bits, and computers can hold a lot of bits.
This leads to bigger problems than just git running slowly.