A project with a single 11,000-line code file
austinhenley.com
austinhenley.com
I guess the moral is, never underestimate what determination and creativity can do, and always be skeptical when someone says there's only one best way to do something.
Best practices were the usual suspects DRY, IOC, SQL + NoSQL, separation of concerns, config files over code, composition over inheritance, unexplainable overlapping annotations, dozens of oversimplified components doing their own thing, and some $something_someone_read_on_a_medium_post
The Single Java File was around 500 lines no db, lots of globals, a dozen or so classes and some interfaces, Threads for simulating event based concurrency, generous use of Java queues and stacks but i specifically made it static with Zero dynamic hashmaps.
It actually runs in my IDE, I can understand what the hell the product is supposed to do what component is doing more than it should and more valuable was to predict what could break if I change that value in the helm chart from 5.0 to 5.1.
It is quite useful and pleasing, I can actually reason about things and I have new found use and appreciation for Type Systems and compile errors. And I can write tests that run in under 3seconds.
These best practices really only make sense in large organizations, i.e. Conways law.
After all, you can't really ask 100 developers to all add code to one file in a couple weeks - they will spend a month or so just resolving conflicts...
100 file repos are designed so that 100 developers can edit them (in theory), and have relatively few conflicts, not because its better code.
As another anecdote, I find whenever I do solo code I can easily spin out thousands of lines of code within a week (includes testing), when doing code on a project with many other devs my rate drops into the range of maybe 200 a week, just because so much time is spent interlinking other code, finding fixing bugs and tests strewn across many files...
(HN has indentation, though.)
It’s important to realize that this is good design. It’s hard to separate yourself from the time you live in, but the rewards are worthwhile.
(def add1 (x|int)
(+ x 1))
http://www.paulgraham.com/bel.htmlI've been implementing it for a couple years now, though not seriously till the past couple months. There are some interesting (and overlooked) ideas in Bel.
Bel is sort of the limit case of generality. For example, you might expect the "type" above to be a separate kind of object, the way that types are separate kinds of things in TypeScript.
But in fact, it's simply a function that receives the argument and can throw an error. So for example, you can do something like:
(def positive (x)
(if (< x 0) (err 'negative) x))
(def sqrt (x|positive)
...)
I just wish he'd solved keyword arguments as thoroughly as every other kind of argument. There are hints that it was always in the back of his mind. Though it's true he never needed them, so that's probably why he never made them.[1] https://docs.racket-lang.org/ts-reference/Experimental_Featu...
There's a big difference between "code being in one file" and "code being in one function." It sounds like the OP had something reasonably close to "one function," whereas the HN code has a lot of (what appear to be) small well designed methods.
Which IDE do you like? I'll see about getting some highlighting for it.
As for actually running arc, it’s hard to run the original arc3.1 due to racket updates. I’ve made a few branches over the years that try to preserve the original spirit of arc (no significant changes) while making it easy to run. Try this one:
https://github.com/tensorfork/tlarc
I believe you can simply install racket, then run make && bin/arc and be dropped into a repl. From there you can follow the arc tutorial, whose link I’ll dig up after I’m finished driving home.
EDIT: That fork is actually a lot different from arc3.1 proper. I’ll try to locate a more faithful one.
EDIT 2: Unfortunately it's quite a lot of work to remove the mzscheme dependency from the old arc3.1 codebase. And I'm not sure it's even possible to install the mzscheme lib on the latest racket (e.g. `brew install racket` doesn't seem to have it).
So the above instructions are the best I can do for now.
I like how the "html" and "css" part was embedded in that "news.arc" file. Do you think that VIM script will highlight and lint the "css" part of an "arc" file?
I try to use Arc for as much as possible. We wrote our TPU monitoring software in it: http://tensorfork.com/tpus
Eventually I became frustrated with Racket's FFI. So I eventually made my own arclike language called elflang: https://github.com/elflang/elf
... which itself is a fork of Lumen (https://github.com/sctb/lumen) by Scott Bell.
The performance is good enough to run a minecraft-style game engine: https://i.imgur.com/iyr0YrB.png which was satisfying.
Nowadays I've been trying to implement Bel, mostly for the challenge of it than for any practical reason.
> I like how the "html" and "css" part was embedded in that "news.arc" file. Do you think that VIM script will highlight and lint the "css" part of an "arc" file?
Nope. https://i.imgur.com/o9aUG6j.png
But it has one very important feature: it can properly highlight atstrings: https://i.imgur.com/wO4f742.png
It's probably hard to tell, but the "@(hexrep border-color*)" would normally be highlighted as if it were a string. Arc has a feature called atstrings, where you can use @foo to reference the enclosing variable "foo". It can also call functions, e.g. "The value of 1 plus 2 is @(+ 1 2)" will become "The value of 1 plus 2 is 3".
That application would likely fall apart if multiple developers of with diverse backgrounds had to maintain it and add new features.
[1] https://docs.microsoft.com/en-us/previous-versions/visualstu...
Like most of the world at one point, it ran on Excel 2003
/s
Ya, you could just go in there and mess with the code all over the place.
For a while, I was the only developer working in a small module of a bigger project: I started the code base, discussed requirements with clients, implemented the needed features, tested the whole product end to end. I developed a very good instinct about it and about what any change would do, much better than any other project. My theory is that the code base matched my way of thinking, so thinking about it was pretty easy.
The other thing is style - when used well, state machines don't require testing. There's nothing to test - either your machine works or it doesn't, there is no point testing state transitions because that is the fundamental job of the state machine. You may as well test that addition works.
Ofc, they must be used well - problems may be difficult or overly complex to model as a state machine or even set of state machines - the pattern excels for small problems, less for large ones.
When working on projects by myself, I like putting everything in one big file too. Trying to find the "right" place for something is some unnecessary overhead, not to mention the navigation cost. It's a different story when a team is involved though.
Suffice it to say, it's a 28k LOC file that was so bad, it could even hold up in court as evidence that a South American company stole the code of Zynga's -ville games. We could reproduce each and every single bug and its effects 1:1 in their games, with all the crashing scenarios that were easy to reproduce, hard to debug, and almost impossible to fix.
Once you dig into the hole of depth sorting and being smart by "just slicing" everything into squared ground tiles on the fly, there's no way out of that spaghetti code anymore.
Fun times, was always a joy seeing people give up to a single code file. The first step to enlightenment was always resignation :)
Learned repeatedly from painful experience.
A lot of edge cases and race conditions would easily slip through, also a different set of edge cases or race conditions you never considered and therefore never tested for in your first version could pop up in your rewrite.
I've dealt with concurrency issues. grep is a handy tool to find related synchronization code, then I try to replace it with an encapsulation. In general, I look for things I can replace with algorithms, and things I can encapsulate. And so on.
Almost nothing is impossible to test yes, however to know and be able to mock the data for each test case can be extremely hard and at some point not worth the effort to even attempt.
Most I have seen these kind of systems doing is statistical testing with reference benchmark/ sample data, and maybe monitor real world feedback either telemetry or user complaints.
Nah, most ML systems (actually doing something in the world) are mostly just ordinary code, which can be tested like any other code (as you put it into functions, etc). The models themselves are pretty awkward, but you can normally freeze the model and just use that to ensure that things stay working while you refactor, and then re-run a few (10+) times to check coverage and intervals and stuff.
It tends to be more difficult, as many DS/data people are not software engineering focused, but it's not impossible.
The model is what gets continually updated and is the critical path that needs coverae, Testing interfaces are trivial and at times not critical to test if already running in production for a while (you probably have already caught most/all issues and know what to test or take care of in a interface rewrite).
It is not about impossible, here is an example, let's say you are working on English speech-to-text model, the next version works better in your set of benchmarks.
It could for example perform very poorly (compared to your previous model) for accented English or mixed with other languages, for older people or in noisy environments like a car, or for for specific subjects like medical/legal dictation and so on and since your benchmarks originally didn't cover these types of scenarios you wouldn't know one way or another.
These were real cases all added to speech-to-text models after user feedback and adequate demand being identified and research effort put in, and now training/benchmark data includes these. There are plenty of scenarios not yet solved (mixing two languages is active area of research) or not included because user feedback didn't capture it, of not yet worth solving.
Neural network testing is hard because by design they have millions(and these days billions) of parameters as inputs and you cannot feasibly test every possible outcome, you will not know what all things to check until people start using your app in ways you never thought off.
NN /ML is not hard requirement this is true for any complex systems. Shazam type fingerprinting for example is just spectrography and Fourier transforms, NN is just newest tool devs use. All complex systems with thousands and above parameters have same problems
And if there is no test suite, there is often very few ways to add a test suite. Poor code has very few points where you can attach a test. If the code contains file databases or structured input of some sort (a web page) you can add some very high level end to end tests. But not all code has easily verifiable endpoints like that. Perhaps my bad experiences comes from "hard to test" domains (Sound, drawing, ...) code, and not "given this input this is written to the database and this is written to screen".
This cannot be done with a file containing 28k lines of code. That is an insurmountable task. They may as well have been asked to start from scratch and build a new engine.
I'm curious what the purpose of this ritual was. Was it just hazing, or was the thought that someone might actually be able to accomplish this?
I agree with you, a rewrite is probably how they should have tackled this one.
What makes testing hard is a lot of side effects, like the program writing to databases or calling external API's. Not the LOC of the source file. You might want to mock an automatic call to the fire department for a fire control program, but for API calls and databases, just have the test write to the prod environment, but include rollback/cleanup in the test. That way you don't need a separate testing environment.
I disagree with that. Automated testing should be done on a test database.
Holding a lock for too long could effectively block the entire production. This could happen while debugging through a test (e.g. by hitting a breakpoint or by just stepping line-by-line). Or by simply having a bug in the test causing a "transaction leak" - depending on the tech stack, this could keep the transaction and the associated locks alive until all tests have finished running, not just the one which had the bug.
Or you could commit instead of rollback by mistake.
Or you could simply put unexpected strain on the database, affecting the performance of real users.
Actually no. Most devs that just got started familarizing themselves with the codebase wanted to refactor the file and came up with the idea themselves. Usually they thought this is a crappy file and this must be an easy task to do because they saw all the nested if/elseif/else statements in the code.
The problem, architecture-wise, was that the road logic was the glue code that integrated a lot of different parts, layers, and NPC behaviours from the rest of the codebase as it was changing the surrounding game world.
If there was a hospital placed with a non-squared ground tile next to it, if it was placed with a 1 offset (roads were 3x3 tiles), if it was placed with a 2 offset next to another road... It went as far as influencing the path heatmap that was necessary for the A* guessing algorithm to make the NPCs walk correctly on the sidewalks. The permutations of possible sidewalks alone were enough complexity on their own...
So in a lot of ways necessary features that historically had no place in the Entity/Component based engine at some point made it in there.
The next best thing (and also spaghetti code) was the Cursor Entity, which had to have line tracing algorithms to be able to select things that are visible under a donut-like shape when the user was hovering the hole, or say, a tree in the game world. Convex and Concave shapes were integrated, and lots of edge cases in there, too, which are actually huge mathematical problems in terms of available performance once you dig more into it, so we ended up with binary height map sprites that helped both the slicing and the cursor at some point.
The important lessons learned from the road logic were very valueable for newbies, as it was teaching the practical problems of isometric game worlds.
So afterwards everyone was able to grasp why the complexity was added, and what was necessary to remove it (in the sprints in the future).
At some point we decided a couple of things because of the road logic and cursor entity for new iterations of the engine, like:
- always use a 1x1 road tile
- always use square based tiles for all objects
- dont make sidewalks, use just road tiles
- dont make trees with holes in their leaves
- dont make trees higher than the buildings
- no artist can ever request crosswalks. Never ever.
...etc
Always tricky though when the hacks have both undefined features AND bugs.
[0] http://dtrace.org/blogs/eschrock/2004/07/01/real-life-obfusc...
This thread made me realize that it's better to have a working profitable project with bad code, than a perfect unfinished project, with meticulously chosen design patterns.
Afraid of being judged for bad code, I could not start until I had the right architecture.
I'm glad I read this.
This is developers therapy.
apparently jwz decided not to be linked from here :/
there's an archive.org link below.
NOT SFW! https://cdn.jwz.org/images/2016/hn.png
It's funny. Why does he hate us so?
For future reference, a non-archive, non-jwz.org link. Straight from the source as that's the author's own site.
Then just copy & paste it.
* Team A writes code quickly. Not bad code, really, but they take shortcuts everywhere they can. They don't have the strongest tests, they don't generalize for all the known use cases, etc. Their code goes to beta and gets users and makes progress.
* Team B deliberates and deliberates. They try to avoid taking shortcuts. But in the end, even their code doesn't have the strongest tests, doesn't generalize for all the known use cases, etc. Team B never gets users or gains momentum, and their code+architecture was probably no better than Team A -- they just took 3x the time to get there.
I had a lot of trouble trying to explain this to juniors.
The most important things is to have code that is easy to refactor when you know what you're doing (i.e. everything is working properly). Juniors I worked with had a nasty definition of a pretty code being split into a hundred files, each no longer than a screen, and each function no longer than 5 lines. The onboarding of new devs to such code was way worse than into a code that would be 10k lines in one file, but with a flat structure and less interdependency.
Very true.
I'll just add that another most important thing is to actually take time to refactor, even when things are busy.
I spend maybe 1/3 of my time refactoring, and that feels good.
Once you know it works take a minute to clean up and make the changes fit your preferred style, extract repeated code into shared methods, comment the tricky bits, etc.
I had an "everything should be broken into a hierarchy!" stage back when I was learning to code, and boy was I off track. In my defense, at the time (and this dates me) OOP was all the rage.
Procedural spaghetti is more manageable, though I once had to update a C app with a 3000 line case statement. Pure madness.
Golang does a good job of addressing this one particular problem.
In a subsequent project we banned subclassing of exceptions.
But... isn't the easiest way to show that there is little interdepency to put them in separate files that don't import from each other?
This whole post is about how refactoring doesn't matter because your project's development lifetime isn't long and wide enough for maintentance to matter.
It makes debugging in Java or C# an exercise in face shredding frustration. Where each class is in a file, it's better to structure the class consistently with other classes in the project, and things like naming conventions become a lot more important.
You could argue that the C/C++/Java/C# languages are fundamentally broken because they don't encourage the succinct, small class methods that Smalltalk did, but you could also argue that those small methods don't necessarily work very well in a different class of languages from Smalltalk and neither approach is really more productive than the other, with the caveat that Smalltalk is largely dead and irrelevent to modern programmers other than a curio.
Just starting and doing it is just unreasonably effective because very few projects actually need novel solutions - most are just fine with off-the-shelf hacked together solutions.
Thinkers are required if the software is actually groundbreaking new work. Almost everyone's work on this forum probably isn't that however (mine included), which is why I agree with your sentiment
This is the dumbest part. You’d think that someone documented something when they originally built it, but nooo. Don’t even know the requirements, just that it has to be the same as the previous one.
It sounds like your CTO did not just operate quickly, but also sloppy and chaotic. From what I am gathering from this thread, the best practice is to move quickly AND organized such that refactoring is reasonable.
Team C works like team A. However every time a feature ships, someone who knows that feature well immediately refactors the relevant code to remove the prototype scaffolding. When code becomes static, an expert adds good quality comments. When a bug is found, it is recreated in a regression test prior to being fixed for good.
Hack at the code for 4 years, collect your options and leave the mess to someone else.
To be honest, a hacky codebase written fast is not the worst codebase to deal with. The worst type is when someone had the time to overarchitect and overengineer things.
Following references across 200 different files, tracing calls through hundreds of microservices. Graphql servers with complex resolver logic.
I concur. Refactoring should be as much about removing unneeded abstraction and features as it is about adding same.
> there is little incentive or bonuses for improving the situation.
Yeah, I just can't seem to believe this in my soul. I just want to fix ugly code and can't stop myself. I get huge satisfaction from speeding up, tidying up or fixing up bad code.
It doesn't help when management wants to minimize time spent on such tidy-up, especially when it's hurting our productivity to maintain it without fixing it.
1. Figure out what needs doing and what code you need to use
2. Refactor the code you're using until the new code or change is easy
3. Make the change
4. Tidy and document.
Repeat
I often also document during step 1, while I am trying to understand code and realise that comments are missing.
The beginner uses simple chords
The intermediate uses advanced chords, crazy fills and runs and riffs.
The advanced uses simple chords
Intermediate devs write code that is over-engineered in places where it could be simple, while still needing a lot of extensive refactoring in the parts that deal with irreducible complexity.
Expert devs think deeply about the problem at hand, understand where the complexity is, and create a solution that is as simple as possible, but no simpler. After the fact, beginner and intermediate devs tend to think of this code as something trivial, that anyone could do.
This is why it may sometimes be tactically useful for the expert dev to sometimes take on projects that some beginner or intermediate team has been struggling with for some extended period of time, analyze it properly, and show how elegant the difficult parts can be solved.
Care is needed, though, as it can affect the morale of the other devs and even cause hostility. If the expert dev has a secure position in the organization and want to keep those other devs around, keeping a low profile as well as letting those devs take as much as possible of the credit is adviced.
Also when someone from Team A initiate the code and Team B takes over, there were few times that Team B feels like the code is damn no good and just massively refactor it with what they think is good (re: the boilerplate) without others' concern. Then when Team B leaves, it goes to 1st paragraph..
I think as a team we need to consider the learning curve of our own code cz we don't code for ourselves.. And it's good to know the tolerance & acceptance to 'structure' of the code from other people..
The new shiny services took much longer to identify bugs and add new features due to the complexity of the design and endless interfaces.
I was so afraid to cook in a dirty kitchen, I ended up not cooking at all.
This thread made me realize that's better to sell food prepared on dirty surfaces with unrefrigerated ingredientes half-eaten by rodents and roaches that makes people sick, than fresh food prepared on clean surfaces with clean utensils.
I'm glad I read this.
This is a restaurant worker story.
Construction industry version: I was so afraid of not using the right construction materials and not building code-compliant structures, I ended up not building at all.
This thread made me realize that's better to sell houses with structural problems and low quality materials that will be unsafe to live in, than houses built according to code.
I'm glad I read this.
This is a builder story.
In any other industry, a person would go to jail for saying that. You won't, because luckily for you, software development is not a regulated activity, and people with your mindset can make a happy living outside of jail. But hopefully one day some types of neglect in software development become illegal."Better is the enemy of the worse" is no excuse to have spaghetti code, or 50,000 lines of code files. It means that good is sometimes more convenient than perfect. Spaghetti code is not good to begin with.
Using bad ingredients in food, or poor quality materials in construction has tremendous impact on the final product.
>"Better is the enemy of the worse" is no excuse to have spaghetti code, or 50,000 lines of code files. It means that good is sometimes more convenient than perfect. Spaghetti code is not good to begin with.
Just calling code good or bad doesn't mean much - ultimately, results matter - if your code doesn't have tons of bugs, if your team can add features without any problems, if you can ship reasonably on time, if your product delivers value to the end user, etc - then you have succeeded. It doesn't matter what outsiders think about the code or what labels people give. Its best to ignore them and continue doing good work.
Failure is cheap in our industry. That's largely a very good thing.
Also all of your examples have wildly different impacts than a dev "portfolio project". They all cause physical harm to people, which a poorly coded website/cli tool/etc almost certainly won't. Unless this person's hobby is writing code for MRI machines, in which case, go ahead and make everything is perfect, but that doesn't seem to be the case here.
Dismissing basic development good practices as "perfectionism" is just gaslighting people into believing that any form of thinking is overengineering.
Ok, so let's continue this analogy on the other one.
It's less like basic hygiene, and more like refusing to cook outside a clean room.
> Dismissing basic development good practices as "perfectionism" is just gaslighting people into believing that any form of thinking is overengineering.
Basic development good practices are something you develop during the "portfolio building phase", not before.
Your entire point rests on a baseless assumption. There's absolutely nothing in the parent's post that indicates that the programs he would create could have the potential to harm humans.
That is a really sad statement that is predicated on the assumption that nothing can be objectively compared and therefore nothing can be ranked, which is also a way to kill arguments that lead to innovation and iterative improvement.
You can measure the complexity of an algorithm, you can measure the cyclomatic complexity of a function, you can measure code in terms of length, you can count references to external functions or modules, etc.
There are many ways in which you can compare code and make decisions about what style is more convenient for your team.
What is clearer for you to understand?
a) 1 + 1 + 1 + 1 + 1 + 1 + 1 + 1 + 1 + 1 + 1 + 1
b) 12
If your argument was true, they would both be the same. We both know that a) is a waste of time.
If you take any style guide, book about good practice, or stuff like that, you'll find that there are some good ideas, and there are some bad ideas. Even with something as simple as code formatting, we still don't know as an industry if it's better to format everything the same, or use formatting to convey information. The debates about OO, FP, static/dynamic typing are endless and there's no evidence about which is better. There is a rough idea that you need more organization when more people work on code (static types, microservices, more documentation, more isolated parts), but even that isn't really clear.
My assumption is not that "nothing can be objectively compared and therefore nothing can be ranked". It's not even an assumption, it's what I've seen in this industry: we lack data to give objective good practice that go beyond anything trivial ("try to make code easy to read and simple", "give meaningful names to your variables"). There is a wide gap between "common sense" and "cargo cults", which is the gap between easy and complex topics. I would really like if in the last 50 years we had learn a lot about how to build software as an industry, as people. The reality is that we haven't.
A lot of businesses were built on PHP this way
But ultimately, Maddy Thorson isn't selling a block of code. They're selling a game and it has extremely satisfying control of the character. And that's all that really matters for a player controller.
Maybe better organization and design patterns could have made it faster to develop? But I don't believe it would.
But also the type of product does matter for this. Celeste had 2 programmers so a lot of the things necessary for a team of 100 devs would just be harmful. If you're making a library/framework to be used and modded by others, architecture matters a lot more. If you're designing an enterprise application that you know will need business logic customizations for 25 business customers it matters more. It's all about knowing the scope of your project. But also until you start getting that many customers, maybe the unsustainable method is what will allow you to reach those first few sales more quickly to be able to stay in business long enough to be bit by the technical debt.
Laster on Noel Berry did give a response explaining the various design choices behind their code:
https://github.com/NoelFB/Celeste/tree/master/Source/Player
Anyways, kudos to the team sharing their code even if it's a bit messy.
I'm so afraid of creating programs in languages that don't enforce a structure at all, even though I know how to write everything from scratch make it work.
If it's some framework, then it'll already be structured somewhat.
In the rare event that I do create something with no frameworks, I ensure that there aren't much global variables.
So very many editors just couldn't open it.
Some would use so much memory that the system would either freeze, or the OS would kill them.
Some would silently truncate at 65,535 lines.
Some would produce a load error.
Some would pop up with an error indicating the developer thought it was an unreachable state. e.g. "If you're seeing this error... Call me. And tell me how the fuck you broke it."
Others would manage to open it, but were completely unuseable. Where moving the cursor would take literal minutes.
There were exactly three editors I found at the time that worked (none of which were graphical editors). And they worked without any increased latency, letting you know that the developers just thought through what they were doing: vim, emacs, nano.
(A few details because people are probably curious - the vast majority of that single file project was taxation formulae. It was National Mutual's repository of all tax calculations for every jurisdiction that they worked in, internationally, for the entire several hundred years of the company. They just transcribed all their tax calculation records into COBOL.)
Vim, on the other hand, does it splendidly: it keeps only a chunk of the text in memory, and iirc the ‘swap file’ that it creates for every opened file, keeps the changes in some kind of a sparse structure, so they can be tracked at various places in the original file. This ‘swap file’ also serves as a savepoint of the editing session, so the changes can still be recovered even if the machine crashes while the user never saves.
Alas, editors still tend to deal badly with very long lines (just in the low thousands characters). IIRC both Emacs and Vim drop into a big think if the user attempts to put the cursor further down that line.
Yeah, essentially this (apparently) mostly occurs when the file has no newlines (like json often does). I think the hacks are around turning off font-lock mode and one or two other things (install long-lines-mode if this is a problem you're having).
This is probably the most current implementation of a ‘view large files’ package: https://github.com/m00natic/vlfi
But don't forget society itself if governed by OOM larger bits of text with no referential integrity, no machine to tell you if it's inconsistent, and no way to test anything, other than making humans write more text to each other and occasionally show up in court. The law itself, even parts of it like the tax code, and regulations on various areas, are a melange of text and cultural understandings between lawyers, judges and government. We collect the data for this machine in the form of contracts and receipts, and it piles up in mountains.
As with code, it's not just legal professionals who have to deal with law. It spills into everyone's life, and there's nothing to do about it other than either guess what to do or pay a pro to tell you what to do.
I hate these words as I type them but the law is also "agile" (ugh). It gets modified as it's used. It does not need high-assurance machine-verified "referential integrity". In my entire course of studying the law I don't think I've seen a single legal dispute over a problem of referential integrity. Mistakes, especially drafting mistakes, are corrected on the fly pretty much everywhere they appear, and then they disappear. For a dev, using the wrong variable name in a bad language could mean you introduce a huge security vulnerability and massive loss of trust. (Or if you write smart contracts, $100M down the drain.) For lawyers, referring to the wrong section has essentially zero consequences. Nobody cares. Maybe you get a funny look from a senior.
Finally re the 10k LOC tangent that this is supposed to be connected to, I'm not really sure what you're complaining about. You get "10kLOC" cases, but you also get well-organised practice guides & bench books. Laws in statute are typically very well organised, in my experience about 5-10x better than the average codebase. Laws are organising large swaths of the sum total of human endeavour, just as code does. I would say developers are behind overall, which makes sense for a discipline that's less than a century old.
I attacked it by printing out and taping together each program into "scrolls" and tracing control flow with highlighters and sharpies. Had them all taped up on my office wall so I could refactor the whole thing from scratch, coworkers found that entertaining. Got a much more readable replacement working nicely. Then a couple years later HR bought a new system and we stopped printing our own checks. I was not sorry to see the whole thing go.
Here's a TV Tropes article on the cliche: https://tvtropes.org/pmwiki/pmwiki.php/Main/StringTheory
IMO, organizing source code in files seems archaic. E.g. tracing the history of a function moved across files can be tedious even with the best tools. I’d like to see more discussion around different types of source storage abstraction.
There are benefits of large source files... When compiling hundreds of thousands of files (like Office), the overhead of spawning a compiler process, re-parsing or deserializing headers, and orchestrating the build is non-trivial. Splitting large files into smaller ones could add hours of just overhead to the full build time.
For example in a video game your player, monsters, health potions, and attacks could all have the code for HealthComponent as a part of their virtual file. And updating the HealthComponent would affect the raw file so the virtual files would have the updates automatically. Yeah you can open dozens of editor tabs or always use jump-to-definition, but just being able to scroll around or ctrl+f within a restricted set of limited files would be nice.
I mean something that would dump them all into, one contiguous "file" from the editor's perspective. Included components (I don't want to say classes because it could work in a functional language or maybe you just don't include full classes but certain methods) wouldn't have to be coupled to the main class you're editing. Like if you have a decoupled event system you could pull in just the events relevant to the idea you're working on. You could have different views depending on what idea you're working on and save them as their own file.
To use the gamedev analogy again you could have MonsterCombat.view MonsterAI.view MonsterAnimations.view which would all expose different subsets of the Monster class and various related methods from other classes/modules.
Like you could have all of the code from class A except the debug related methods, a few methods from class B, just a few functions from a static MathUtils class, so on.
Maybe your ClassAUnitTest meta-file could include some of the ugliest methods from the class you're testing but not the entire class.
https://github.com/dotnet/runtime/blob/main/src/coreclr/gc/g...
https://github.com/python/cpython/blob/main/Python/ceval.c#L...
Then, reimplement as a simple state machine but this time, fill in all the transitions (event+state => new state + action)
One was an Infiniband code base from the vendor - a 'computer scientist' had written several layers to do what one or two could accomplish. Another, the Windows CE DHCP client (went from seconds to choose an address to milliseconds). Then there was an HDLC modem protocol - I got done, that was sped up a multiple and no longer crashed.
I can't understand them by just reading. I had to make a road-map of all the states, events, actions and interfaces. Design a new code. Then make sure every function of the 'old' code was represented in the new code - line by line. So nothing got dropped.
Satisfying. But more like turning the crank and making sausages than design or architecture.
If you divide the single 11k-line file into a thousand 11-line files, it may become objectively much harder to understand, but it'll also receive much less flak, guaranteed.
I suspect this is also why Architecture Astronaut-ery can be so successful within a company. If code is chock-full of superficial signs of order and craftsmanship, such as hierarchy, abstraction, and Design Patterns(TM), it takes a lot of mental effort to criticize it, and most people won't.
is it different from regular bikeshedding? or are you saying that the dark twin is the evolutionary process of eg. architecture gaining complexity until it becomes difficult to criticize..
I have a 4000-line script in a single file that has served me very well. It's perfectly organized and modular. I thought about breaking it into more files but it seemed pointless. It's very convenient for jumping through every mention of a variable, for example.
A thousand 11-line files? You definitely could not make that guarantee of the people I work with.
When I was finally able to retire the project several years later, I first replaced the home page with this picture: http://2.bp.blogspot.com/-6OWKZqvvPh8/UjBJ6xPxwjI/AAAAAAAAOv...
The first project I inherited was PHP app that used a custom UI framework created by an agency that didn't work with us anymore.
One file had 7000 LoC and it would generate hundreds LoC of with-sprinkled JS code and send it to the browser on every click.
Debugging that thing was a nightmare.
I'm still not a programmer.
EDIT: I'll go even further. Programmers who don't like long files are probably using the scrollbar to navigate around the file. Vim saves me from that bad habit.
$ wc -l perl/lib/Sys/Guestfs.xs
11930 perl/lib/Sys/Guestfs.xs
Worse still, this expands to C which can be large and takes a noticable time to compile: 30019 perl/lib/Sys/Guestfs.cAbility to search through your codebase by file name
Ability to hide irrelevant information and expose a higher level API through private functions
My mortgage is paid by a 50 kLoC C program with a single 11000-line function. I'm always blown away by how many so-called "code editors" can't give me a simple list of C functions in the file, the way BRIEF could in 1987.
Few things annoy me more than having to trudge through a codebase with hundreds of .c files, inevitably all with 8-character filenames. Any day when I have to break out Eclipse to navigate an unfamiliar project is an official Bad Day At Work.
If you split it, it's crucial that you're splitting the logic in the right way (if the modules are too small, they'll just waste your time) and that you're making sure references can be easily traced (eg. if you have modules with some DI system which prevents references from being recognised, as it happens frequently in certain node.js enterprise applications).
If it were a huge, single file, with very understandable modularity within that file, likely nobody would've bothered to write a blog post about it :-)
We have a few multi-kloc legacy monsters where I work and I quite often completely lose my place when working on them (and, by association, my train of thought), even though they’re actually structured somewhat reasonably.
Unfortunately, at one point I got so used to navigating with the outline that I ended up making a 1500 line function in C (I was an even worse C programmer then than I am now). Because of the outline, I could read and follow it easily, but anyone with a different editor was royally screwed :-(
If you're interested, the editor is LEO (http://leoeditor.com/) it's been mentioned on HN a few times
* Compiler support for function-level incremental may not be taken granted
* Editor shows a nice file tree (although you can do that with symbols too)
* Working with git is easier
* Reading code on site like GitHub is easier
I agree that obsessing over file length is it’s own kind of anti pattern. I have had colleagues who insist on putting every little thing in a different file and that is its own special kind of hell.
Similarly having the discipline to separate your programming logic into different files will force you to think about the architecture(or lack thereof) of your program. This is a good thing. The Java "one file for one class" model is overkill IMO, but it does force programmers to discipline themselves by thinking in terms of namespaces/classes when they write code, which for beginners at least is not a terrible thing.
Obviously version control is another reason. Hard to get work done if everyone is working on the same file.
It also will make it easier for someone else to grok your program. When I git clone someone else's large project, I start trying to understand the project by writing down what each folder and file is designed to do before I go any further. I suppose if everything was in one file I would just have to do the same for functions, but what if there were thousands of them? Imagine if a large program like WordPress, Doom, the Linux Kernel etc was a single file?
TLDR: For small one person projects, no big deal. Otherwise, it's just a bad/unscalable practice.
I just searched for the largest code files on my system and found a 100k file
I opened it on the online repo and gitlab did not want to show it at first. When I clicked on show anyways, Firefox broke trying to load the website and I had to restart it. (then I could not post on HN anymore due to the noprocrast setting there)
When I opened that file in an IDE, it was shown quickly without any issues. But there is a notable delay when typing, so 100k lines are too much.
Other IDEs might already fail with smaller files
Although when the IDE has a "search in the open file" and not a "search in all files of the project" feature, one file is much easier to use than multiple files
The outer loop was something like 4 kLOC, and consisted of blocks where there were first 20 lines of loadTable(filename) calls, then a call to calculateLosses( <all the tables just loaded> ) and then freeTable( <just loaded tables> ) calls. The inner loop was a little bit of setup and then a very long part where all those losses would be subtracted from the spectra.
The funny thing was, that once you got the structure, that code was actually not that bad. However, I told my boss several times that the second something comes along and doesn't exactly fit into that pattern the entire thing will blow up, and was always told that they maintain that code for 15 years and that didn't happen yet.
In a week, she came back and said "OK I've finished the prototype." I thought no way and I asked her to demo. Try X, try Y, try X + Y, etc. -- it all worked.
Then I looked at the code.
She had written the API handler as a one gigantic function, presumably because Twilio gives you a single API callback on an incoming call. It was a maze of nested if-statements going 10+ levels deep, subroutines relentlessly copy-pasted inline throughout the whole thing. Then she manually tested by dialing the phone 100s of times, putting in hacks throughout the if-tree.
Her prototype ended up being pretty easy to refactor and was ultimately the basis (at least logic-wise) of what we put into production.
its fine
you just binary search your way into it, put print "AAAA" in the middle, see if its printed, then put it in the half of the half and try agian.
emacs couldnt even find the bracket ending of the if condition (not the block, the condition..), have you ever seen if conditions(again, not the block) that spans your whole screen?
its not as bad as you think, it made me realize we take code very seriously, but its actually ok, 10k line file 100k line file, whatever.. its all the same
I prefer many short files and folders structured hierarchically and grouped semantically. I have no proof this is better so I would probably just leave it to a vote with the team.
In the end I think they is how a lot of this should be viewed until we get proper research. How do you WANT to code? TDD? No tests? One giant file? It should be a team and executive decision.
If you don't like the style on your team, and nobody wants to change it, move on or adapt.
Technical debt is like a superfund site. It renders the real estate worthless and poisons the rest of the company.
It does matter. My current gig is hemorrhaging money because we can't keep devs even though the pay and benefits are great. We cannot execute on mission critical initiatives.
We cannot adapt our product to meet the needs of the market in an agile way.
This is due to people saying "a working product is more important blah blah.." for years. I would argue there is a balance to strike and you can do both with a good team and realistic planning. But there is always the nay-sayer who is willing to step in and say whatever product wants to hear.
It is so bad we cannot train people to use the software anymore. It is too poor quality and were can't on-board them before they decide to go elsewhere.
Everyone who knew anything has left and there is too much of it. So the remaining devs get overwhelmed, they leave... It is a vicious cycle.
The funny thing is the money machine works, but it is so frustrating to see all of the extra money we could be making and having to leave it on the table.
One of my first jobs was as a maintenance programmer for a 100KLoC (or more) single-file FORTRAN IV (1970s vintage) application (a proto-email server).
Three-character variable names, no documentation, and having been stepped on by every junior programmer that went before me.
My best debugger was a Ouija board.
The original author was a ringer for Donald "Duck" Dunn. Interesting chap.
It taught me the value of writing good code.
It sucked.
It was great (because of the lesson learned).
Thing is, there's a great knowledge of the Android OS in there and the app works great when I use it. I think the OP blog post is correct 'Users don't care about the technologies or code.'
All of which is to say, by all means argue about whether colossal files are acceptable software engineering, sometimes that fight takes a back seat to "a double-digit percentage of the company's CPU and memory are wasted on parsing and loading this file in literally every new process".
https://github.com/microsoft/TypeScript/blob/main/src/compil...
Does anyone know why and how they maintain it?
Some documentation from Orta Therox on the checker:
https://github.com/microsoft/TypeScript-Compiler-Notes/blob/...
What does that even mean? It seems that typescript uses an alias for string as __String in the source but then a bitwise operator with string?
Fix your include/import systems before preaching for modularity.
After the better part of a week becoming acquainted with the code, I found a suitable integration point. Luckily for me, the new feature being requested didn't depend too much on the existing code so I didn't have to make too many modifications to the existing code. I added the entry point to the new section along with some comments describing how things worked and some ascii art of a dragon. In the end, the new feature worked great and the customer was very happy with the results.
Some years later, I was working for a different consulting firm and that project surfaced again. This time it was being re-written in ASP .NET after being passed around to a couple of different off-shore development teams. My coworker was working on it and asked me if I had written a specific piece of code in index.asp. I took a look and we both had a laugh, because my ascii survived after all those years!
Is that the same for all languages?
Is it just me that get piqued by the sound of it? I often spend my night figuring something intriguing and chasing that aha moment. I also like play puzzle games, and to me, its one way or another to spend the time. Then if I get to clean the mess up, usually that's another rewarding effort. However, I definitely understand the frustrations if one's hands are tied or there is a deadline that you just want it disappear.
I love to clean up my own mess as well. We all make messes, at different levels and at different perspectives. Just like playing a game, it is just boss fight at different level. Novices make mess, veterans make mess as well. Usually novice's mess are easy mess.
(The conditional in that if {} always evaluated to true).
var a = true;
if(a == true) then return true;
else if(a == false) then return false;
Or something like that.
This is why I’m always loathe to criticise stuff I see on WTF.
public static bool ConvertToBoolean(this int number)
{
var TrueOrFalse = false;
if(number != 0)
TrueOrFalse = true;
return TrueOrFalse;
}But the thing as a whole didn’t follow any best practices I’ve ever heard of: the project also had what looked like a bizarre attempt to reimplement the concept of properties(!), in that the UI classes had fixed length arrays of all their subviews, which were accessed by constants. Each section of the UI had its own god class, where different views in that section were all the same class, called with a constructor that determined if a view ought to be created for any given constant.
There were also something like 20,000 blank comments. No idea why, the guy who added them didn’t even understand my question:
int something = foo();
//
double baz = bar(something, 5);
That kind of thing.(The project is no longer available and the business who made it has since closed, before anyone asks).
[1] Merriam Webster definition of edifice: a large abstract structure
[2] Source: Me
* Or inherit and maintain
It's main file is 88kloc Fortran 77, started in the eighties, still actively developed.
https://www.iap.kit.edu/corsika/index.php
Currently a rewrite in c++ is underway.
My favorite quote has got to be: "Unit tests aren't meant for you now, it's insurance against a future developer".
If this app was factored into 300 different files, it would still be an impossible mess. The redundant and buggy logic would just be in different files.
Someone had a good idea to make a header.jsp template for common header stuff.
But it was hilarious. The file essentially became a giant if-else condition with a few hundred conditions like "if path == 'some_page'" followed by CSS and sometimes JavaScript for that page.
Absolutely horrendous.
The existence of large files is mostly just a style issue.
Text editors with different or lets say more semantic interfaces to the code would not care about file size.
You could then happily have 1M loc in a single file.
You would care about it as much as you care about how the code was laid out on sectors and pages on your HDD.
A 10k single line program can be easier to understand and better organized than an overly abstracted mess strewn across multiple files, but which checks all the "best practice" checkboxes.
As a first-order guess, I would guess that no code has its cyclomatic complexity measured. As a fraction, I'm probably not too far off.
I suspect there are a lot of younger programmers with these views, who don't know elsewise.
Testing suites, having dedicated test/dev environments, all of these things are relatively new across most programming on the web. "Best practice" has changes more times than I can count in the last 20 years, and we've gone full circle from "monoliths are bad, split everything into microservices" to "maybe try combining them to reduce complexity".
I'm not saying this style of coding is the best, but automatically assuming the current practices are the definitive best ones in all cases is silly, and the idea you have to refactor because it doesn't meet those practices is - in my view - insane and wasteful.
One factor in favor of mono-file is its often easier to navigate around fast and do search/replace in a single file then across many. Believe me, I understand all the benefits of mutiple files too -- not my first rodeo -- but under right conditions not always a sin to have a "fat" file.
generally after a project has left its prototype stage and become more of a stable thing, primarily under maintenance, it should certainly be split out into smaller files with sensible module boundaries. Best in long term, and plays with version control better, and large teams doing concurrent mutations on the source tree.
https://www.felienne.com/archives/2974
https://www.microsoft.com/en-us/research/blog/lambda-the-ult...
No metion of a cell-u-lite version.
I ran wc on FreePascal to search for some. There are a few, but not as many as I expected.
9k file, data structures for the compiler itself: https://gitlab.com/freepascal.org/fpc/source/-/blob/main/./c...
30k file: Pascal parser/scope resolver: https://gitlab.com/freepascal.org/fpc/source/-/blob/main/pac...
And the record:
119k file, Sharepoint API (but it seems to be autogenerated): https://gitlab.com/freepascal.org/fpc/source/-/blob/main/pac...
As far as libraries go, this is one of my favorites:
23k file, regular expression library: https://github.com/BeRo1985/flre/blob/master/src/FLRE.pas
I searched my own files and found a 197k file to parse HTML entities. But that was an autogenerated trie (one switch/case for each letter)
Yet! Those non-programming people somehow managed to add their little requirements over the years, without breaking the other forms?
There is probably more to the story. Like that, I suspect, users who needed to produce a certain form probably had private, years-old copies of the program that they used, impervious to subsequent changes.
It'd be pretty unmanageable without Visual Assist (a plugin for Visual Studio that does fairly fast searching for symbols and files).
My horror story was being asked to do maintenance on PHP sites written in the early days of PHP. Hundreds - maybe thousands - of PHP files with copy-pasted HTML and intermingled logic. As far as I can tell, the idea of instantiating a whole MVC framework from a single entrypoint file came after that particular site was created, so every possible page was its own entrypoint with its own boilerplate. Source control also seemed to predate this project, so you had plenty of .old.php and .old.v2.php files.
Programming at webdev agencies is a challenging experience.
Oof. That tbh sounds worse
There was a file in that source that had over 900 cyclomatic complexity.
The reason was that at that point Android had a limit on their method count. Called the dex limit. You could only have 64k methods, before multidex was introduced.
The engineers couldn't add more methods, so they had to jam code into existing methods. No DRY. No refactoring.
It was fucking impossible to understand at the code level. You had to just use the app for years to understand it.
That analysis got a bunch of us targeted by the senior developers because it made them look bad. They fired multiple people, including me, who pointed this out. We said it made it impossible for new engineers on the project to succeed, and the incumbents didn't want that to be known, less they lose their bonuses.
This shell script was 9,000 lines long. Each new "menu" after entering a number corresponding to your selection from a previous menu took 60-120 seconds to appear. Needless to say, configuring systems was a painfully slow process.
I quickly found it practically impossible to extend this tool to meet the new requirements, so I quietly rewrote it (also still in shell script, because of host limitations). The new version was 1,000 lines, and menu changes took < 2 seconds in most cases (thanks to appropriately placed caching).
Management was not happy that I decided to rewrite (since the manager had written it himself originally), but the users were ecstatic about the performance increase. I would not be surprised if that tool was still in use today.
What year is this? 70's and mainframes? Because no way in hell you cannot, on large organization that had "Jeff in marketing", since 80's and PC's, duplicate the environment. Especially given this was used by almost everybody in the organization.
And once you're done duplicating the production, and create proper test environment, you can start actually refactoring and creating a beautiful app out of it instead of just "To this day I sometimes lie in bed wondering what could have caused this".
Conclusion - article's author is a whiner instead of a solver, no better than the ones before him that "copy and pasted at some point then later diverged"
But one antipattern might be that the file is being modified by other people, even if on paper he is the only owner. So while you create a test, someone you don't even know exists modifies prod right under you.
Another antipattern is forever-changing ownership. You own it for 2 weeks. Then someone anonymous in some layer above you decides something else, and of it goes. Maybe they did bother to tell you you're not the owner anymore, maybe they only told the new owner. In these 2 weeks, you've got ownership of 3 new programs and lost it for 2 others.
I've seen it happen. There is no way to build something stable in these circumstances. Management will need to provide some stability before the underlings can do any work. If you're living there, run away, you can't save them.
Yes, there is. Since he clearly stated that started as a job assigned to his team, he could've version it right there where it was officially staying. Then it would've be very easy to see changes done to it by somebody else, even if who was that somebody else was unknown to him.
"Then someone anonymous in some layer above you decides something else"
He also stated in the article that he had a lot of time in his hands and tried a refactoring. You don't have time to do a refactoring of a file with 11k lines of code in only 2 weeks, so he clearly had at least several months.
My conclusion still stands.
But the moral of the story in the article is that unfortunately some things break much less easily than we'd like.
It was possible to use include files to split up a single unit, but that wasn't used much either because lack of cross-file navigation features in IDEs of the era made it really difficult to manage many files.
https://gitlab.iap.kit.edu/AirShowerPhysics/crmc/-/blob/mast...
I see lots of people getting confused these days about modularity and people missing the point altogether. What is it you're after? What's most important to you? That's where you need to focus your modularity efforts.
About 10 years ago I was shorty working on an app for bank tellers.
The app was in production for 2-3 years at that point, 4 people where implementing new features.
The whole app was 4 Java files/classes.
2 files had about 1.000LOC, 1 had about 15.000LOC, and last one had 80.000+ LOC.
The big one did it all, UI, db calls, forms for printing, network calls.
Most of the variables where named something like “c12”, “bkf22” etc.
Either figure out how to improve the situation, or leave.
My, how our knowledge of how to do things has changed!
Countless non-software developers, ranging from IT support to business analyst able to implement business requirements. Good for business if you ask me.
Your boss will still expect you to be able to fix bugs in an hour or two without introducing any new issues.
It was an interesting exercise but I’d never do it again.
1. Who blog about how bad a 11,000 line project is
2. Who shipped the 11,000 line project
> I have no idea.
Sometimes successful software just grows.
Perfect code you never shipped.
Single 11,000 file code you shipped.
A seasoned programmer should have no problem navigating a multi-million-line codebase, that's just routine.
There isn't anything that special about a 10k file.
There is no need — at all — to be evangelical about it one way or another.