Cold Showers
github.com
github.com
LOL, this hits close to home. My company had a modeling specific VM set up to run our predictive modeling pipelines. Typical pipeline is about 50,000 to 5 million rows of training data. At best, using an expensive VM, we managed to get 2x training speed from lightgbm on the VM vs my personal work laptop. We tried GPU boxes, hyper threaded machines, you name it. At the end of the day, we decided to let our data scientists just run models locally.
AWS EC2 has a VT1 instance family that enables high-speed A/V encoding via a Xilinx media accelerator card.
8 channels of DDR4-3200 only provide 200GB/s bandwidth. RTX 3090 has 936GB/s. So even 4 socket Xeon won't catch up.
Many if not most other algorithms are iterative. Hell, even sorting is.
I guess it's time to invent a blockchain that trains ML models as PoW :)
As a sidenote, I have rented servers with GeForce cards from multiple providers in multiple countries, so this rule doesn't seem to be respected very much. And since it's part of the driver EULA, nvidia can't legally go after server providers, since they don't install any drivers, just build and rent out the hardware. For all they know, all their customers are running noveau.
[0] https://www.datacenterdynamics.com/en/news/nvidia-updates-ge...
GPUs are like heavy flywheels. Getting them up to speed takes some time (copy data, compile and copy the kernels, kickstart everything, etc.), so you need to start them once to get the performance benefits. Otherwise CPU is much more nimble since they're closer to RAM and made to juggle things around.
After that we colo'd 4 big Xeon servers for about $1,600/month total. Looking back on it now it's just so insane...$30,000? no way.
There's just so many better things for your company to be spending the money on.
We don't let pencil pushers with MBAs anywhere near what we're doing, and it's going great.
I know this isn't the most usual configuration but if undervaluing my skills and trying to bottom dollar on them is going to be their rules, then I'm just going to do my own thing, and they're just going to have to scrape the bottom of the barrel for talent.
I hope the zeitgeist changes any time soon. God knows how many unicorns have been sacrificed with that kind of paradigm which could be successful companies by now.
I don't know how that changes the equation. No one is undervaluing your skills; it's would you rather spend your time driving to a colo center to replace a RAID array or working on $product.
With AWS you are outsourcing an IT team, not just processors and how you approach pricing should reflect that.
If you’re doing colocation to save money, you’ve also figured out that going to the datacenter sucks and it’s a terrible place to do work.
You’re not building your own servers from scratch, you’re generally purchasing them from a vendor who offers a warranty and optional on-site service.
Or you’re leasing them from a hosting company who will take care of those pesky RAID alarms for you.
You (or your hosting providers) have likely outfitted your server with remote out-of-band access to allow you to get into BIOS or the RAID controller without physically being in front of the server.
And finally, you have remote access to power cycle the server (or a batphone at your hosting provider to do it on your behalf).
I want to say that these datacenter-visit-prevention techniques have been near standard practice for a decade-and-a-half.
Or is it just me and my circle that do this?
Nope, this seems to be the norm. I've worked on a couple colo servers that nobody at the company had ever actually seen in person. They figured out colo in Germany was the best deal, so they had some servers delivered straight to the DC and the staff there installed them and plugged them into an IP KVM. Not sure if this is a standard service most providers offer, but I'm sure a big enough cheque would convince most - and considering the cost of transporting both the hardware and engineer to install it, that cheque can be quite large.
Our in house data centres we visit more often, usually to add new equipment, doesn’t take long to walk down stairs.
I can’t think of the last time a hard drive failed
Take all those things you just talked about, and expand them horizontally and vertically up the stack, and you have 'AWS'.
So not just 'a guy to replace the hardware' - but now it's software configurable, has all sorts of other, fancy things.
Time is money, and it's expensive to pay people to mess with things if they don't have to.
It's like this:
If your company needs 3 cars, you rent/lease them. You do not hire your own mechanics, even if technically speaking "we could change the oil for so much less!"
If your company is in the business of transportation, and you have thousands of trucks, you may want your own repair/maintenance team etc. instead of paying some service company a fat margin to change the oil.
For most things 'local prem' is an optimization that usually needs on some degree of scale to justify, or, you have a peculiar setup i.e. a couple of well versed hardware and networking guys who have no problem with a bit of a physical setup. Which can be a bonus.
"I hope the zeitgeist changes any time soon."
No, it won't, it's going in the 'other direction' forever, because the 'economies of scale' at Amazon, it's incredibly difficult for individual engineers to compete with those efficiencies.
Just the opposite of 'being a problem for startups' , the 'cloud' has basically made entire swaths of types of startups possible where they wold not otherwise.
Like everything, you have to use think about it a bit but their costs are really, really transparent (imagine Oracle trying to do it ...).
What a glib and senseless follow up.
Also, it's not peanuts. How many extra developers are $28400 per month?
This then means that while op wouldn't immediately add 1M ARR, the original dev and the new devs could soon add 5M ARR (stupid extrapolation, it won't be as much in practice, at least normally)
According to my napkin calculus, you could get about 4.5-9 developers (for 1-2x minimum wage) onboard here in France for 28k€/month. I'm betting in other countries with low salaries and a vast talent pool, 28k€/month would get you even more employees.
In Austria, the "IT Kollektivvertrag" [0][1] (think contract for the whole IT collective/unions) demands a minimum of €2503 brutto (before taxes) per month for developers (those are normally falling under ST1 category), and that 14 times a year, and that's for entry level (as in, not first job but starting at a company). Note also the 14 times a year, where the last two extra salaries (Christmas pay for December and vacation pay for June) is taxed much less (note, we can get bonuses on top of that too).
[0]: https://www.wko.at/service/kollektivvertrag/kv-abschluss-inf...
Note for above PDF (only the short money table the full one can be easily found via searching "Austria IT Kollektivvertrag 2022"):
- The ST1 is for devs, and LT1 for leadership roles.
- "Einstiegsstufe" is Junior, "Regelstufe" is "normal" and "Erfahrungsstufe" is Senior
So if you get hired as senior in a leadership role you'd be entitled to €5521 Brutto salary, 14 times a year, or more depending on your experience/knowledge and your negotiation skills.
In France a young developer easily costs €100k/year to employ, even if he earns a third of that before taxes.
No, actually they don't. In the time I've been there they've 10x the size of the company. Maintained majority control through multiple rounds of funding. Significantly increased salaries. Provided a great work life balance. Etc.
Why are developers obsessed with how many additional developers they could hire with hypothetical savings?
It will be an interesting next few years.
ARR could be whatever. They are not interchangeable or mutually exclusive.
There are some data sets where you really do need big-data tools, but it's for when you have petabyte-scale data, not megabyte/gigabyte-scale data.
However the complexity of the algorithm many times scales with the size of the dataset, either the full corpus or the size of individual examples.
I think data scientists just aren't really hugely concerned with programming optimizations or bottlnecks or whatever. Most of them are just intermediate-level python programmers, and that's completely fine until they think they need a hadoop cluster for whatever they're doing and the costs start piling up.
It get's you ~10x speedups for batch predictions, more if your model is big. It's not complicated, it ended up being <1K lines of Python code. I heard a couple of stories like yours, where people had multi-node spark clusters running LightGBM, and it always amused me because by if you compiled the trees instead you could get rid of the whole cluster.
At least to me, the big advantage of static typing is not that it (allegedly) reduces bugs, but that it aids my understanding and helps in navigating the program. It's a tool for thinking and communicating.
That’s a dynamic typed language with comments
I remember working 2012 on a SaaS app, and I wasn't the only guy anymore doing frontend stuff with JS. I knew my objects, but my colleagues did not. How to you document object APIs? TypeScript really shines in large projects with lots of devs.
All the other claims from readability to understandability to refactoring to less bugs, all come with an “it depends” caveat. Sometimes the claims are true, sometimes they’re not. It’s also not possible to say “but in most cases claim X holds”.
The thing I’ve never understood yet in this debate is in my experience, the people who have argued about correctness have universally been below par at getting to the bottom of requirements. Which leads to “great, you correctly built the wrong thing. And you took forever to do it.” Which isn’t doing our profession any good in the eyes of other professions who depend on us.
Where this falls apart, the more verbose writing style hasn't been proven to convey more information or in a better way. That's an assumption still tossed around.
And typed or untyped, you’re only ever reasoning about the types in the context you’re working in, not the entire program.
A Python function can be called from within or from outside of your codebase with different callsites passing different types.
This can make it much harder to actually reason about the code, while making it seem easier to reason about. Most people would agree w/ your reasoning on a short piece of logic, which then at runtime spectacularly fails because the inputs don't adhere to the types you expected. In a statically typed language you would not even have gotten it to compile and while it might not feel like a bug is being prevented and actually feel tedious, every time your IDE (or compiler) tells you that the type on something is wrong, you've prevented a potential bug.
Let's say we compare Javascript and Typescript (as they're so close but one has static typing.
const myFunc = (param) => {
doSomethingWith(param?.property);
}
Easy, right? Well, does param actually have `property`? No idea. What type is `property`? Does the function `doSomethingWith` take that kind of input? No idea. Now I have to check that function, which might be coming from I don't know where, I might not even have an IDE that can reliably determine where `doSomethingWith` is coming from exactly. Even if I can navigate there now I have to check that piece of code and any other code it calls with `property`. Maybe `property` itself is an object and `doSomethingWith` assumes it has yet another property. This can easily go quite deep and I will not be able to easily reason about this at all. You can't tell me that someone can have all possible runtime combinations of this in his head for any reasonably sized program.Now let's take something that is almost equal but slightly longer to read and write, same thing in Typescript. I've had to define the types of these things somewhere once. Big deal.
const myFunc = (param: SomeType) => {
doSomethingWith(param.property);
}
Notice how this is really not much of a difference. Just a type declaration and it gives me a lot of safety. Let's assume SomeType defined `property` as non-null, so no `?` needed, I know my inputs have already been checked. `doSomethingWith` also defines its parameter type correctly and we know what `property` is or isn't. No need to know anything from the top of my head or spend time digging through code myself. The compiler knows that I am passing the correct type of object along and I won't get a runtime error (well, OK, it's Typescript, so let's also assume I'm not in a mixed TS/JS code base where I might very easily get `any` kind of object.Now syntax will be a little bit different, but I would argue the exact same thing in say Java or Kotlin is equivalently short and readable (yes even in Java!) while benefiting from even more type safety:
public myFunc(SomeType param) {
doSomethingWith(param.getProperty());
}
Didn't really hurt much, did it?But these are super simple example. It get can arbitrarily complex.
But Java, Kotlin, and Typescript types are very weak sauce. When types can do more, we can do more with them.
I associate verbosity with object-oriented programming, whether statically typed or not.
Here's something I can express with static typing that I can't express with dynamic typing: "this function returns a function which returns an integer for every input". There's no test you could write to verify this property. So I'm inclined to say that static typing is more expressive, since it gives me a way to express and verify properties like this.
Without compile-time types, you are not equipped to express serious compile-time work.
Removing traffic lights and stop signs actually reduces accidents because drivers are more careful when driving through intersections which reduces speeds and drivers become more alert.
Developers will adapt to their toolset. If you have a statically typed language, you trust it will deal with type related issues and you become more lax with testing things related to types. When you develop in non-typed languages like Ruby, you tend to write more tests and not trust your compiler (because you don't have one). This is why you will find most Ruby developers are really good at writing tests and embracing TDD.
A type system can keep you from having to write those tests.
Because with a proper static lang (hint: not Java, not C#), nil doesn't exist? Right.
The claim is true: a type system _can_ prevent null-related issues and eliminate the need to account for them in tests. That's not the same as saying every type system does.
Static typing gives you assurances and tools with which to test your assumptions in the code, for those times when reading the whole stack is cumbersome, and you need to defend against less careful developers. It also transfers a bit of knowledge between developers in a trivial way that would otherwise be a pain to communicate.
In my experience similar arguments hold for software developers. Especially caring can be a big factor; i.e. the "move fast, break things" mentality.
I've been back and forth between typed and untyped languages (somewhere in the range of haskell and tcl) and personally prefer less typing when hacking things together and more typing for high quality software. I'm currently working an infra job where we use both ansible and terraform. They're not direct competitors, but I tend to prefer terraform over ansible when possible, as terraform gives me more "static" guarantees, which translates to more confidence when we apply our code.
However, refactoring code in C# is much easier than refactoring ruby because you can lean on the type system there. However writing new code in C# is often much harder to do in C# because of the constraints of the type system. So really, it ends up being a wash for me.
See, When I'm throwing together apps to clean up configurations, I am Pythonifying XML often. And when handling different return values, reshaping it into the useful components I need and trying to analyze data (and dealing with different return formats depending on number of results, aka a dict if there is one value, or a list(dict) if there are more) I have to constantly remember if I am going to be getting a list(dict(dict(dict(str)))) or just a dict(dict(string)), and so on. But that's me cobbling together scripts and not understanding the API by heart well enough.
That's the point though. With dynamic typing you would only (hopefully) catch this with manually written tests. With static typing you get that feedback for free at build time.
How would that pass any code review, regardless of static or dynamic typing?
What kind of clown show of a programming org are you working at?
This is morally equivalent to "There's no point to having a safety on a gun, because the safety won't stop you from bashing someone in the face with the gun." If you really want to, you can throw exceptions or crash the process or call exit() or call system("shutdown -h now") anywhere in your codebase. That has nothing to do with a type system.
If you have not seen them, the reason is probably that the code was tested well enough before you looked for the bugs.
You clearly have little experience with such languages, then.
It makes you actually have to consider scenarios where a variable can be none or not and try to push the validation up closer to where it entered the system.
This is false.
TypeScript, Swift, and Rust are commonly used and support non-nullable references.
> those languages, don't help you because you are dealing with real world data where inputs to your system can be null or not so you end up using some type system escape hatch anyway.
You don't need escape hatches to deal with "real world data" that may be missing some values. This blog post is my favorite detailed comparison of handling "real world data" in static vs dynamic languages: https://lexi-lambda.github.io/blog/2020/01/19/no-dynamic-typ...
Rust's serde_json docs on "Operating on untyped JSON values" are also a pretty good description of working with "real world data" where you want to examine an arbitrary document: https://docs.serde.rs/serde_json/#operating-on-untyped-json-...
The only requirement for working with "real world data" in a static non-nullable language is to choose whatever kind of behaviour you want when working with the data. Everything you can do with null references, you can do better with option types; there is nothing that null references uniquely permit.
1. Python with mypy has `strict_optional`. On by default.
2. C, being “portable assembler” is not really statically typed.
3. Java has had Optional for years, although it’s not the most pleasant to work with it does exist. And JVM languages like Kotlin go well beyond this.
4. C++ has `not_null`.
5. C# supports type-system enforced non-nullable types since 8.0.
You said:
> languages that disallow nulls, if you are one of the 10 programmers on earth working in one of those languages
I hope it’s clear that you are simply incorrect. There are plenty of tools to eliminate nullable references in modern mainstream languages.
Most importantly in C++ only pointers can be null. You can return value types that are not nullable.
There have been people writing at least two of those languages everywhere I've worked for a while. Most of my professional colleagues can write at least one of these comfortably. I'm extremely confident in being able to hire programmers for all of these. They're all in use at every major tech company.
If you really want to stick your head in the sand and cry about how nothing can be better until they're literally top of the charts, I can't stop you, but they're certainly not rare. There's good stuff out there. Lots of people are using it. You can too.
If you'd rather trade links to charts, I trust Stack Overflow's developer survey's methodology a lot more than TIOBE's. 30% of respondents said they've worked with TypeScript, and that jumps to 36% in the professional developer subset. Rust is 7%/6%. That's a hell of a lot more than 10 developers.
https://insights.stackoverflow.com/survey/2021#technology-mo...
They also got 15% of developers who aren't using TypeScript want to use it, and 14% for Rust:
https://insights.stackoverflow.com/survey/2021#most-loved-dr...
My country has about 15% black people about about 7% asian people. My country has about 4% LGBT people, and my city has about 15% LGBT. It would be really weird to hear someone say that black, asian, and LGBT people are not common, especially after knowing and working with plenty of them.
Also in the opposite direction, many dynamically typed language allows specifying types if you want to including python.
x still has static type, the compiler just infers it based on the assignment, the type information is still there. Agree that implicit/unsafe casting is still and issue in some languages though.
And yet, per TFA, it’s not; at a minimum it’s clearly not “self evident”.
Why do we developers value our personal experience above studies, while dunking on average citizens for doing the same?
Guess we’re just as human as the rest of humanity; subject to the same urge to trust our own beliefs over contrary evidence.
Cold shower attempted, but the plumbing was busted?
Such studies invariably wholly miss the point: when you have a language with powerful type support, error checking is the least valuable work you get out of them. Types do serious heavy lifting expressing semantics.
And at the end of the day you have to contend with being in a work environment where politics and personalities rule, not science (or engineering).
That said I do wish more devs would take an interest in the available quality literature. Unfortunately I'm far more likely at work to run into an Uncle Bob recommendation at work, than a recommendation of ACM's Digital Library.
In a sense "type coverage" is analogous to test coverage.
Static typing reduces bugs because it aids your understanding.
Not sure though.
I suppose it's not even necessary to argue about experience fixing them or not, just the fact that those are runtime errors rather than compile-time (and so we presume not shipped) shows it reduces bugs doesn't it?
These are two very different outcomes.
If the remaining bugs are unrelated to the class of bugs that were eliminated entirely, then the difficulty in finding them has little bearing on the outcome, since we’re now talking about an entirely different class of bugs.
- in JS, your code will run with the bug then do something catastrophic during runtime that you can then notice and trace to the core issue
- in Java it won't compile, so you fix it so it compiles and runs, then it'll hit you in like 2 hours of runtime with a NPE or something and you'll have no idea what caused it
Maybe Kotlin, Rust, and the like solve that sort of thing better but I've yet to be convinced.
It seems you're describing an orthogonal issue, and it's unclear why type checking is a Bad Thing or even related to the NPE at all.
Let's say I work on an assembly line, and must place physical parts into a machine that assembles a larger part. There are many ways this machine can break down - I could put the wrong parts in, leading to a complete failure, or some part of the machine could malfunction independently.
- We could implement part validation on the assembly machine to make sure it's impossible to insert the wrong parts. This eliminates failures related to incorrect part insertion.
- Unrelated to this, a drive belt starts to wear out and slips every so often, leading to a slight slowdown in a conveyor belt, which ultimately leads to a botched item.
The way I read your argument, you would say that part validation is bad, because it's easier to diagnose a meltdown when incorrect parts are inserted by the operator than it is to determine that the drive belt is slipping.
Except the drive belt slipping is not related to operator error, and would have happened whether part validation was happening or not.
This is hopefully obviously nonsensical - better to reduce the overall error rate by implementing part validation than to leave two avenues for error. Before part validation, the machine could fail because of operator error (common) or drive belt failure (uncommon). After part validation, only the uncommon error occurs.
This is better than no validation at all, even if drive belt failure is harder to identify than the machine screeching to a halt when the wrong parts are inserted.
What am I missing?
And Unlike typescript my code doesn't need to be transpiled at all since it is already vanilla JS.
I want a shirt with this on it.
It may not be very comprehensive static typing but it is static typing none-the-less.
def add_item_to_cart(item)
vs void add_item_to_cart(IItem item)
They are equally easy to understand. The first is easier to read.I guess you now need to read through the implementation or docs. The first is much easier to read incorrectly.
The only citation I have is the tenuous grip I have on my own sanity - I could have more correctly talked about the incredible amount of mental overhead this has _for me_, but read the rest of the thread and you’ll see that this isn’t an uncommon experience. As I said, if you can work around this then you have my respect.
What you see in a statically typed language: IBlaha blah. What could blah be? An IBlaha. What could IBlaha be? Anything! The type has not gained you anything.
> The only citation I have is the tenuous grip I have on my own sanity
That's an argument from authority where you are the authority. It doesn't work on HN since we are all skilled developers. I've also been a software developer for decades and I can count on one hand the times static types has provided a tangible benefits.
Of course you could say the same thing about sane naming dynamic naming conventions for your declarations in a dynamically typed language - and you wouldn’t be wrong, but a compiler won’t help you in the case of human error. All I’m interested in is offloading as much complexity onto the tools at my disposal, so I can focus on what’s important.
On my citation … that was tongue in cheek and I thought it was obvious. I don’t have a citation, this is all my own experience. For the third time, if you can work your way around this you have my respect.
Dynamically typed languages are very popular so it seems that many developers can work their way around dynamic typing.
That definition is rarely as accessible as an explicit type though. For example take an API response or any third party library. Determining the data type isn't as quick as simply scanning a function for the object definition.
- Dynamically typed languages are very popular so it seems that many developers can work their way around dynamic typing.
As someone who has spent a fairly even mix of their career using typed/untyped languages, I think this is due to a few reasons:
- Lower initial learning curve.
- Lower barrier to entry.
Those are real benefits, but I would argue most projects quickly hit a point where they benefit from static analysis.
Having worked with 100s of devs at this point, I'm yet to meet one that after learning a typed language and using it for a sufficient period of time (more than a few months) wants to use an untyped language for anything outside of small scripts.
You're confusing Java-type extreme (and also mostly strawmanned) application of OOP with static typing.
Not every type in your program has to be AbstractFactoryProxyBeanInterface, and if you don't write code like that it's either obvious or some kind of extension interface for non-core code.
You've never worked with vaguely named variables? What you are suggesting is guessing the data type based off the name.
- What type of object? The type that can be added to a cart.
Okay sure, but what precisely is that?
- Nobody just throws random objects at a function.
I couldn't agree more - so the follow up question is what is the fastest way to get familiar with what type of input or output this function returns?
- They are familiar with the code in general and they know what to do.
For very small projects with very small teams after some onboarding time perhaps, but outside of this I would disagree.
Code changes over time, parts that you use to know intimately get changed subtley and erode knowledge away. Having types in place highlights these changes if your assumptions are incorrect.
It doesn't matter "what precisely" is the the thing that you are adding to the chart, and a static type system won't tell you that either. There could be be any of 1000 things that implement IItem. And probably half those things just throw exceptions for methods they aren't actually able to implement.
You can do that with Python (sometimes) because many libraries have type hints today, so even if you don't use types yourself, the type checker can infer them in your code and help you out.
The same way Rust checks for object lifespans with the borrow checker, which is distinct from the compiler and type system.
The same way valgrind for C can check for use after free.
The same way errorprone can look for null checks in Java.
This is a well tested and proven technique. Static code analysis is a staple of the industry, when it comes to automated code analysis.
All of the examples you gave are from static languages, where the information is known at compile time (except for valgrind, which requires a runtime). The parent to my original post was claiming that you can have the same tooling for Ruby.
Also, you're wrong about Rust. Lifetimes are part of the type system.
You need the same tests from a typed system in a non-typed. You _don't_ need all the tests from a non-typed system in a typed system.
Writing tests to enforce types just hand-rolls a type system, in my experience.
You do though, because invariably people violate the LSP and just "throw Unimplemented" in the methods required by the interface they can't figure out how to implement. In other words all system are duck typed in reality.
Not sure what typing system you're referring to, but it sounds very half-baked at best. I'm using Rust fwiw.
I cannot access any field or method that does not exist. Even dynamic traits are compile time enforced, but i think we can largely have this discussion around static dispatch.
addItemsToCart(items)
Vs Type ItemCode: string;
Type ItemDetails = {...};
addItemsToCart(items:ItemCode[])
or, for a slightly different implementation: addItemsToCart(items:Record<ItemCode, ItemDetails>)
If you only use trivial examples, types seem silly. But in real examples they become more useful. In this case looking at the function signature give you immediate information about the implementation that is missing from the untyped version.EDIT: please excuse formatting, I'm on mobile and cannot get it to add spaces before the last code block
void add_item_to_cart(auto item)
will still statically verify that item has the required syntax; C++ is very poor on this aspect on only doing the verification at instantiation time, languages with more sophisticate typing systems can infer the correct type from tome add_time_to_cart definition alone.My working day jobs have been mostly C++, and these days C#. Periodically, I will temporarily inherit some of my younger colleagues' projects, if they move on to greener pastures in different companies, with the charter of "can you do something about the long-running issues this software has been having?" My go-to solution is to go through their typescript and add return types to their functions, and replace their anys with interfaces. After having done that, I fix the bugs that revealed, and then I'm usually done. Recently when I did that, I came across a central class/data structure, which turned out to exist in no less than 5 slightly different variants. i.e. different parts of their code adhered to 5 different assumptions about what fields would exist and be populated (but all expressed on the blank canvas of 'any').
The problem is, other people with just as many credentials as you have the opposite experience. From an outsider's perspective, two people with equal authority say opposite things, what can they possibly do except an independent study?
Also, note that there's a reason anecdotal evidence is not always reliable. E.g. the famous story about fighter pilots and the "regression to the mean" hypothesis.
In this scenario, I honestly don't think it matters whose objectively right. Software is not a clean, normalized and organized set of use cases after all, maybe static typing works for person X and doesn't for person Y because of their background, or preferences, or codebase requirements, and so on.
Maybe one day we can conclusively prove that on aggregate static-typing/{insertThingHere} is overall less buggy, but even if we did, it'll still change depending on circumstances.
On the other hand, though: have you worked with large, thoroughly tested projects in a dynamic language? Personally, I find that good tests catch 99% of the bugs that static types do, plus quite a lot of other bugs as well. Arguably, you ought to write tests anyway to find those other bugs. Since they also find your typos etc, you get to enjoy the ergonomics boost of dynamic typing almost for free.
That's my (also anecdotal) argument for doubting static types.
function myFunction(user, security) { }
good luck finding out what user and security actually is. In static typing, it's all there.
And when we understand this, we can weigh it up with alternative tools for thinking and communicating!
Would this 10 line shell script be better in a statically typed language? Well maybe not because I can hold all of 10 lines in my head, there's nothing else to communicate.
Would this CRUD app using Django/Rails be better with static types? Well the framework has defined a structure that communicates properties of the code to me, I don't need types written down because I already know them.
Would this complex parsing process of untrusted data into a trusted and verified format benefit from static types? Yeah probably, testing will be tricky and code review for security is hard, types will help reason about the possible states of the system.
There are lots of alternatives to static types: documentation, testing, frameworks, design patterns, code review, pair programming, error messages, and so much more. I'm generally a fan of static types and find them very useful in a lot of development, but they are a tool in a big toolbox.
It’s only hype because it’s imprecisely stated. Static type systems make entire classes of bugs impossible at runtime. The stronger (read less permissive) the type system, the more classes of bugs cannot occur.
Everything's a trade-off, the question is which approach is best for your application. Your average website doesn't warrant as much rigor as a Mars rover.
Dynamic typing only increases development speed for the first few thousands lines of a solo programmer project. After that it is, in my personal experience, a significant drag on development speed.
Furthermore, dynamic typing makes modifications and new features significantly harder to write. Turning compile-time bugs into runtime bugs is a catastrophic decrease in development speed.
That’s my personal experience at least. YMMV.
> dynamic typing makes modifications and new features significantly harder to write
Depends on what you're doing I suppose.
Adding a parameter to an object? With dynamic typing you just add it to the object in literally any location, no issues. With static typing you might just need to refactor half your codebase if you have lots of interfaces. Have fun resolving merge conflicts with your team.
Not only that, but (in Java as an example) serialization codes for objects will change unless you planned for that previously (you didn't), making old objects impossible to load. A completely new bug that's created solely by static types. And it's not the only one.
Maybe it's an issue with Java. In Rust if I add a new field, I just wrap it in an Option.
When de-serialized, old objects have a None, new objects with the field have a Some. When serializing it's just `Some (5)` instead of `5`, if it's an int.
No runtime crashes from nulls, either.
1) Statically typed languages with inference don’t require time spent writing signatures.
2) I know I’ve spent time chasing down bugs in dynamically typed software that would have been caught by a type checker.
3) I also know I’ve spent time writing tests for conditions in dynamically typed code that wouldn’t pass a type checker.
What are some strong examples of this? Haskell does an amazing job with it. Java technically supports some amount of inference, but it doesn’t reduce verbosity by all that much. Apart from those I haven’t run into it.
Type checking also introduces its own set of additional bugs by the virtue of object incompatibility, that do not exist at all in dynamically typed languages (or are handled correctly every time by the compiler/interpreter automatically).
Take as an example exchanging objects over sockets, rest, files, etc. Whenever the object definition changes in another piece of the software stack the statically typed parts will crash upon receiving the updated objects, even if it's just one new param added that would've been fine otherwise if dynamically typed (or say change from a float to a double which can be irrelevant). A nightmare in systems with lots of moving parts.
One might say that's working as intended, and it of course is, but it also forces you to fix and recompile all of that for no real net gain. Hence the longer dev time I mentioned.
I've spent years working with statically typed languages, and I honestly don't think I'll ever go back to them in any professional capacity.
Anecdotally, this is not true for C++ using JSON or msgpack, since those are self-describing formats where extra fields are safe to ignore.
And it's not even true for Rust using serde, which writes the serializing / de-serializing code for you. serde_json will also ignore unknown fields when parsing, and you can preserve the original object as a `serde_json::Value` in case you want to pass unknown fields downstream as opaques.
Protocol buffers and Flat Buffers also have solutions to this, and all 4 of these formats are pretty popular in both static and dynamic languages.
Even if you write a custom TLV format, this is not that hard to deal with.
Was this common in the static code you worked with? You weren't just casting objects to `char *` and doing a `memcpy`, I hope?
Now, being able to assert that a field is present in an object is a basic and valuable use-case for static typing. However, the developers felt that even this basic level of static type checking added too much friction whenever they had to update systems.
https://capnproto.org/faq.html#how-do-i-make-a-field-require...
You appear to completely misunderstand what static typing means. Runtime crashes due to type mismatches are exactly what static typing prevents.
(Well you could, but then you'll be introducing a whole set of problems if the other system has a static type system that behaves differently from your application's static type system)
Aaaahg it burns. I basically am going through and first annotating, then upgrading.
Interestingly enough, py2 was smart enough to automatically load things in the correct format. Loading a binary file? Get bytes. Loading a text file? Get a string. Py3 has worse functionality now, all for the sake of consistency.
I am not sure that that is correct. Type inference significantly decreases development time imho, and access to compiler errors means significantly less testing time because a certain class of errors is avoided.
Personally, I find that my productivity w/ Haskell was significantly higher than with python exactly because of the type system even though I have 3x the experience in years w/ Python.
I have inherited (hah) a project at work and had to introduce all sorts sanity checks (mypy, pylint, etc) as pre-commit hooks to make my life easier wrt bug hunting.
To use that review as an authoritative statement is ingenuous, to say the least.
> Shower: Researchers had programmers fix bugs in a codebase, either with all of the identifiers were abbreviated, or where all of the identifiers were full-words. They found no difference in time taken or quality of debugging.
That's a very weird take on the statement. The downside of using abbreviations is probably dominated by the difficulty of fixing the bug. The problem I have with abbreviations is that they just become another thing that you have to spend your mental resources on. “What is this variable exactly? Oh, I see.” is just wasted effort.
Usually I don't have long loop bodies, so if the loop body fits in 24 lines, `i` is perfect for the index.
If I'm locking a mutex, doing something, and quickly unlocking it, `l` for the lock guard is fine.
But if it won't fit on screen, it needs a longer name.
And if there's something like a top-level `App` struct, I just call it `a` because its scope is really just `main`, even though its _lifetime_ may be the entire process.
Like, isn't it fine to change:
var basicAuthenticationController = new BasicAuthenticationController()
to: var ctrl = new BasicAuthenticationController()
I find that if code authors do work upfront to contextualize things for me that it helps tremendously. Like do I need to know at every point that this is a BasicAuthenticationController? Or is this the basic auth module and we're always dealing w that controller? I prefer it when engineers set the table for me like that, it helps me narrow my focus to the purpose of the code.wait what. I live in SF and everybody just says SF
I doubt it would be controversial if programmers adopted a similar convention.
Usually it doesn't matter WHAT things are, I need to know WHY you have a variable, and what it's intended use is. That can be explained in the variable name a lot of the time.
I have the same argument against the no-comments evangelists, and wanting to squash all commits into one when they're merging. Yes you can read the code diff, but that only tells you what, not why, something was changed. Why did the API endpoint change? Why do we have to call the payment processor before this event rather than after all of a sudden? Why did the add to cart button move up one div? All very useful information when you have to come back to fix things.
For the no-comments evangelists, I understand the idea is to make the code as self-documenting as possible, and that's awesome, but you can still miss out on the why, and sometimes the Why is entrenched in external business requirements that aren't in the code.
This stuff all feels better suited for a commit message. Conversely if I found all this in code I'd delete it.
I'm an almost never comment person (I use them when things are weird or inconsistent, but otherwise view them as noise) but I've been persuaded by John Ousterhout that code cannot adequately describe abstractions, and you need comments to fill in the gap.
Similarly with variable names, if you don't know i is index, please don't try working on this code. I get some code is a soup of inscrutable variable names and flow control to the degree it looks decompiled, but that's a higher level composition problem, not a "use more nouns" problem.
For me it's more about not wasting people's time. Names are opportunities to divulge minor contextual information that you spent time figuring out, and it would be going backwards to then compress that into shortened names.
i for index is fine, but if you have nested loops or multiple loops it gets annoying.
An example of a comment I'd value is something like a "why bother" at the top of a file doing a lot of in-depth algorithmic work, like a note accompanying a Chesterton's Fence. It makes some sense (though I quibble with this all the time) to design things so as to be as understandable as possible, and then only start optimizing when necessary, and when that happens it's useful to put a "this is pretty complicated, but it cuts CPU time down by 30%. [here](link) are the tests, and [here](link) is the naive equivalent implementation".
Names evoke ideas, and this is very subjective. So "self-documenting names" can turn into "misleading names" in a hurry. The name "x" never misleads because it evokes nothing.
This matches my experience. People say otherwise might need to put more effort on writing better tests.
That said static typing usually catch minor problems like typo faster. But for dynamic typing it's going to catch by tests a bit later. But usually tests run faster with dynamic languages, so it's a tradeoff.
My favorite approach is the mixed approach like TypeScript:
1. Faster feedback loops: Usually statically typed languages compiles slower thus slower feedback loops. But languages like TypeScript can skip type check and emit only to the runtime, making test watcher very fast - as soon as I hitting Cmd-S I can see the result.
2. Optional typing: Sometimes a function signature is going to be 10x larger than the body, which is really a hassle. Sometimes I just skip them, or skip typing module private functions as long as it's properly tested.
Another thing: people working with dynamic typing often forget that often times they are interacting with a system with extreme opinions on types, namely their relational database, which is often times literally the heart of the business. If your language of choice has typing, you can sync that type info (smart ORMs, codegen), and thus reduce friction. With that said, I develop in Ruby and my understanding of where there might be "type friction and potential for bugs" has improved a lot over the years, so I don't mind the lack of types until we get optional typing sorted out.
Shower: keybored, your shower is not cold enough to induce stress and you are too lazy to modify the showering experience in any way
Caveats: Other people may be anti-lazy enough to modify the showering experience
Notes: Check yourself
https://en.m.wikipedia.org/wiki/Diving_reflex
Edit: This is also the time of year to do this, while the tap water is probably warmer and your house is warmer. I wouldn't try to acclimate in winter.
Edit: looked it up. 30min-2hr for hypothermia at that range. Typically I take a 15-20 minute shower, so only flirting with hypothermia.
I love a hot, hot shower. Best advancement in technology in the last thousand years
Interestingly, one of the theories behind why cold showers might be good for you is that they induce a heat shock response, same as a very hot shower.
Granted this is personal experience and I may be simply not careful enough to do untyped refactoring, but few people I talked about that shared the experience.
Studies insisting static typing has no effect always assume the runtime-typed program is less buggy than it really is.
It can still be fine if you're being meticulous and know the codebase well enough, but otherwise, you'll miss corner cases in parts of the code with less test coverage or typical interaction. When you accidentally occasionally have a string in place of a number deep in some data structure, you might not notice for a long time.
Static typing is objectively better than dynamic typing for the vast majority of cases, but you can't capture this in a scientific experiment or study.
The only thing you can do is find people who are similarly experienced, break them up into groups (n=1 is also ok) and ask them to complete a specific project, and see how much time it takes. But even then there are so many caveats. The experiment must be set up such that the presense of extensive library support for a particular task is not a confounding factor. There's also the open question of how do you measure these people's skills before the experiment begins?
Static typing is good because dynamic typing measurably makes me sad. Therefore I want the company to use static typing.
Your feelings are often a good heuristic. If something makes you feel depressed, it's because your mind has gathered enough information from past experience to know the situation is hopeless. It's your mind discouraging you from wasting your energy on a pointless activity.
Oof, that one is a huge can of worms.
My only experience with dynamic typing is really node, which improves a bit with typescript. But I'm still extremely wary of it because it doesn't give you runtime guarantees. It irks me a lot that you can just do JSON.parse, declare any type on it and call it a day. I firmly believe that static typing increases code readability and catches low-hanging fruit, but I'm going to have to sieve through that research it seems.
As the article says, this does not appear to be actually true when you go and check.
Either that means it's hard to measure, or it means the effect is not actually there.
Either way, this means that we cannot say that static typing reduces bugs.
> all massively complex systems that I personally have worked with that run for decades transacting billions$ etc without flaws are all statically typed
Did you try building the same systems without static typing as well and see what the difference was?
Or are you just saying that you built systems with static typing and they were successful? That tells you nothing about static typing except that it doesn't prevent successful programs, which is a very weak thing to be able to claim!
Lots of people, like you, think static typing is very important for developing software, but when someone challenges this and says 'can you actually show that?' they have never been able to. At some point you need to reconsider if it's actually the case.
I lean more to dependently typed languages than dynamic as I simply saw only misery, but, again, this is not saying anything, it is just my experience. I am a bit afraid though, because there is no clear measurement, it is just anyone’s experience.
Edit; which actually would be fine; if it works for you and your team, company and business goals…
If you think you really can see evidence for it then I'd encourage you to write up a paper and submit it for peer review. I guess when you sat down to write it you'd suddenly realise that you don't actually have any hard evidence.
> cannot really imagine writing large systems in not statically typed systems
Ok but lack of imagination is not science. Some people cannot imagine a spherical Earth, but that doesn't make it untrue.
For example I work on a system in a dynamically typed language (Ruby) that successfully handles tens of billions a year, so we know that it is possible. (We are adding optional static typing to it, but it was written without it.)
Sure, but that’s what I said in the first place. It all seems there is no evidence, not even empirical for either way. I was looking for any, if there was. MS or whatnot must have something no?
I would say in summary that no, we don't have any evidence. Evidence that people have tried to present has been found to be flawed in peer review.
Why, if not for reducing bugs?
Maintainability, tooling, and some people would argue for reducing bugs - but I'd challenge them in the same way I challenge you - can they prove it? And I would guess that they cannot.
If the person I was replying to was saying that they could not imagine a large system without types, and well they don't have to since I gave them a real example.
Of course not, to do that you’d need a way to detect all bugs in a program.
If we had an effective way to that, we’d be using to remove all of the bugs from our programs instead of looking for ways to avoid bugs.
(only slightly joking)
Many older statically typed languages (Java/C/etc.) don’t enforce null safety. Tony Hoare called it his ‘billion dollar mistake’.
https://en.wikipedia.org/wiki/Tony_Hoare#Apologies_and_retra...
Thankfully more modern type systems (Swift/Kotlin/F#/Rust/TypeScript/etc.) now do, hugely reducing the likelihood of runtime errors.
As a mostly dynamic language developer I never saw the benefit of Java-like static typing but more modern type systems seem much more useful.
You say this like it's a fact... but again nobody has been able to show this in a proper scientific study.
Write up your experience with null safety as a paper and submit it, because this would be a world-first if you can demonstrate it.
We do have proofs that certain classes of bugs are impossible given a certain type system.
I'm with GP, I'd really like evidence of this before people continue shouting things which aren't proven.
So if I do need a null value, I can make the field optional and the type checker will help me figure it all out.
The only way in which strong static and dynamic typing could produce the same number of bugs would be if strong static typing resulted in introducing other bugs, ones which wouldn't be introduced in a dynamically typed language. Proving that would probably require a scientific study ;)
The function will generally only work on a subset of types for given variable.
If I don't check the type of the variable in the function, the function will not behave as you might expect, e.g. silently fail or crash.
If I do check for every possible type for a given variable:
I may not have a good way of handling certain types being passed in, I may be forced to either log something out, create a run time crash or even have the function silently fail. All 3 are bad run time behaviours.
If I am checking for every type in the functions, then using static typing would cause the failure at compile time so the bugs could never exist, but also being significantly less verbose than the dynamic language equivalent.
I'm saying it because it's a fact. I'm not giving you an opinion - it's a falsifiable fact that you can verify for yourself - we have as an industry not been able to give any good evidence for static typing reducing bugs that has stood up to peer review.
You're presenting arguments for why you think there should be evidence... but when people look there isn't actually any evidence. Maybe your arguments are not sound for some reason that we don't understand, or maybe we are unable to measure the effect.
https://www.microsoft.com/en-us/research/wp-content/uploads/...
I'm not sure what evidence your fact is based on, feel free to present some evidence for it.
The linked meta-study in the article we're commenting on.
https://danluu.com/empirical-pl/
"under the specific set of circumstances described in the studies, any effect, if it exists at all, is small"
And this question has generated the most rebuttals and retractions I've ever seen for flawed studies in computer science - it's notorious.
Remember this?
https://dl.acm.org/doi/10.1145/3340571
> Well, this paper found that a conservative underestimate of 15% of bugs would be saved through static typing
Yeah does look that one shows a larger effect and has not been rebutted or retracted. One positive measurement against many negative measurements.
We haven't measured an effect but that doesn't mean that we can't reason about this in other ways. The lived experience of people that do this thing for a living can be an important resource. You could argue that we are getting in to soft social science here, but is there value in what the collective of trades people think about their tools?
We didn't know the science behind steal for a long time, but we still figured out how to make it and that it holds an edge really well.
Empirical measurement is not the only way to discover value.
Lived experience works to some extent, eliminating failures and making progress, like with your steal example, but that doesn’t mean it can pick optimal options from a list of options that all work to some reasonable extent; aka dynamic typing does also build software.
This is true for many functions defined in statically typed languages too. Just because your function says it works with an integer, doesn't mean it's necessarily going to work with _any_ integer (think `1/n`). Very often the type used isn't narrow enough.
Modern dynamic languages understand this, see for example Clojure's spec[1].
Does planning a project mean faster delivery?
Does thinking about architecture up front reduce refactors and technical debt.
Does estimating lead to faster software development?
There is a lot we have to decide without evidence in how we develop software.
Maybe we just pick the things that feel right and make us happy.
Does that plan include the client changing their mind 10 times before delivering? Or finding out that some approach doesn't work and that pivoting is needed halfway through?
Likely depends on how smart the planning is.
> Does estimating lead to faster software development?
Unlikely, because engineers are then busy ass-pulling useless time figures over and over instead of actually working on the project.
All of these points have high random factors attached to them regardless, so you'd need a pretty big sample to say what generally works best.
We have to carry on regardless and make decisions.
The nuance is what you do, why you do it and is there cultism or cargo cultism to the decisions.
Quality of life benefits are where static typing (and in some cases performance) really shines.
You should look at zod [0], which validates data with inferred types based on the validation having passed, so you do have the guarantee of the type being correct at runtime.
[0] https://zod.dev
But clojure’s REPL, babashka, and clojure.spec are sorely missed.
To understand your whole program all the time at scale is probably something Rich Hickey is capable of but I definitely am not.
1. If I don’t understand what code does just delete/noop it and see what tests fail.
2. Write my assumptions into new tests to be very loud people if someone wants to break them. Easier said then done with frameworks/callback patterns that want to call my code from god knows where but really powerful when you’re pulling a value.
3. In the case of callback spooky object at a distance stuff be strict about what you accept and fail loud.
That said, I certainly agree it comes up less than it does in, say, JavaScript
It is very hard to make large programs simple, but lisps are some of the best PLs for that, especially if you stick with immutable data.
Bugs are reduced by monitoring, testing, and reducing how much code there is so there's less surface area (which also makes monitoring and testing easier).
You reduce code by refactoring.
For some definition of "instant".
I can often run my Python or JavaScript test suite faster than Haskell/GHC or Rust/rustc can type-check a module.
And it can take a while to understand Haskell, Rust, or C++ error reports.
(Of course, Python and JavaScript startup time and execution speed can hurt refactoring too. And AttributeError-s and random undefined-s aren't always easy to debug.)
This is a great point.
I wonder if there's a study on bugs-per-line-of-code, how this changes across project sizes, and whether refactoring changes bugs-per-line-of-code.
There are so many variables involved that I see this being difficult to prove empirically, but anecdotally it doesn't ring true.
Most names aren't particularly good - especially when someone tries to make them sound like a full sentence. My experience is that at around four words they start getting less accurate. At seven I would probably read more from an interpretive dance about the function in question.
I do not get this one. It's incorrect by definition: "Big data refers to data sets that are too large or complex to be dealt with by traditional data-processing application software. "
If you do not need a big data system, then it's not big data.
> Shower: A review of all the available literature (up to 2014), showing that the solid research is inconclusive, while the conclusive research had methodological issues.
Which research is supposedly solid?
Skimming through the article it seems to me that proper controlled study would be prohibitively expensive and that no controlled study comes close to real world conditions of multiple people developing and maintaining a larger codebase. Conditions, where supposedly (according to anecdotal evidence) static typing would make a notable difference.
It’s been an immensely enjoyable experience getting into it, but with a non trivial startup cost.
My team has generally adopted it (at least to some degree, and in a language which half supports these patterns), but I’m sure this coding style erodes in favor of something that feels more imperative when we move on.
True and not a hype: Compared to other languages, Go has a runtime and built-in concurrency primitives that makes it easier to write concurrent code.
> and is less prone to bugs and memory leaks.
Who is actually claiming this? I've got a collection of good Go resources and nowhere something like this is stated. In contrast, some actually make it very clear to be very careful when sharing memory or actually better no sharing memory at all.
Concurrency is conceptually hard in Java, C#, Javascript, Python, C, C++ ... there's no reason to bully Go.
Rust makes the claim of "fearless concurrency" and the compiler indeed saves us from race conditions, but those are not the only concurrency related bugs, e.g. deadlocks are very well possible in Rust. Therefore Rust would be a proper candidate on this "Cold Showers" list.
https://www.agitma.nl/wp/wp-content/uploads/2016/07/Dilbert_...
> Hype: "Static Typing reduces bugs."
Just because the author has not found proper research proving it doesn't mean it's false.
A great type system with sum types prevents tons of bugs. I wonder how anyone can question this.
> Hype: "Identifiers should be self-documenting! Use full names, not abbreviations."
Same as above. To fix bugs you need to read and understand the code. Not having to map abbreviations to your mental model reduces overhead. I've seen code bases where u was used as an abbreviation for user and users in different methods. That was not a fun code base.
Some developers and teams don't need the crutches. Type systems are also an encumbrance that comes with a non zero amount of issues. They slow development velocity down for a theoretical trade-off boon.
I've seen a type system take down a production system where a simple coercion would have functioned just fine. Literally the only thing wrong was the type defined and the code refused to run.
I would argue that static typing reduces bug complexity/creep. Mainly from working in environments without proper testing during the js days, TS was rough at first, but did help.
Interesting though, I would say that dynamic typing allows you to shoot yourself in the foot a bit more, especially over teams who might need to interact with source later. I agree with you to keep the dynamic ones (that are inevitable) simple though.
He treats them like an incomplete requirements document, but instead they should just contain the bullet points so that the people who talked about the issue still remember what they talked about. In a typical requirements document the communication form is written, whereas in case of user stories the communication form is verbal and the document is just there to help people to remember.
Shower: Unfortunately, there is no conclusive peer-reviewed evidence that it is, in fact, a good morning. A randomized trial found that many mornings are bad.
Caveats: Only applies to mornings. No rigorous paper exists for evenings, so this remains an unknown.
Claim! Thing good!
Wait! Thing has downsides!
It's not thorough comparison, it's just showing that everything in life has tradeoffs, which is obvious.