It's an illusion to think that in a language with a garbage collector you do not have to worry about memory management. To give an example: take a function that takes a substring of a string.
1. If a substring function returns a string with the same underlying character array (just with different start/end indices), you may be wasting memory with very large strings of which you only care about small substrings.
2. If a substring function returns a string with an underlying character array that is a copy of the particular substring, you may waste a lot of memory when you have a lot of (partially) overlapping substrings. Moreover, the substring operation is O(n) rather than O(1).
In other words, in any non-trivial application working with strings you want to know the allocation behaviour of common string operations.
String someLargeString = ...;
//...
String interesting = someLargeString.substring(10, 20);
someLargeString = null;
Now assume that someLargeString is garbage-collected after the last statement. Is the backing character array of interesting 10 chars long or the size of the backing array of someLargeString?You don't know, unless you know what the implementation of String#substring does. Suppose that someLargeString's array was a couple of megabytes, you could be using 20 bytes of memory or megabytes for interesting.
(Note that in the case of Java they switched from a O(1) index-based slicing to O(n) copying some time ago.)
Can you give an example of a (somewhat widely-used) library/runtime that does this?
I think this is harder than it initially seems. How does the GC know that if you have some type with a backing array and two pointers (or indices) that I don't want to move one of the pointers to some other valid position later?
I am not arguing that it is not possible, but this seems difficult in the general case, unless you can explicitly provide hints to the garbage collector.
At some point this for considered for JS engine in Firefox, but that was before introduction of ropes, small strings that holds string chars inline and various other optimizations like recent switch from UTF16 to byte per char for strings with Latin1 characters. That and the fact that Firefox flattens all strings that name properties of objects, this optimization is unnecessary on the current web.
Right, but that is a specific case where a built-in data type is optimized. It's not difficult, because the GC knows the semantics of a specific data type. But this does not work in the general case of slicing of library/used-defined data types.
A good one can detect that only a range of allocation is alive and release unused part of someLargeString.
I do agree that ropes are nice, though they may not be acceptable if you want O(1) indexing or slicing.
Exactly. Which is the core of "not worrying about it". When it becomes a problem (Such as an out of memory exception) you worry about it. There is no illusion. The time you don't worry about it is when you write it. I'd write it exactly like you did above (except I probably wouldn't null it), and wouldn't think of it again until it became a concrete problem.
You are absolutely right that a GC didn't magically make all memory issues disappear. It just hides them.
But that tradeoff of memory footprint vs performance implemented by Oracle is not in the API documentation[1] for Java 7 nor Java 8.
One would have to stumble across other sources such as unofficial blogs[2] written by non-Oracle employees like Mikhail Vorontsov. That blog article then triggers an explanation from the actual coder that's buried in a reddit thread[3].
Hunting down why your business-domain java code is now suddenly slower because of hidden GC tuning is taking time away from "solving the problem."
Your rebuttal consists of:
-- repeating the "premature optimization" meme which is not relevant in this case.
-- recommending to look at the API docs, which doesn't even have the information!
[1]v7: https://docs.oracle.com/javase/7/docs/api/java/lang/String.h...
v8: https://docs.oracle.com/javase/8/docs/api/java/lang/String.h...
[2]http://java-performance.info/changes-to-string-java-1-7-0_06...
It's not a contrived example. When Java changed this exact behaviour it caused a lot of pain for people. Code that previously ran in a reasonable amount of time suddenly became impractically slow:
https://news.ycombinator.com/item?id=9862556
>may require a quick look at the API docs
The behaviour isn't specified by the spec. That is why they were able to change it from one release to the next.
>premature optimization and all that.
The "premature optimization" meme you are parroting contains the implication that at some point optimisation will be necessary. When that happens you have to think about memory management, even in a GC language. The string example given by danieldk is just one of many ways in which that can happen.
That's exactly the point: caring about the actual memory used by interesting may not be as important as other parts of the application, or may not be important at all for now. You should only care about that when it does prevent pose a problem (performance, impossibility to evolve features, ...) (all while making the initial code flexible enough to be potentially changed in the future)
Then you are setting up yourself for a lot of trouble.
Another part of memory management is thinking about ownership (to which the substring discussion is very much related) and it pains me to see how many bugs are caused by code like:
Foo getFoo() {
return myFoo;
}
Where Foo is some mutable class and the invariants of the wrapping class suddenly don't hold anymore because some outside code is mutating the Foo instance through the pointer/reference.Throw in some amount of concurrency and you get something that is not only hard to debug but also difficult to reproduce.
tl;dr: in a GCed language you need to think about memory management and ownership as well.
(Note: I am not arguing against garbage collection.)
Ah, there's your problem :)
[1]: At least between GC and not-GC. Rust is a new third option IMHO.
We're not getting more cycles any time soon for sequential programs.
Other peoples problem constraints are not necessarily the same as yours.
Most GC problems can be pinpointed to poor application behavior and many are due to bottlenecks in GC (e.g. too much work done in sequential phases, etc).
I'm talking 10-50x performance increases here. Unfortunately managed languages(and GCs in particular with everything being a ref) make this much harder than it should be.
In this settings GC is not necessary beneficial as it brings an extra unpredictable latency/pauses. All is fine when it does not affect the application, but when it becomes the problem, it is hard to fix and may require a significant rewrite or a very time-consuming tuning of GC parameters.
Even for most normal web apps a full stop the world GC is a common occurrence and a deadly one.
I love C#, and I love not having to write code that worries about memory (99.9% of the time anyway).
But if I was writing the .Net Runtime then I wouldn't be writing that in C#, because the runtime _does_ have to manage its own memory - and Rust looks like a very good language for writing that kind of thing in.
Oberon's GC is written in Oberon:
https://www.inf.ethz.ch/personal/wirth/ProjectOberon/Sources...
Module SYSTEM is like Rust's unsafe {}.
Singularity and Midori GC's were written in their C# dialects.
With luck C# 7 onwards will get all the nice features from Sing# and System C#, as per Joe Duffy's blog.
Then using C# will be no different than Modula-3, in terms of features for systems programming.
There is a huge category of bugs related to sloppy fudge-factor-based programming. C++ has great tooling but it's a nightmare on an organizational scale. If I was building an engineering org and had no other alternative but C++ I would be in shambles worrying that no-one does anything horrible, no matter how good the engineers were.
I write C++ for my day job :)
We can easily fix tools and processes, we will never fix people. Sure, all languages allow to create bugs, but some orders of magnitude more then others.
> kernel is written in C and is the most successful/widely used piece of software
end every release finds bugs in filesystem drivers, and buffer overflows and security issues and breaks display drivers. Seriously think of all the things kernel developers could do if they didn't waste years of work on tracking invalid pointers and off-by-one errors. Over and over and over again.
It allows managing arbitrary resources more easily, so it is clear who is control of things at any point of time, e.g. if a function takes a File, who should close it (if the callee closes it, or if the caller needs to) is just a matter of documentation, while it is expressed naturally in the Rust type signature, and is hence checked by compilers. This isn't unique to Rust, e.g. it's basically the same as C++'s RAII but it is checked far more precisely by the Rust compiler: it is up to the programmer to not make a mistake in C++. http://blog.skylight.io/rust-means-never-having-to-close-a-s...
It is also an important part about getting guarantees for concurrent code, as it allows fairly precise control over sharing (or not). This not only means Rust code without `unsafe` is free from data races, but also means that, for example, one can be sure a program built on message passing (ala CSP/Go) is not accidentally sharing things that shouldn't be shared. http://blog.rust-lang.org/2015/04/10/Fearless-Concurrency.ht...
Writing in a language with proper types helps a lot. Being forced to acknowledge every error returned is great. The only time I had to think about memory management was during integration of database connection pooling and the web framework, but that's just because I was one of the first ones to try. Otherwise memory management doesn't get in the way at all.
My hypothesis was that the type system is be able to benefit practically all domains. Personally, so far, I've become a cautious believer.
> To be clear, a major goal of this project -- outlined in the earlier blog posts -- is "GC as Interoperation Feature". That is, we want to embed Rust code into contexts with existing GCs (such as the V8 or SpiderMonkey Javascript engines), and so Felix is trying to find ways to accommodate a relatively wide range of GC approaches, with maximally safe, efficient, and ergonomic support on the Rust side.
Might be a while off before good support for custom GCs for pure Rust code, but it's definitely being thought about.
[0]: https://github.com/rust-lang/rfcs/pull/1398
[1]: http://blog.pnkfx.org/blog/2015/11/10/gc-and-rust-part-1-spe...
[2]: http://blog.pnkfx.org/blog/2016/01/01/gc-and-rust-part-2-roo...
[3]: https://www.reddit.com/r/rust/comments/3zqtoc/gc_and_rust_pa...
Jai [1], Jonathan Blow's language in progress, has a ton of customizable features. For example, the memory allocator is quite straightforward to swap natively without any hacks.
[1]: https://github.com/BSVino/JaiPrimer/blob/master/JaiPrimer.md
I'm also interested in one particular kind of memory abstraction: abstracted access to memory mapped files. Here the problem is that the base address of the memory mapped area might change whenever you perform the mapping. And you would like to use standard data-structures (like maps, lists, etc.) inside the memory mapped area.
A library that can do that in Rust would be great.
The main problem is not opening the mmapped file. The problem is having a kind of pointer that allows arbitrary offsetting, and using this pointer inside standard data structures.
I can see why you might want to know in the case where you're representing an external resource, e.g. an open file, but otherwise, I don't see any reason to care about it.
In face, with persistent data structures, you can't know. The same is true of share (immutable) data with multiple threads. You can't know the lifetime.
In the space of core engine stuff, I agree.