Loci: A C++-like systems programming language
loci-lang.org
loci-lang.org
What I think is particularly need is making lvalues and rvalues explicit constructs in the language: http://loci-lang.org/LvaluesAndRvalues.html, as well as being able to redefine the behavior of certain implicit operations relating to copies and moves: http://loci-lang.org/ImplicitOperations.html.
class SomeType {
public:
SomeType(int a) : p{std::make_unique<int>(a)} {}
private:
std::unique_ptr<int> p;
};
This version supports move operations as well.I do wonder how they're planning on avoiding overhead with the garbage collector, though. Once you have a single garbage collected object on your heap, it seems like you'd have to start scanning the whole thing to make sure the garbage collected object is still alive.
L = total size of live objects in heap
H = maximum size of heap
P = time between collections
R_a = rate of allocation
R_s = rate of heap scanning
It's easy to define P and R_s (assume a mark-sweep collector so we can ignore copy reserve): P = (H - L) / R_a
R_s = L / P
Substitute and simplify: R_s = L / ((H - L) / R_a)
R_s = L / (H - L) * R_a
In other words, it doesn't matter how big your heap is. If you make your heap twice the live size, the amount of heap scanning you do will always be equal to your allocation rate times some constant. If H = 2L, that constant is 1.So the way to improve performance is to cut down on that allocation rate. If you use RAII for most objects, and GC only occasionally, your allocation rate, net of non-GC allocations, will be very low, which means your scanning rate will be very low. You'll have to traverse all live objects every GC cycle, but you'll rarely have to do a GC cycle.
R_s = (L + H) / (H - L) * R_a
For a copying collector: R_s = L / (H - 2L) * R_a
If H is a constant multiple of L, then your scan rate will be a constant multiple of your allocation rate, independent of heap size.In practice, you won't sweep the whole heap after each GC cycle, but will do it lazily: http://www.hboehm.info/gc/complexity.html. The point in that article about marking time dominating allocation/sweep is even more true today. Allocation and sweep access memory linearly and are easily parallelized. Meanwhile, marking accesses memory randomly, and the maximum concurrency is dependent on the shape of the live object graph.
Say we use your example of H = 2L. Then we have R_s = L / (2L - L) * R_a, or R_s = R_a. Sure.
Now say we double the heap: H = 4L. Then we have R_s = L / (4L - L) * R_a, or R_s = R_a/3. For a fixed allocation rate, increasing the heap makes the scan rate go down, i.e. we scan less often.
It seems like you're assuming that the live set is a fixed fraction of the heap size, but that doesn't make sense to me. The size of the live set is a function of what the program does, not its heap size.
* Use the garbage collector. * Don't use the garbage collector.
I'd expect most users to choose option (2), however the idea is to support garbage collection for those users who find it to be worthwhile.
I say this choice is theoretical because the current implementation doesn't yet have a garbage collector (this is not a high priority task); ultimately I want to provide a second implementation of std.memory that supports garbage collection and when users build their project they can choose whether they want the garbage collected implementation.
I haven't got any experience with Go development, but from those I know that use it, I hear nothing but promising things about Nim. "Modern" is a word that gets used.
Between Rust and Nim, we're starting to see some really clever new languages that fix common pain points. Both of these are so neat it makes me feel as though we're in a programming language boom right now.
It gives you a high quality set of components that cover a big bunch of the things you need when implementing a compiler. Energy can be spent in the language itself.
Oh, and I guess the constructor syntax, which again I wish Python would adopt. I end up with too many "self.x = x" lines in my Python __init__ methods; pure syntactic ceremony.
The docs [1] describe that objects (of class type) only contain data (i.e. object fields), and no vtable pointers. Only when an object is cast to an interface type is a vtable generated and passed along with the object pointer in a fat pointer.
However, instead of just structuring the vtable as an array of pointers to methods (ordered e.g. by name), Loci instead generates a hashtable with method resolution stubs in case of conflicts. I don't understand why - since the target interface's type (and methods) are known at compile-time, it would be just as easy to fix the ordering of the methods and use the array approach (like in C++), instead of using a hashtable.
"This design also differs from C++ in that vtable pointers are not stored in objects, but are part of an interface type reference. This decision is particularly appropriate to the language, since Loci doesn’t require classes to explicitly implement interfaces"
I hope this helps.
Basically I've had a lot of experience working with C++ on various projects and while I like it very much I've also observed its weaknesses (right now I'm struggling against slow build times); after a while it seemed logical to build a language that would solve a lot of these problems so that I wouldn't have to face them over and over again for each project.
The name mangling scheme is similar to (one of) C++'s so that could probably be made more compatible. However the vtable mechanism is completely incompatible and given the reasoning in the docs (to support structural typing), I don't think it's a big stretch to say that it's a fundamental incompatibility.
TBH I'd be quite surprised if many (any?) languages will aim for C++ ABI compatibility that's much beyond simple extensions to the C ABI. Even ignoring the lack of a standard C++ ABI, once you start getting into virtual functions, exception handling and, especially, templates, you end up constraining your language to a point where you're probably better off just using C++.
I may well be wrong on this, LLVM helps a lot with the grunt work of name mangling, exceptions, trampolines etc. so it may be possible to get to a sweet spot that supports a significant number of important libraries without hamstringing the language too much. Not going to hold my breath though...
I digress... just the usual problems with wiki-type sites.
Could someone expand on this for me? What is the performance overhead typically associated with C/C++?
In C++ virtual functions are implemented as vtables attached to each object. As such you need to retrieve the object (possibly from the heap), there's a function pointer indirection, there may be trampolines if you're using multiple inheritance. That sort of thing.
The author has a nice piece in the documentation regarding how dynamic dispatch is implemented in loci in order to support structural typing. He explains the design choice, how it differs from C++ (a vtable at each interface reference call site) and gives some qualitative estimates as to how the performance would differ. It's worth the read.
C doesn't have dynamic dispatch built in to the language. If you want it, you'd have to code it by hand.
I found that the best thing when I needed to be "productive".
syntax is a huge help, it's a shortcut to being able to understand what the program does. syntax is good.
Execution environments, on the other hand, are tailored for a class of problems, incorporating useful abstractions. Yes, I know it is not so evident in a world where we have general-purpose programming environments by the dozen; when we talk about domain-specific languages, this makes a lot of sense. They can only gain expressiveness if they can represent bigger abstractions with fewer words.
> A language that doesn't affect the way you think about programming, is not worth knowing.
- Alan Perlis