Checked C: extension to C that adds static and dynamic checking
github.com
github.com
Checked C adds new types to C:
• ptr<T>, the checked pointer to singleton type,
• array_ptr<T>, the checked pointer to array type,
• T checked[N], the checked array type,
Many people have been down that road. GCC had "fat pointers" for subscript checking years ago.[1] That never caught on. Walter Bright had a proposal.[2] I had a proposal. Microsoft had "Managed C++". It's technically possible, but retrofitting old code rarely happens.[1] http://williambader.com/bounds/example.html [2] https://news.ycombinator.com/item?id=8509155
Hence why Microsoft went C++ and .NET Native, keeping C compatibility just as far as ANSI C++ requires it. And support C++ kernel code since Windows 8.
Why the NDK is such a pain to use on Android, with a very tiny API surface.
Why all Objective-C improvements are now only related to improve interoperability with Swift.
So maybe if Microsoft decides to also force Checked C on Windows, it might eventually work.
Azure is a money printing machine, the "Year of Desktop Linux" will never happen, hybrid tablets/laptops with Windows are being bought instead of Android ones, IoT devices for healthcare or ticketing machines run on Windows, on the enterprise Linux mostly matters on the server room, they own PC/XBox gaming.
So just like Apple and Google, they can do whatever they feel like it, because non-technical consumers aren't going to suddenly start buying alternative systems, just because devs don't like Apple, Google and Microsoft roadmaps regarding which programming languages one is allowed to use on their systems.
how does this happen without diverging from the history of backwards compatibility?
Yeah. A more realistic solution might be auto-translation of existing code. Here[1] is an example that replaces declarations of arrays, and pointers used as array iterators, with macros that enable the code to be compiled as either unsafe C code or "safe" C++ code (bounds checked and iterator use-after-free safe).
Determining in general whether or not a pointer is being used as an array iterator, or a pointer to an array is not trivial. It sometimes has to be deduced from context. Which may be why it doesn't seem to have been done already?
[1] shameless plug: https://github.com/duneroadrunner/SaferCPlusPlus-AutoTransla...
Very interesting to add additional compiler support to am existing language, in a sense. Sanitizers and static analyzers have made C/C++ much more bearable, and this adds to the pile!
Actually, that's exactly what happened.
Has that situation improved?
Here's some of the better ones I've found:
https://two-wrongs.com/tags.html#ada
The official reference manual is freely available too, but it's quite dense.
>>> fit the the same niche (high performance, real time, and very important software
I don't see any community emerging from the web any time soon :-) To me this niche means "avionics, military, nuclear, medical, spatial". Not your week-end project.
C's history is tightly coupled with that of Unix, it's always been rather open and "hacker-friendly".
I'm not sure I agree. Sure, it's reasonably easy to write your own implementation of something that closely resembles C, but C is also infamously dependent on implementation specifics and it, last I checked, had a paywalled standard. That's the opposite of open and hacker-friendly.
The primary reason C feels open is, I suspect, precisely its Unix coupling. If you want to read the sources to your Unix-like OS, you're likely to find sources written in C. You'll also find system libraries with C calling conventions. If you crack open your OS, you find C. It's easy to see the connection between "open", "hacker-friendly" and "C" – but the connection is between your OS and your knowledge of C – not C itself.
You're right (and it's a bummer that the C standard is behind a paywall) but don't you think that having a large open ecosystem is more important and significant than having an open technical standard, especially when several high quality open source compiler implementations already exist? Having an open standard won't do you much good if all the tooling and libraries are closed source.
To start, it isn't really a C-like language, is it? It looks more like Pascal (or Modula), which is a different offshoot from ALGOL than C is.
Wikipedia tells me that Ada-0 was based on LIS from the mid-1970s, so a couple of years before the K&R C book came out.
The Ada '83 Rationale, at https://web.archive.org/web/19970407040255/http://sw-eng.fal... , mostly compares Ada to Pascal equivalents, though many other languages are mentioned.
The bibliography lists those languages: https://web.archive.org/web/19970407041806/http://sw-eng.fal... . You'll note that C is not present in the list.
The Ada 95 Rationale does reference C and C++ ("Ada 95 incorporates the benefits of Object Oriented languages without incurring the pervasive overheads of languages such as SmallTalk or the insecurity brought by the weak C foundation in the case of C++" - http://www.adahome.com/LRM/95/Rationale/rat95html/rat95-p1-2... ) but as you can see, it also refers to Smalltalk. The references include Modula (though not Modula-2) and Wirth's paper on "Type Extensions".
It comes across like they drew much more from languages other than C and C++, and were in the niche of "high performance, real time, and very important software", but from a different path than C/C++.
...actually, this is something I'd really like to do if I had more time...
Those seem to be ahistorical comments.
It's no surprise that languages used in embedded and real-time systems would cluster together. But that doesn't reveal intent and influence.
Edit: Silly me. I know why. I visited the HTTP version of the site on those two devices...
They do NOT want an autonomous Ada community to pop up, they just want to attract programmers they can contract. Since they're basically the only noteworthy contributors for Ada, they're impossible to bypass, hence using Ada means submitting to Adacore's whims.
Ada is indeed a very well designed language, and with some recent language revisions, so have you ever wondered why it did not get any traction at all? Well, that's why.
Every 6 months or so on HN or Reddit, someone rediscovers Ada, gets psyched and tells everyone ... and then discovers the licensing / controlled ecosystem issues, and walks away. And we don't hear about it. For about 6 months.
https://www.ghs.com/products/ada_optimizing_compilers.html
https://www.ptc.com/en/products/developer-tools/objectada
https://www.ptc.com/en/products/developer-tools/apexada
Ada it is pretty live in European universities, a constant presence at FOSDEM during the last decades, and high integrity computing conferences.
Also, if you want the industry to care about any C or C++ feature, you have to buy your seat at the ANSI/ISO table.
Moreover Adacore is the only one to have a general-purpose OSS (albeit kind of captive) ecosystem.
Your ANSI/ISO remark is off-topic. You don't need a seat there to use the extensive C++ ecosystem. With Adacore it's either GPLv3 all the way or you pay up. Also, to underline their gatekeeper status even more: they could one day decide to stop releasing their compiler to GNAT, without giving any reason, and there's nothing anyone could do about it.
Thanks to GCC and Sun, and later Apple.
GCC was a toy compiler until Sun decided to start charging for their UNIX development SDK, which made many companies contribute to GCC's development.
Likewise if it wasn't for GPLv3, Apple would probably never bothered to create clang.
> With Adacore it's either GPLv3 all the way or you pay up
Actually I do have issues with people earning money with the work from others for free without contributing anything back.
GPLv3 is quite appealing for free software.
Want to earn money without contributing anything back? Use commercial licensed software.
> Also, to underline their gatekeeper status even more: they could one day decide to stop releasing their compiler to GNAT, without giving any reason, and there's nothing anyone could do about it.
There is always the very latest GPLv3 version, that everyone willing to contribute can carry on using.
With the BSD license, you can fork it all you want, but you have to maintain that fork. For most people/companies the feature is not a competitive advantage. So they donate the code back, and it means they don't have to spend so much time maintaining a fork.
The way I see it is, I'm a working stiff that's used a lot of opensource in the past. I don't mind releasing things as BSD clause to help another working stiff, as long as it doesn't contain my company's "secret sauce."
One would think this is a reason not to inform people of the existence of Ada – but I actually see it the other way around.
With a large enough community, the influence of Adacore would diminish and we would see more independent contributions. Ada has all the right things in place to create a great FOSS community, with an open, published standard, democratic contribution model, GNU support and so on. It's just missing the people to make it truly free.
Also, if you have to recreate a collection of lib from scratch, I have the feeling people would rather contribute to a newer language in the same (or same-ish) space doing just that, such as Rust -- which I think is what is happening, even if Ada still has some distinctive features.
Also, Adacore is not the reason Ada didn't catch on. There were a lot of limitations in the standard library as of Ada95, because Ada focused so much on the embedded space. Because of this, some things that were easy in C/C++ were unnecessarily difficult in Ada.
Frankly, I think Rust achieves better ergonomics without Ada's verbosity, but Ada still has some great features that I wish were in other languages, like ranges.
Disclaimer: I work at AdaCore. Both forces are working in tension in the company: Some people push for the community, some for commercial interests. Of course as a company we want both to coexist, which is the difficult balance to find. The community advocates are getting stronger and trying to push their visions and make it understood. This is a long process but it's making headway. But, anyway, saying that we don't want an autonomous Ada community to pop up is dead wrong.
The libraries we release on GitHub are generally under GPLv3+runtime exception, so you can actually dynamically link against them in code with a different licence. https://github.com/AdaCore/gnatcoll-core for example.
It is very possible, if not ideal, to develop native software in Ada today, with whatever licence you choose.
But a plain *char in C makes no such distinction.
You call free(foo) not free(foo, 12) or free(foo, strlen(foo));
But there's no standard way to ask for this size. Why doesn't strcpy just check with the allocator about how much memory is available and refuse to write past that.
And if the memory allocator doesn't know about the string then you could just revert to the existing behaviour so you're no worse off.
static char buf[128];
and use strcpy() at any point to copy into it, so I don't think statically allocated strings must be literals.The point is not that allocators don't know how much memory has been allocated, but that C has no idea about the underlying allocator, so it cannot ask it about the length of the string as it was proposed.
How much C code is out there that only uses malloc/free? All of those programs could be made safer if the default string methods were able to prevent writing past the end of the array.
Would be nice to be able to tell if something is allocated on the heap or stack in a portable way.
As in being able to ask if and object is thread local or not. The language and run time could support that but studiously does not.
Just like the language could allow you to directly monkey with the ABI but does not.
I looked into that, not available on many many platforms.
I think the inclusion of strdup() shows that it was a mistake to make C strings completely allocator-agnostic.
It is not the case that C has a benevolent dictator with a particular vision and drive for the language. It's more like a small group of people who don't want to mess things up. And so not much changes. This is neither good nor bad to me--it simply is; and anyway, these days, you have quite a few options that do offer automatic strings that can also compile into machine code (Go, Rust I suppose?).
Getting rid of standard string functions would be the best thing to happen to the c language in the last 25 years.
There are two types of c programs. Those that scrupulously avoid standard string functions and brittle programs shot through with security holes.
In other words, if your pseudo-c candidate already suffers from most the interop problems that Rust, D, Nim and Go do, why not just use one of those and at least reap the other benefits they provide?
Your suggested dichotomy is, of course, a little bit false. But I'm sure you knew that when you wrote it down.
It's entirely possible to write secure programs in C, even with standard functions. Writing your own code does not somehow confer a level of security-consciousness that you lacked when sticking to strings.h. (It does give you a wonderful opportunity to write your own security holes that no one has discovered yet!)
I mentioned this somewhere else, but we're in a pretty good place right now with languages; we finally have really solid alternatives to C that can compile to machine code, in both Go and Rust.
You realize that most standard string functions are outright banned by organizations that care about security. As in you're not allowed to use them not even if you pinky swear to be 'careful'
That said, several “high-level” string libraries for C exist that make it safer, more efficient, etc. Some of them are even pretty clever, allocating extra space right before the C pointer to a null terminated C string that it gives you, so it can do its own bookkeeping.
Edit: Actually seems I confused Fortran and Pascal. Also, Fortran uses an internal length variable.
The reasons aren't technical or based on merit but on stubborn ideology.
a) What's involved in getting the C code to be 'safe' - how much annotation overhead, etc
b) How this interacts with C Preprocessor
c) What the runtime overhead is (usually the killer)
We have been trying to publish a paper on this, perhaps one day we will be successful! Our runtime overhead is very good, and our interaction with the preprocessor is minimal.
C, not even C++, is still the only common denominator between languages. The fact that I cannot create a shared library in Golang or Python that is callable from the other is very disappointing.
Maybe I'm wrong and hopefully someone can correct me; I don't have deep export knowledge outside of simply using C exports, but it seems this decision has had long-standing repercussions for people wanting to use FFIs outside of C.
Do try it out though!