The Ć Programming Language
github.com
github.com
The idea here is to let the quickest, shallowest reactions go, and wait for the more interesting ones to show up. With a post like this one, that would mean focusing on the details of the language. It's not that names are irrelevant, but we shouldn't focus on the surface at the expense of the depth—doing that leads to exciting-in-the-short-run, but boring-in-the-long-run results.
Are there any samples of what the generated code looks like in various languages?
I would also like to suggest a project name change as soon as possible.
Also, you could make jokes about the holy see programming language.
> cito has no own garbage collector. You get what the target language offers. If it's C#, Java, JavaScript, Python, then there is a GC. In C, C++ and Swift, there are stack variables and reference counting for dynamic allocations. In OpenCL there are only stack variables.
So you can (and probably should) do almost any memory management as you would in C++. Except I don't see any alternative for `weak_ptr` at the moment.
Well that's already out the door because I almost never use shared_ptr in C++ code. The C++ core guidelines recommend using unique_ptr whenever possible. If you're going to use shared_ptr literally everywhere you do dynamic allocation, you'd be better off using a tracing GC to avoid the extra pointer indirection (and cache miss) with every dereference.
It seems to me that Cito should have manual memory management because a manual memory language can be trivially mapped onto a GC language (just turn every free or delete to an nop), while the inverse problem is intractable in the general case.
A shared_ptr can be optimized to unique_ptr if it's not copied.
Not sure what you mean by "a GC language". Most cito targets are garbage-collected.
Even though I put HN at 120% font size
At least the github page heading has a large enough font size to see it
I don't have one anymore to test though, but I used to use the mac layout on Linux for a while exactly for composing using alt.
Or maybe this has changed (back) in a recent OS update? I'm running 12.0beta on this machine.
The dead-keys like Opt-E for acute only work with a small set of letters in the standard US keyboard layout, although it's possible for a different layout to support more -- subject to the combinations existing as precomposed characters in Unicode. (For other combinations, you'd need to enter a combining accent after the letter, and rely on the font to supports placing it properly.)
edit: Looks like it's supposed to work, I don't know what I'm doing wrong: https://help.ubuntu.com/community/GtkComposeTable
edit: hold on, that's only in Firefox. Anywhere else, comma+c does ç and apostrophe+c does ć
If you’re not in en_US you can find other locales and their Compose sequences in the nls directory.
Yet your sentence, clearly written in English, contains all these characters.
You are supposed to be able to type them if you write in English. Also, if you are in polite company, it's alright to drop the occasional Greek word. So here we are: if you can fully write in English, you can use all these characters.
Ascii-only is the modern equivalent of all-caps. Something archaic, still used by old people with hopelessly limited input devices.
Forgive me if the above is incorrect, but in case it is:
The majority of your world's fellow humans, wishing to be comp developers, already spent months or years (combined) of their individual lives in order to learn English as non-native-speakers to do just that.
Your "whining" (yes, it's intended to be somewhat impolite, but I don't think it's unreasonable given the above) that, in case Ć becomes the next-gen well-known programming tool, you'll have to spend maybe a couple of hours to update your keyboard mapping or learn some other way of quickly typping 'Ć' invokes 0 of my (and many others') empathy. Though many were helpful to suggest you how to do that quickly/correctly.
I hope my words will be deemed as 'fair enough given the context' :)
And would the onus not be on the screen reader to not produce gibberish for a simple letter?
> creating a lookup table for all the unicode material out there might've been considered impractical or performance-hitting for the developers.
just doesn't ring true to me in any way for current software. I understand that people can be using older software, which is why I strive to restrict myself to ASCII as much as possible for the widest possible support for my users, but my software also supports unicode identifiers, up to and including a whole unicode table to talk about confusables[1]. And not all TTS software "ignores" characters, which is why people advice against using 𝑓𝑎𝑛𝑐𝑦 unicode because it doesn't get read as text but instead each character is described individually. (This is also something that TTS software should support for their users' sake, but I digress.)
To be clear, it is reasonable to be practical and cater to the software as it exists, but that doesn't mean that we shouldn't ask for better software.
[1]: this is thanks to the crate unic-udc containing this information: https://github.com/open-i18n/rust-unic
I use such characters every day because my native language has them, so maybe that's why my tooling is naturally chosen/adapted to support it.
The simplest example is string interpolation, I've just submitted an issue: https://github.com/pfusik/cito/issues/26
However, I don't think a project may simultaneously have dangling references and claim to follow "Principle of least astonishment". Dangling references are unsurprising if you've been using C++ a lot, but I expect them to be very surprising to people coming from basically any garbage-collected language: Java/C#/Kotlin/JavaScript.
On the other hand, if you only know "safe" languages and target "safe" languages, there will be no dangling references.
By POLA, Ć doesn't reinvent keywords (compare to Rust) and the translated code is meant to look "obvious" compared to what you wrote in Ć. It's not "knowledge of C# is sufficient to write code correct in eight other languages" (that would be quite astonishing actually).
I think they’re challenging to popularize, though, because people already like their favorite language and love all its unique features, while these projects need to be a lowest common denominator by definition.
@pfusik please maybe put that example also in `ci.md`.
Why? Everyone else I've seen is not recommending the BOM and trying to get rid of them. They cause issues in various systems and you could just infer UTF-8 from the file extension.
This is just a loose recommendation, cito does accept files without the BOM.
It seems to me that if you say Ć program text is UTF-8 then that is an explicit encoding, and that if you feel it isn't explicit enough, an actual way to write out the encoding unambiguously is needed instead, which a BOM doesn't provide.
I am a little concerned by the wording "cito does accept files with the BOM". The BOM was chosen because it's a zero-width non-breaking character and so doesn't really mean anything, if cito thinks it means something, that's likely to be a problem elsewhere. For example if I concatenate two related Ć files, that ought to be fine, but I wonder if there's a BOM in the second file its presence in the middle of the concatenated file causes trouble.
It used to be a somewhat popular option for Haskell as well.
Have a look here:
https://github.com/pfusik/cito/issues/21
I generally check-in just the Ć source and not the translations, but if you want a quick look at the generated C code, here's some: https://sourceforge.net/p/asap/code/ci/master/tree/asap.c
> if there's a error/bug in the generated code am I actually going to debug and parse it effectively? Generated code may be "readable" but it is it understandable?
I had no problems with that so far. Nim adds a lot of boilerplate code in C output. cito sometimes adds a few lines here and there, but mostly it looks like the code you would write directly.
> Memory management is native to the target language. A garbage collector will be used if available in the target language. Otherwise (in C and C++), objects and arrays are allocated on the stack for maximum performance or on the heap for extra flexibility. Heap allocations use C++ smart pointers.
I have a set of runtime environments that only support lua but this Ć language looks interesting.
Obvious aside: I do agree with the many others that the funky line accent above the C doesn't help with typing during either discussions or searching, I'd expand the name purely for clarity. This isn’t actually bikeshedding - a name is used many many times in many contexts. It matters.
I recall taking a course in university about model driven programming - the idea of creating an abstract representation of logic, interfaces and other system components and then generating either full implementations or stubs in multiple languages was an interesting one, even if implementations were really hard to get right.
In practice, i've mostly only seen one language specific model driven design tools, like JHipster (https://www.jhipster.tech/) or the likes of JPA be reasonably successful, since there's a lot of problems with supporting abstractions across different languages and runtimes, but what has been the experience of others in that regard?
But in my experience, you oftentimes have to work with libraries and ecosystem components which are platform dependent. In those cases, the cross platform code would only work when these ecosystem components are also written in that particular language.
Now, it can work with Xamarin Forms, React Native or similar technologies with very specific use cases, but it feels like any other, more generic attempts at getting something like that working are doomed to fail (e.g. when you don't have a cross platform solution in place, for example, different ways to access DB on different platforms, or different ways to make web requests etc., or even interact with device capabilities or handle permissions).
Even if there is a little duplicate work, it's still better to have the whole app implemented in the native language and frameworks, rather than having some odd bits which don't work with the same tooling for example and are harder to debug.
Swift and kotlin are so similar i wonder if there isn't a possibility to generate at least the type definitions from one to the other. That would let me ensure that the code design at least stays in sync (at least for the model layer)
I can imagine there might be some serious tradeoffs in terms of making something which has to be interoperable with all these different target languages.
There are more subtle ways. For example, overflow of signed 32-bit int is:
* Well-defined in Java. I assume in C# as well.
* Well defined differently in Python: it starts doing arbitrary precision arithmetics.
* Well defined differently in JavaScript, because there are no integers, only IEEE 754-like 64-bit floating point numbers.
* Completely undefined behavior in C and C++. See the "Signed overflow" section at https://en.cppreference.com/w/cpp/language/ub and https://godbolt.org/z/y4vIi1 specifically, if you assume reasoning about undefined behavior is allowed in general case.
Using int32_t and uint32_t, try adding -2147483648 and -1. The arithmetic result should be -2147483649, but a program converting back and forth between signed and unsigned will produce 2147483647 -- wrong answer and wrong sign.
Of course you will never be able to obtain an unrepresentable value (such as -2147483649) through a workaround, and if you cast back to signed you get the wrapped result (which, indeed, may have an unexpected sign). But the point of transpiling to operations on unsigned numbers is to avoid UB, not to escape basic computational bounds.
https://www.stroustrup.com/whitespace98.pdf
>Generalizing Overloading for C++2000
>Bjarne Stroustrup, AT&T Labs, Florham Park, NJ, USA
>Abstract: This paper outlines the proposal for generalizing the overloading rules for Standard C++ that is expected to become part of the next revision of the standard. The focus is on general ideas rather than technical details (which can be found in AT&T Labs Technical Report no. 42, April 1,1998).
>Introduction: With the acceptance of the ISO C++ standard, the time has come to consider new directions for the C++ language and to revise the facilities already provided to make them more complete and consistent. A good example of a current facility that can be generalized into something much more powerful and useful is overloading. The aim of overloading is to accurately reflect the notations used in application areas. For example, overloading of + and * allows us to use the conventional notation for arithmetic operations for a variety of data types such as integers, floating point numbers (for built-in types), complex numbers, and infinite precision numbers (user-defined types). This existing C++ facility can be generalized to handle user-defined operators and overloaded whitespace.
>The facilities for defining new operators, such as :::, <>, pow , and abs are described in a companion paper [B. Stroustrup: "User-defined operators for fun and profit," Overload April, 1998]. Basically, this mechanism builds on experience from Algol68 and ML to allow the programmer to assign useful - and often conventional - meaning to expressions such as
double d = z pow 2 + abs y;
>and if (z <> ns:::2) // …
>This facility is conceptually simple, type safe, conventional, and very simple to implement.At least I'm not the only one who fell for it:
https://groups.google.com/a/isocpp.org/g/std-proposals/c/uTO...
In Kotlin you can write a library usable from C/C++/ObjC/Swift/Java/JavaScript.
But how much does it limit its own future just by having a name that isn’t distinguishable in speech from a much larger brand?
Though this makes me rather curious about products that succeed despite their branding. What happens if we make an unpronounceable language and it takes off? Would it just be given a de facto name by the community?
Leslie Lamport
Btw the GitHub repo is named cito .
The bigger problem is probably that English speakers can't pronounce it :) On the other hand Chinese should have no problem.
I mean, I probably never pronounced Python or Ruby in the same way a native English speaker does but I can write them and Google them. I have to copy and paste Ć or look for it in my keyboard's hidden characters (easier on my phone.)
Yes. Nobody calls ECMAScript by it’s official name. We just call it JavaScript.
If this becomes ever possible, it is going to be revolutionary.
Of course, it's possible to introduce a type like 'any' which can refer to any object, but then you essentially get untyped Ć which results in untyped Go. Which is not really Go anymore, it's more of a Python interpreter.
Go source code / binary in this case are of less importance for code readability, because they are meant for production deployments. Something that happened in GWT: write in Java, compile into JavaScript.
The written representation of this sound might vary - Polish Czeski and Czech Česky but the sound mostly remains same.
Discouraging bigots from joining a programming language community is a good thing for that community.
Super interesting. Could something like this be created for declarative UI paradigms like SwiftUI and WPF?
And we all use JavaScript NOT implemented in JavaScript. ;)
I did not look closely but writing parser without any tooling is asking for troubles, from parsers that enter infinite loops to not handling parse errors properly. ANTLR is way better for writing parsers for simple languages.
Says who?