Neovim: input encoding and you
aktau.be
aktau.be
In the first piece of code, which is apparently new to Neovim, there's a bug waiting to happen: casting size_t pointers to int pointers will fail on big-endian platforms where sizeof(size_t) != sizeof(int), and in general, the number of casts sprinkled everywhere is indicative of sloppy coding.
The whole area of code also seems to be poorly reimplementing iconv, which is actually supported in the code, but never enabled in Neovim for some reason, according to the footnote - why have two code paths to do the same thing?
> The function is short enough to paste below. It is pretty unique to Neovim, as it was wrought for its brand new event-oriented nature:
I believe the problem is the legacy vim API takes an (int ), so that code will also need to be corrected:
char_u *string_convert_ext __ARGS((vimconv_T *vcp, char_u *ptr, int *lenp, int *unconvlenp));You're correct. It's new code that's calling into old, unrefactored code. Hence the casts. convert_input() is new, string_convert_ext() is old. As is mentioned in the article, I'm working on a refactoring of string_convert_ext() that will obviate the casting (and incidentally also some allocating).
To be honest, my memory is failing me, so I would appreciate it if someone with a fresher understanding of the standards or BE compiler behavior would clarify.
Does a cast from (pointer to int32_t) to (pointer to int64_t) change the memory location referred to? My assumption would be not. If it does not, then there is a problem with this code on big-endian systems that does not exist on little-endian systems:
On a little-endian system, you are only reading/writing the low bits of the lengths. If anything won't fit in an int, that can be a problem. It seems plausible that other things (buffer sizes?) guarantee that can never occur, though it would still be better code not to rely on it.
On a big-endian system, you are only reading/writing the high bits of the lengths as if they were the low bits. That will almost always be wrong.
In fact, we have sections of the projects wiki devoted to this type issue: https://github.com/neovim/neovim/wiki (the specific section: https://github.com/neovim/neovim/wiki/Integer-types-refactor...).
It's not perfect, but you'll see we've put at least some thought into it and we welcome new contributions.
It's a problem of manpower and/or time, not necessarily of misunderstanding C and making rookie mistakes about types as the grandparent seems to imply.
There, had to get that off my chest. (I'm not seeing Neovim's code, old or new, is perfect, but it's probably better than what the tone of the comment implied).
int len = strlen( some_c_string );
Without getting a warning! I don't really need to be reminded that the random little strings I create in my code need to be less than 2 billion characters long!At first I thought you were referring to casting `size_t` to `int` (and the converse) in general, but now I notice you were referring to pointers to `int` to pointers to `size_t`. You are right, of course, and we should be more careful about that. Nevertheless, this specific case will soon dissapear.
To answer the last part of your question:
> The whole area of code also seems to be poorly reimplementing iconv, which is actually supported in the code, but never enabled in Neovim for some reason, according to the footnote - why have two code paths to do the same thing?
I alluded to it in the article itself, and I may only guess as to what Bram's goals where, but I reckon:
1) Speed, for an important conversion like latin-1 to UTF-8, Bram might've wanted some extra juice. This is not uncommon in the Vim codebase. Possibly, one can't (or couldn't when Vim was made) rely on iconv being well-engineered. Also note that Vim was made to run on some pretty bare bones systems. 2) Compatibility: as it stands, (Neo)vim can work just fine without iconv and still do the most useful transforms. This is actually good because iconv is not available everywhere or buggy (switches for disabling it have been requested in vanilla Vim).
I haven't grokked the other encoding sides of vim (file and output), but it's possible I'll discover some things that might preclude an iconv-only approach.
I'd use that for running ctags on save without blocking the editor, for example.
I do like my gui, so I'll have to wait a bit more
We have refactored some of the code base using modern coding standards, which makes it more hackable. Some of these systems have become more efficient by, for example, using pipes over temp files.
We accept upstream patches in order to stay on top of vim, so I hope that no one will ever not use neovim because of a feature lacking (aside from a GUI--see below).
neovim creates a socket for each instance that external programs can communicate with, allowing one to write a plugin or extension in any language, providing it has a msgpack implementation. These plugins may eventually control things like the GUIs, spell checkers, syntax analysers, allowing nvim to become more focused, like a micro-kernel.
As of right now, many people have switched without having any real problems.