What about K?
xpqz.github.io
xpqz.github.io
I almost wish this link was to a blog rather than to a book about K, for which I only have a perennial curiosity.
Here's to hoping they consider writing said blog. I notice they have one but it only has 3 posts, all of which are about past Advent of Code puzzles.
Guessing you meant to say "peripheral curiosity" here? Perennial would mean you have a long-lasting and/or continued interest/curiosity.
Uiua[0]'s stack model is much more annoying to work with, but I really appreciate its embrace of unicode glyphs. Every other derivative of APL throws those out at the first opportunity, but when you have a lot of glyphs, you stop being so tempted to make different arities cause the same glyph to mean wildly different things, when the arity is not actually written down explicitly and depends on whether the next thing to the left is a parameter or another function. Once you can See The Matrix, this is the chief thing that still does make K and friends objectively unreadable in a way they don't have to be.
[0]: https://uiua.org
K has "traditional names" for all the primitive operators which appear in reference cards and which are typically used when discussing code aloud with other K programmers. Q and Lil, which are both K descendants, outright replace some symbols with those named keywords. Named keywords can make the primitives superficially easier to remember, at the cost of making idiomatic patterns in the language less visually apparent.
To be frank, your quote is mind-bogglingly stupid. How easy do you think Java is to read to a native Greek speaker with no English language knowledge? Would the Java standard library be as easy for you to read if it were written in Greek? Would java still be an easy to read language if all the keywords and library were in Greek?
If your only definition of a good programming language is one written in your native language and a PL becomes bad if written in a different language, then your criteria is terrible/useless. And right now that's your criteria.
I don't know about the qualities of k itself, but I think the idea of having a common practice for experimental programming languages to be grouped under a single name like "E" with a number is quite attractive.
There are lots of students, hobbyists, researchers, professional devs and companies who are developing their own working programming language. There are a million of them, all with their own names. 99.9% of them are ignored, or criticized unfairly by others expecting fully fleshed out features.
I can imagine a GitHub repo where you can register a new language "En" (with n being a number) rather than it living in obscurity on a random website. Then others can jump in and experiment with the language and give it feedback, fork it, etc.
This isn't just for toy languages, but for big organizations like Google. Instead of naming a not-fully-baked C++ successor as "Carbon" and getting flak for it not being ready for real world code yet, they could simply call it "E321" and the status of the language would be self-explanatory.
Then if one of the E languages gains enough traction, it could "graduate" to its own named language.
I also like the cred that an "official" E language could get when a dev talks about it to others. Everyone would immediately know it was experimental and where to see the code.
That said, some form of array language more suited for stuff like that is a somewhat common question; maybe one day someone will figure it out.
Vanessa McHale is doing some interesting work on a typed compilable array language, Apple[0].
[0]: https://github.com/vmchale/apple/?tab=readme-ov-file#apple-a...
While we have CUDA being polyglot, it is still pretty much C and C++, or shader languages, hence why I keep thinking why not an array language that is also a kind of shader language DSL.
Thanks for the heads up nonetheless.
What problem is K trying to solve? What does a K program look like?
So I think k, q and kdb are fun to work with, but one of the major components of its success is that it allowed a community (in finance) to evolve that can earn 50-150% more than their peer groups who do the same work in Java or C++. 10 years ago a kx course cost $1500 per person per day.
>What does a K program look like?
You might want to check out https://news.ycombinator.com/item?id=40335921
beagle3 and geocar both have various comments you might want to search for.
With an Oracle-style DeWitt clause[1] prohibiting public benchmarks.
[1] https://mlochbaum.github.io/BQN/implementation/kclaims.html
[1]: https://shakti.com/ -> Compare -> h2o.k
You can link to the subsections: https://shakti.com/compare/h2o.k
As a sanity check I just cloned https://github.com/h2oai/db-benchmark, ran the data generation script and ran on a 64 core AMD EPYC (AWS c7a.16xlarge):
import polars as pl
lf = pl.scan_csv("G1_1e9_1e2_0_0.csv")
print(lf.select(pl.col.v1.sum()).collect())
The above script ran in 7.58 seconds.If I change the collect() to collect(new_streaming=True) to use the new streaming engine I've been working on, it runs in 6.90 seconds.
I can't realistically time the full "read CSV to memory" with this 50 GB file on this machine as we start swapping (this machine has 128GiB memory) and/or evicting data from disk cache (this machine has a slow EC2 SSD attached to it), so we do have a blow-up of memory usage (which could be as simple as loading small integers into an 8-byte Uint64 column). I think it's likely that on K's machine the "read full CSV to memory" approach also started swapping, giving the large runtime. However, in Polars you'd typically write your query using LazyFrames, which means we don't actually have to load the full CSV into memory.
EDIT: running on a m7a.16xlarge with twice the memory (256GiB) once the CSV file is in disk cache Polars can parse the full CSV file into an in-memory dataframe in 7.68 seconds.
K's claim that it parses the full 50GB CSV in 1.6 seconds if true is very impressive regardless.
(But even then, 1.6 s would be quite a feat. It makes me wonder if the K implementation is partially lazy, as you say typical Polars usage is.)
It might be worth speculating, or at least optimizing the serial chunker more. You could theoretically start a second serial chunker from the end working backwards but that would not be wise with our ordered streams, as the decoded data would have to be buffered for a long time.
Similarly on the new streaming engine, each thread is active ~half of the time, except the thread running the chunking task: https://share.firefox.dev/3WQV9og.
Note that in a lot of realistic workloads on the streaming engine compute can happen in between decodes, completely hiding the bottleneck. Also all of the above is with the file being completely in file cache, if fed from a slow SSD it's not a bottleneck whatsoever.
Or since newlines in strings should be rare, maybe it works to save the index of every newline and tag it with the parity of preceding quotes in the block. Then you get the true parity once each thread's finished its block and filter with that, which is faster than going back over the block unless there were tons of newlines.
But we only have a finite amount of time and tons and tons of work, so no one has gotten around to it yet. At least now we know that it might be worthwhile for >= ~32 core machines. PRs welcome :)
Also, if you encounter a double-quote character anywhere with a comma on one side and neither a newline, double-quote nor comma on the other, you immediately know 100% whether it starts or ends a string.
I get the idea that one either already knows one needs an array programming language, or doesn't grok why anyone would need one
But, this sort of language is more about writing and reading from the disk efficiently, right? I guess SIMD type optimizations would be less of a thing.
Here's a program in k. I'm not sure exactly what it does. I think it might be a json encoder/decoder:
Two reasons k folks like k: first, if you believe that programmer working memory, as in the number of chars or lines of code you personally can hold in your head is limited, then it might make sense to be as terse as possible -- this will significantly increase the range of things you can reason about.
Second, if such a language were to focus more on array and vector-level manipulation, then for certain sorts of math tasks, you might be pretty close to grad student nirvana -- programming looks like using a chalkboard to work out a strategy for some processing, and then straightforwardly translating this strategy without mucking around with all the 100s of lines of weird shit say python or java make you do to process something in bulk and in parallel.
On top of this, whitney is a mad genius, and his k interpreters tend to be SCREAMING fast, and, like a couple of hundred kilobytes compiled. Over time the language has built connections to large-scale data processing jobs (as in, you run a microsend-or-shorter-timeframe strategy based on realtime depth data from 500 different stocks, say), and it has benefitted from the path dependence you get there.
Anyway back to the top - it exists as both a rallying cry for and a great tool for a certain sort of engineer that wants to make millions of dollars and refer to him/herself as a "Spartan" of coders.
Note, from wikipedia: Q serves as the query language for kdb+, a disk based and in-memory, column-based database. Kdb+ is based on the language k, a terse variant of the language APL. Q is a thin wrapper around k, providing a more readable, English-like interface.
Coding Style The q gods have no need for explanatory error messages or comments since their q code is perfect and self-documenting. Even experienced mortals spend hours poring over cryptic q error messages such as the ones above. Moreover, many mortals eschew comments in misanthropic coding macho. Don’t.
A more enjoyable read than the parent post.https://code.kx.com/phrases/wikipage/
It's primarily used for trading research and surveillance, not live trading. And I've never heard of anyone running it without an OS.
kOS is in development though current status is unknown.
(https://gist.github.com/chrispsn/da00835bb122c42f429a084df83...)
heh
For those curious, what they're actually using is FPGAs and custom silicon.
How does it compare to R/tidyverse?
In my opinion, it's very cool, but Python's ecosystem (and R's) is just so much better with scientific libraries and charting and all that. Kdb+ (the database) and K the language are likely much faster than R for general analysis type stuff. R is also free and Kdb+ is not.
I'd be really curious to know if they really are baseless. It's very very difficult to imagine that K developers can really read a mess like this as easily as one might read Go or whatever.
https://github.com/KxSystems/kdb/blob/master/e/json.k
Has anyone tested this? Take a K program and ask a K developer to explain it? Or maybe introduce a deliberate bug and see how long they take to fix it compared to other languages. You could normalise the results based on how long it takes them to write some other code.
Free research project for any compsci researchers out there... (though good luck finding skilled K programmers).
水落石出。
> Has anyone tested this? Take a K program and ask a K developer to explain it?
I am not sure what you're asking. Do you want me to read it to you?
Here is me reading some other people's code:
https://news.ycombinator.com/item?id=8476633
https://news.ycombinator.com/item?id=22010223
Do you want me to read to you the JSON encoder (written twice) and the decoder in this way?
> Or maybe introduce a deliberate bug and see how long they take to fix it compared to other languages.
https://news.ycombinator.com/item?id=27209093#27223086
> You could normalise the results based on how long it takes them to write some other code.
It seems to me that it has a lot of the same properties as regex. Looks like gobbledygook at first glance, but after learning it I can write them, and read them with some effort (depending on the complexity). However nobody would describe regexes as "readable", and they're quite error-prone. I definitely wouldn't want to write a whole program in regex.
Regexes shine most when they're used interactively, e.g. in one-off greps, or editors. There readability doesn't matter at all, error-proneness doesn't really matter, and terseness is important. The problems start when people put those grep commands in scripts where the output isn't supervised by humans.
I wonder if the same is true for K - it started as a query language for one-off queries & investigations, and then people started saving those queries and making bigger programs?
I would, and do, and I urge you to be less judgemental about things you do not know anything about because you will never learn anything new with that attitude.
> I wonder if the same is true for K - it started as a query language for one-off queries & investigations,
Why do you wonder this? I don't think it's true, but so what? Did you not read what I wrote? Seriously: Why do you put so much effort trying to talk yourself out of learning how to do something that is obviously amazing to you?
I am telling you I can read this. I like this. I am not nobody, just someone you did not think existed. And I am telling you it is possible for you too.
The thing is, I am extremely familiar with regexes (I've even written a regex engine), so I know exactly how readable they are - even after knowing them really well. So the fact that you think they are still readable suggests to me that your judgement of K's readability is also suspect.
> Why do you wonder this?
It would be a reasonable explanation of why K exists.
Your experience writing a "regex engine" once upon a time led you to believe regular expressions are difficult to read.
My experience maintaining a few million lines of perl over a couple of decades has led me to believe that I can read regular expressions with no discomfort.
The Real™ thing is you can get better at anything with practice, even this, but listen I also think K is more useful than regular expressions and I would have used less perl had I learned K sooner.
> So the fact that you think they are still readable suggests to me that your judgement of K's readability is also suspect.
It should make you suspect whether or not you have any idea what an expert actually is. I mean, the inventor of regular expressions tinkered with them for decades, and new advancements are still happening sixty years later!
You don't know what you don't know, and there is very little you can do about that except pay attention to people who can do things you do not know how to do yet, and reserve your judgement about how they do it until you can do it better.
The only thing k has in common with regular expressions is your claim they are both difficult, a claim I disagree with.
> > Why do you wonder this?
> It would be a reasonable explanation of why K exists.
You misunderstand me, perhaps on purpose, but I hope you and others will think about this: Why do you care why it exists when I have shown you something so much more amazing than an opinionated history lesson?
I think k exists to make programs that make money. Forever. Because a little bit of money from a lot of programs over a long time is worth a lot, k is fast to write it. Because sometimes getting the answer faster makes more money, k runs fast too. Because people are trusting their money with it, k runs very predictably. Because sometimes your vendor just changes the input format on a Friday night, it's important that it is easy to read and make changes to k programs.
Arthur said it was the keys to the kingdom.
What do you have to gain from this stance, and why don't you believe people who tell you otherwise?
Either everyone who uses array languages does actually find them readable, or they're all persistently lying for... what reason? And forcing themselves to use something they don't find readable? Why would anyone do that! Especially considering a lot of array language users are hobbyists, who have chosen to use them, it's not like they're forced to.
Here's a gentle guide to APL by the same author (me):
https://xpqz.github.io/learnapl/
Dyalog APL is likely the best supported in terms of tooling, debugging etc. If you're looking for static typing, you're in the wrong place.
The one that will jump out at most programmers who are familiar with mainstream languages is that J, k, q and Nial use ASCII characters while APL, BQN and Uiua prefer glyphs. q and Nial additionally favor words rather than shortened abbreviations, and Uiua has plain words that auto-format to its glyphs to aid in typing. The other glyph-based languages rely on custom (software) keyboard layouts or input methods to let you type the symbols they need. You do not need a special keyboard to program in any of these languages. ASCII-or-not is not a decision that any of the array languages have made lightly or for purely aesthetic reasons, it has deep consequences for how the languages feel that won't really make sense until you get some hands-on experience. As a beginner you'll probably gravitate towards one of the sides without understanding those deeper implications, and that's totally okay, but please keep an open mind.
If access to a high-quality open-source implementation is important for you, your options narrow a bit. J, BQN, Uiua and Nial all have a primary implementation that's open source. k has implementations that are open-source but the official versions of k that most people use "in anger" are commercial products with a limited free trial, and afaik there's no mature open-source versions of kdb+/q, which are kind of k's killer app. There are many implementations of APL but Dyalog is the clear leader and it's a closed-source commercial product with a personal/non-commercial free version. I wish this was less of a factor because it's so hard to get people interested in languages when the best versions aren't available to them, but it has gotten better in recent years.
Regarding tooling, you should go in with minimal expectations. Some of the tooling is quite good (particularly J and Dyalog APL, in my opinion) but it's heavily biased towards the specific type of iterative, interactive development that nearly all array programmers favor. Debuggers are sometimes present but usually not a primary tool. None of the major array languages have static typing. There are some array-adjacent languages like Futhark and Dex that do, but they're very different than the "Iversonian" array languages you asked about, and are also active research projects.
(Edit: Also worth mentioning that package managers and build systems are not common in the array world.)
There are many other differences that matter immensely to the array community but you won't have context for as a beginner, so I'm not going to go too deep into them, but if you're curious, https://github.com/codereport/array-language-comparisons has some comparison tables and example code written in a variety of languages. code_report/Conor's Youtube channel at https://www.youtube.com/@code_report/ is also an excellent place to get exposure to various array languages and concepts.
All that said, in my opinion the easiest languages to recommend to get started are BQN and J, depending on whether you want glyphs or not. If you're comfortable using a closed-source tool with restrictive licensing, Dyalog APL is also an excellent choice. Any of the three will show you both the joys and pains of array programming if you put time into learning it, and give you enough context to make an informed decision about going deeper or finding another array language more to your taste.
The documentation is pretty decent compared to the other members of the Iverson gang and the libraries one can install with the desktop version makes it somewhat batteries included, at least it's easy to suck in a file and start rendering plots.
Maybe BQN can compete on these things nowadays, I'm not sure.
Thoroughout the article it's spelled k consistently except at the start of a sentence. This is weird. The language is K not k. Nobody spells the C language as c.
I hope not.
And if you have to handle arbitrary language user input, there's basically no operations you can/should actually do anyway. Uppercasing/lowercasing? Doesn't make sense on CJK languages. Reversing? Completely meaningless. Trimming to the first N chars for some visual display/summary/preview? Even grapheme clusters won't help avoiding a character with ten thousand combining components, and you'll have to do language-specific logic to not cut in the middle of a word for languages where the display of a prefix of a word may change depending on later letters! And forget about spaces meaning anything.
Basically the only string ops I can think of that make sense for non-ASCII generally would be splitting/joining on newlines and escaping for JSON/HTML or whatever, which'll work completely fine on a byte list anyway.
There's perhaps some middle-ground of doing things for a specific set of languages, but even for such you won't care about the storage format anyways, as what matters for you is just whether operations you use (presumably using some library; and even if you write a manual uppercase for French specifically or whatever, you'd notice if you implemented it wrongly) do the thing they should.
So a list of byte chars is just fine for anything one would actually do, providing optimal access to ASCII, and not actually making things worse for non-ASCII.
Places where ASCII-only is a known expectation and there are meaningful per-char operations are plenty; that's what using a list of bytes provides. Indeed you'd probably want to use another abstraction if you have non-ASCII. And for such you could use something to do the form of iteration or operation you want just fine, even if the input/output is a list of byte-chars representing plain UTF-8.
Even if ASCII is appropriate in some situation, this should be stated within the program. Requiring people to be explicit about the data they produce and consume is important and useful. A user might decide that UTF-16 best serves their need (or be working on the Windows platform) in which case code which works with strings as linear sequences will be able to operate on their strings without issue. Code which assumes a UTF-8 byte representation will require an the entire string to be allocated, converted, then reallocated and converted back. Huge overhead and potential incompatibility for no reason.
I assure you, 99% of people won't handle this correctly even if given a cluster-based interface (if they even bother using it). And this still doesn't handle the question of cutting words in the middle of some languages resulting in broken display of the non-cut part (or languages without space-based word boundaries to cut on). So the preferred thing is still to use a library.
I don't think anyone in k would use UTF-16 via a character list of 2 chars per code unit; an integer list would work much nicer for that (and most k interpreters should be capable of storing such with 16-bit ints; there's still some preference for using UTF-8 char lists, namely, such get pretty-printed as strings); and you'd have to convert on some I/O probably anyway. Never mind the world being basically all-in on UTF-8.
Even if you have a string type that's capable of being backed by either UTF-8 or UTF-16, you'll still need conversions between those at some points; you'd want the Windows API calls to have a "str.asNullTerminatedUTF16Bytes()" or whatnot (lest a UTF-8-encoded string makes its way here), which you can trivially have an equivalent of for a byte list. And I highly doubt that overhead of conversion would matter anywhere you need a UTF-16-only Windows API.
I doubt all of those fancy operations you'll be doing will have optimized impls for all formats internally either, so there's internal conversions too. If anything, I'd imagine that having a unified internal representation would end up better, forcing the user to push the conversions to the I/O boundaries and allowing focus on optimizing for a single type, instead of going back-and-forth internally or wasting time on multiple impls.
I'm really not seeing the issue here.
I'm sure there was a time it was best in class and even now maybe it's the best for a few niche use cases, but unless you're absolutely certain you need it, I would flee from it and save your sanity.
Yeah...Oracle licensing sounds scary and having to pay to fix their own memory leaks sounds frustrating.
Thanks for the experience.
As somebody who hacks on, around and in esoteric languages for fun; I must object.
Author didn't expect to end up on HN then :)
Similarly, the inability of a person to write machine code directly is a property of the person, not the hardware. Yet some of these people admit their limitations and use K.