How to program computers (kOS)
youtube.com
youtube.com
And why doesn't kx have binaries available for BSD? They have Solaris. Burning questions I have always had. Maybe just based on demand from their licensses?
k and kdb are some of the very few things that keep me interested in programming and computers.
clone, execve(!), epoll_create, epoll_ctl, epoll_wait, dup2, stat, rename, unlink, getcwd, chdir, fstat, getdents, open, close, read, write, ftruncate, mmap2, munmap, ioctl, gettimeofday, socket, bind, connect, listen, acccept, socketpair, setsockopt
execve just runs another k interpreter (regardless of what you tell it to do) though, so it can just be used to run more k programs.
But I also recall reading the line "no linux", which really piques my interest. I thought that just meant no Linux userland.
Now I like this project even better!
Sounds like you have access to the kOS source. Lucky you. Enjoy.
https://github.com/kevinlawler/kona doesn't run edit.k successfully.
So for all practical purposes kx.com and kparc.com are distinct languages/dialects.
Edit: I should clarify that I'm trying to make sense of the specific code in the OP, in particular to play with the views idea (the :: operator). I should have probably put this thread under https://news.ycombinator.com/item?id=8743325
Views are available in K4/Q/KDB+, so the version on kx.com does work with views. If you want to play with the views idea, try to solve a problem and see where a view helps you.
Here's one I did earlier this week; I've got this tool that takes a whole bunch of logfiles of userids and counts; they look like this:
5033703751425413371 1
5409789758109122623 4
10102846067816284236 4
I call these columns `s(sessionid) and `c(count); I can load them with this: G:{g::select from(+`s`c!("*I";" ")0:"S"$":",x) where (c>0)}
Now I say something like: G"trk/tnl5/20141207/TRK_ITM"
to get one of these structures into the g variable.I actually have a bunch of filenames:
t:{R:{(x,"/"),/:$:!"S"$":",x};,/R',/R'R"trk"}[]
So I can load the first one with: G t 0This is convenient, since these files are big (1Gb each or so) I'll want to work with a subset:
g:10000#g
This isn't part of my program, it's just something I do while developing. If I run "G" again it'll get the entire file again.Now, I want to match these against another list of users. I'll do a similar trick:
D:{d::0:"S"$":",x}
and now I can do: D"users.txt" to put this into the d variable.My program needs to find the rows of g that d match.
Some of the sessionids in the logfiles are corrupt. They look like this instead:
7241110807448127691320303160517889791474867567241110807448127691320312936057606400732600664532969577042775 4
62349552899220301013203170219134834182361546234955289922030101320303170219134834182361546234955289922030101 4
These are wrong. I originally tried to figure out them by finding the longest string that was also at the beginning of the sessionid, but after talking it out with Oleg and Pierre I decided to try to simply match them against the lengths of d: ?:#:'d
said I only have two lengths (19 and 20), so this is a much smaller search than what I was trying before!To do this, I used views:
K::?:#:'d
k::"S"$"S",/:$:'K
f::k!+(K#\:/:g.s)
Now f is an index on g; instead of a `s column it has an `Sn column where n is length of the key; i.e. I have a `S19 and an `S20 column with the first 19 characters of `s and the first 20 characters of `s accordingly. s::g.s[&|/f[k]in\:d]
c::g.c[&|/f[k]in\:d]
Now I can look at the `s and `c vars and do what I want to do; get my unique userids and my counts.More importantly, I can apply this to all my files (I have a lot of them):
n:0;m:0;{G x;n+:+/c;m+:+/g.c}'tIt's almost as if bloat is a meaningless snarl term with a mostly illusory definition beyond "bad".
Redundant code has performance and resource ramifications: Programs that are bigger load slower (more bytes to pull from disk) and run slower (more trips in and out of cache memory).
Redundant code also has other ramifications: Two modules in the same source tree may attempt similar (or identical) algorithms and miss an opportunity to consolidate. One of them might hide a bug, where the other code path was executed more frequently and the bug fixed.
One product had minimal redundancy. It used a plug-in architecture for the entire application - not just for extensions.
For example, why have Bentley when we already have Ford? Multiple car manufacturers is surplus, waste, and bloat.
Car manufacturers do not exist inside the source tree.
They aren't even obviously unnecessary; After all, people have different tastes about whether they like a Bentley or they like a Ford.
The computer (however) does not have taste, and it doesn't care whether you like your strcmp or my strcmp.
Is it something we can develop in people? Teach? Practice?
It may be rare that people will do it and invent it on their own, but I know there was a point in my life that I did not want to write code this way, so I know at least one person can change.
The speed at which code runs? The size that the binary is? The size of the source file? How long it takes to build software that runs correctly?
These are things we can measure, and while I suspect once we optimise programming languages for these things we will end up with something that looks like K, I do not think that it will be K because I have noticed that K is very bad at some things that I like to do.
Yes. It might be.
Only one person was hostile. They wanted to point out that there was no point in writing it all on one line "just to be unreadable".
I told them I thought the readability was improved by having all the words together, but maybe it made it difficult to learn. I noted Arthur had recently annotated edit.k[1] and asked them if that layout made it easier for beginner programmers to learn to read.
I don't think they liked being referred to as beginner programmers because he mentioned that he lectures at Cambridge, so of course they know how to program.
But most people were very encouraging: A lot of people wanted to share something they didn't like about different programming languages, and wanted some ideas about how they can learn more.
I got asked to read some more snippets of code [2] so that people could see this and get a better feel for it.
I got a couple of questions about syntax highlighting and development tools, and if there might be some middle-of-the-road solutions that were terse but not-as-terse.
Also: Beers.
The issue, for me, is not readability but context. Even a word or two (e.g. cuts, begins, dims) helps greatly to establish the context of the code. Like having a map before hiking.
When I write k/q, I use the same layout as Arthur with the additions of comments for each definition on the same line beyond column 80 or 132. Works well for wide monitors.
E.g.
c::a$"\n";b::0,1+c;d::(#c),|/-':b / (c)uts, (b)egins, (d)ims
I've also experimented with folding editors to keep comments above their respective definition while only taking up one line when folded. Some editors (e.g. - jedit) can display the folded text as a tool tip without expanding the fold. / Docs
(c)uts
(b)egins
(d)ims
c::a$"\n";b::0,1+c;d::(#c),|/-':b
I hope this is useful or helpful.Can you talk more about context? How exactly does it help you?
I don't think:
c::a$"\n";b::0,1+c;d::(#c),|/-':b / (c)uts, (b)egins, (d)ims
is more clear than: c::a$"\n";b::0,1+c;d::(#c),|/-':b
Or for that matter: ⚁::⚀$"\n";⚂::0,1+⚁;⚃::(#⚁),|/-':⚂
because they are just symbols to me, and one symbol is as good as another (except for the amount of space they take up on the screen and the ease in recognising it; I don't have many variables named lowercase L).Context is semantics/meaning. I.e. - hints, documentation, use case, approach, etc.
It's not about the code but the description of the code. The above code is the same, however, the comments loads the basic intentions into my mental cache. I can then read/use the code without having to analyze every definition completely.
This is especially helpful for k operators that can be overloaded by type or meaning. E.g. x!y can be create dict, table -> keyed table, keyed table -> table, integer -> enumerated value, ...
A hint (outside of the code) helps me out quite a bit.
Btw, the code above can be written as: d::(#c),|/-':b::0,1+c::a$"\n"
Which seems more clear, given the context.
What does `(c)uts` tell you?
I'm not actually familiar with this term.
> This is especially helpful for k operators that can be overloaded by type or meaning
We know the types of the data when we're tracing the code path.
> Btw, the code above can be written as: d::(#c),|/-':b::0,1+c::a$"\n"
You should let Arthur know.
It's hint/reminder of the purpose of the definition.
N.B. - This feature (i.e. str$chr) is undocumented in k.txt. I would normally use &a="\n". However, as soon as I saw Arthur's comment, it made sense (and uses less memory).
> We know the types of the data when we're tracing the code path.
My goal is to avoid having to trace the code path and to understand the definition in isolation (or as isolated as possible).
E.g. - you were able to add HOME/END functionality for your Mac without understanding the entire code path.
> d::(#c),|/-':b::0,1+c::a$"\n"
> You should let Arthur know.
Will do. He definitely appreciates concision.
Btw, I'm using k5 for daily use. Are you? If so, we could share notes.
Interesting.
I'd like to talk more about this part.
> If so, we could share notes.
Sure. You want to email me your details?
I felt as though you were about to make a very critical point and I would very much like to know what it was.
I had a go here:
https://news.ycombinator.com/item?id=8476294
I'm still working on it though...
Sorry about the screen visibility on this, as it was being live coded it was a bit hard to film.