Why Tcl?
gist.github.com
gist.github.com
Ruby https://github.com/ytti/oxidized/blob/master/lib/oxidized/mo... or TextFSM https://github.com/google/textfsm
are much better options if you need to do a /lot/ of parsing
https://wiki.tcl-lang.org/page/everything+is+a+string
Yes, there is some detail about internal dual representation.
Maybe, but it's clearly been updated after Tcl 8. And while the page itself isn't calling it something that hampers tcl, it does cover that it's somewhat unusual, that is has implications, and so on.
It works well as long as your goal isn't to totally eliminate boxing of values.
Yes, it is hard in many languages, having experienced it first hand. Just wanted to point out my experience with Elm (I know) 200k loc, I can refactor the codebase fearlessly, knowing that once it compiles it will "mostly" just work.
I fondly remember porting Elm 0.18 to 0.19, which was a big change in the language itself and the libs, me and my co-worker working across time zones almost 24/7 for over a week to make the damn thing compile, it was like struggling in the darkness, unable to see "anything" in the browser for like 7 days and when it finally compiled.. it mostly just worked!
It's sad that to see Elm getting so much negative publicity, but I do enjoy working every day on it.
Elm is a great language and I've written extensively[1] about how nice it is to use in production. But if you want to get involved in the community, it feels like there's not much to get involved in. The fact that even on the Elm Discourse[2], posts auto-lock after 10 days means that the forum is now pretty dead compared to what it used to be in 2018-2019.
Interpreted languages like Python and JavaScript do suffer when code bases reach a certain size.
The codebase is too big to rewrite but we're slowly untangling and hacking bits off to write in golang, but most of the business logic is in that DSL.
E.G, in the Counter snippet, why make some complicated code when Python is designed to make it simple?
import sys, collections
for p in sys.argv[1:]:
try:
with open(p) as f:
for k, v in collections.Counter(f).items():
if v > 1:
print(v, "\t", k, end="")
except Exception as e:
print("dup error: ", e)
I don't buy the argument of simplicity here. Also, if you are going to just do "get", you don't need requests, the stdlib is fine.And I don't get the speed argument either, given his script uses most_common(), that basically sorts the entire dict, which none of the other codes do. Yes Python is likely slower, but let's not push it.
Also the Go code is more complicated, but you can cross compile it and ship it as is. And it's fast.
I get it, TCL is cute, but making up arguments is not selling it.
Personally, TCL fits a niche like QBASIC did. QBASIC was the shortest path from brain -> code drawing on a monitor. Literally 1 line to draw a shape, no initialization and not even an entry point function.
TCL similarly lets you just get on with what you’re doing when you need glue code, GUI’s, and DSL’s. It is easily (even trivially) able to do things that are just a pain in other languages, especially all at once:
- interop with native code (no limits on who calls who or in what order)
- define GUI’s that behave well without it being a huge pain to make anything non-trivial (lookin at you, UE5)
- create novel control flow that feels built-in
- implement an interactive GUI debugger with breakpoints & variable watch in ~200 LOC (saw this once, it’s amazing what you get “for free”)
- save or load a running program’s entire state, including code defined at runtime
- detour any function, allowing you to optimize or patch on the fly
- run from tiny, standalone executables so your users don’t have to install a ton of crap just to run your widget
There’s more but you get the idea. It’s been around for decades for a reason :)
(each ((file *args*))
(catch
(dohash (k v [group-reduce (hash) identity (op succ @1)
(file-get-lines file)
0])
(put-line `@v\t@k`))
(error (e)
(put-line e *stderr*))))>The requests here is a third-party package that required installation and a virtual environment to use. I don't know how to catch more granular exceptions in this example.
so I guess they even didn't RTFM:
import sys
print(*sys.argv[1:])EXAMPLE 1: Echo command-line input
console.log(process.argv);
EXAMPLE 2: Count duplicates in a file import { readFile } from "fs/promises";
readFile("0.txt", "utf-8").then((text) => {
const frequency: { [line: string]: number } = {};
text
.split("\n")
.forEach((line) => (frequency[line] = frequency[line] + 1 || 1));
Object.entries(frequency)
.filter(([key, value]) => value > 1)
.forEach(([key, value]) => console.log(value, key));
});
EXAMPLE 3: GET request fetch(process.argv[2]).then(console.log).catch(console.error);
EXAMPLE 4: GET parallel requests process.argv.slice(2).map((url) => {
const start = Date.now();
fetch(url).then((res) =>
console.log(
`${(Date.now() - start) / 1000}s`,
res.headers.get("Content-Length"),
url
)
);
});
Using the author's criteria, relative to Tcl, these implementations (to me) are:1. Significantly shorter, just as expressive, and familiar to millions of (JS/TS) devs
2. Fast, owing to Node.js/V8's insane optimization level
3. Strongly typed
But you know what? I don't care much about terseness or speed when writing scripts. Honestly, the real reason I use TS for scripts is simple: comfort. I'm comfortable and productive with it. Everything else is just post-hoc rationalization. I'd wager it's the same story for many of these "Why <language/tech>?" posts.
It's perfectly fine to use something simply because you like it -- you're in good company. :)
For example, 'console.log' will pretty-print the data structure, rather than a single space separated line, the duplicate counter only reads from one file (with a hard-coded name) and will not correctly handle errors and move to the next file, and your parallel get doesn't seem to print one duration for all requests to complete.
js will only run on users machines or inside a vm without even kvm because we've seen plent ecosystem hijacks and exploits in the wild already. and you can only fool me MAX_INT.
Where security is concerned, I can't say that Node.js has been more problematic than any other language I've used, but that's anecdotal and YMMV.
But you need to convert the TS to JS, and the TS compiler takes like 3 seconds to start on my laptop (orders of magnitude slower than any other compiler I know, although you can speed it up a bit if you disable type checking, but that, well, removes type checking).
It sucks that your compiler's slow! Check out `swc` if you prefer speedy compilation.
Other languages can be compiled (perhaps "transpiled") to JS, some of which have static type mechanisms. One of those other languages, TypeScript, has static types. There are tools to compile TypeScript into JS: `swc`, `tsc`, etc. There are also tools to statically type check TypeScript (`deno`, `tsc`, etc).
Personally, I use `swc` for speedy TS compilation. I use my IDE's `tsc` language server for the separate job of highlighting type issues. But everyone has their preferred workflow!
Or use deno? Or bun?
.split("\n")
It looks like this has the same issue as the Tcl code, where the empty string after a terminal "" is treated as a line. I think that is an error in the Tcl.You'll also need a .catch(err => console.log("dup:", err));
And the console.log(value, key) needs a tab separator.
And a loop over the argv input files.
All minor tweaks which don't detract from your conclusion.
We need to be tolerant to files that are not terminated by a final newline.
This regex seems to do the trick (under the TXR Lisp tok function):
1> (tok #/[^\n]*./ "")
nil
2> (tok #/[^\n]*./ "no newline at end of file")
("no newline at end of file")
3> (tok #/[^\n]*./ "\nno-newline")
("\n" "no-newline")
4> (tok #/[^\n]*./ "line\nno-newline")
("line\n" "no-newline")
5> (tok #/[^\n]*./ "line\nline\n")
("line\n" "line\n")
6> (tok #/[^\n]*./ "\n")
("\n")
7> (tok #/[^\n]*./ "\n\n")
("\n" "\n")
8> (tok #/[^\n]*./ "\n\n\n")
("\n" "\n" "\n")
A line is a (maximally long) sequence of zero or more non-newline characters, followed by a character.LOL; I don't think I've ever used regex-driven tokenizing to recognize Unix-style lines in a text stream.
Is there a take-home message I'm supposed to get from your comment?
My comment was meant to be a light-hearted observation that you're overthinking the problem, since your language of choice likely has this built-in.
Or a boast that Python has a built-in feature for what requires a custom TXL and/or TypeScript solution. ;)
I have split and tokenize functions which have the same interface.
Split places the focus on identifying separators: the "negative space" between what we want to keep. When we use that for lines, it has problems with edge cases, like not giving us an empty list of pieces when the string being split is empty.
That is, splitting with regex look-behind/look-ahead, like:
(?<=\n)(?=.)
with a suitable definition of dot, also matches "negative space", and produces the same results as Python's str.splitlines(True), except for the empty string. >>> import re
>>> p = re.compile(r"(?<=\n)(?=.)", re.DOTALL)
>>> def check(s):
... lines = p.split(s)
... assert lines == s.splitlines(True), lines
... return lines
...
>>> check("no newline at end of file")
['no newline at end of file']
>>> check("\nno-newline")
['\n', 'no-newline']
>>> check("line\nno-newline")
['line\n', 'no-newline']
>>> check("line\nline\n")
['line\n', 'line\n']
>>> check("\n")
['\n']
>>> check("\n\n")
['\n', '\n']
>>> check("\n\n\n")
['\n', '\n', '\n']
>>> check("")
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
File "<stdin>", line 3, in check
AssertionError: ['']One of the most mind-blowing things to me as a (perpetual) Tcl noob is that when you look at
proc my_function {x} {
puts $x
}
You are lead to believe it is an ordinary C-inspired language with an odd syntax. It took me an embarrassingly long amount of time to realize that proc my_other_function "x" "puts \$x"
Is equivalent to the former. Other little oddities like `upvar` are beautiful once you understand them. It has some warts but Tcl has never had the luxury of under the scrutiny that languages like Python or Ruby have had. This is one of the reasons I would really like to see a "modern" version of Tcl more suitable for programming in the large---something that happened to Python 3 almost as an accident, imho.Would lua fit that definition? Between game engines, redis, openresty (nginx w/ lua) and others I feel like it fits a similar need (adding code to existing applications).
It was a fun OO system to write. It's much more dynamic than it appears to be at first glance; it's first glance looks much like many other languages object systems with a few slight oddities (no new operator, but rather classes have a new method).
#!/usr/bin/env tclsh
puts $argv
I kinda wanted to stop reading after this example. It feels dishonest to start with an example for which there is a builtin. Why didn't you show instead how to output arguments joined by a comma or some other separator instead of space.
If I were to update the Go or Python examples to use another separator, it'd be just a 1-2 character change, whereas for Tcl I have no idea how that would look.The Tcl variant to output "comma" as separator (note -- this is not "CSV").
puts [join $argv ,]
If you wanted "comma space" you'd do: puts [join $argv ", "]
If you actually wanted "CSV" then it would be (assuming tcllib is installed): package require csv
puts [csv::join $argv] puts [join $argv]
to avoid that little issue. Doing "puts $argv" triggers Tcl's output of "lists" in a special format that allows the list to be parsed from the text again at a later time.I learned Tcl / Tk back in 1995 because I wanted to write GUI programs and Hello World in Motif was about 2 pages of boiler plate code while in Tcl / Tk it was about 2 lines.
Then I started in the chip design world and Tcl took over so it was good that I already knew it. I was teaching the older engineers Tcl so it helped me gain their respect.
Their code sucked like crazy. Usually in such teams you are not evaluated by the quality of code.
Ever since then, any time I'm job hunting and I see a listing for a SW programmer that mentions Tcl, it's always been for EDA tools and I run in the other direction.
I worked in the semiconductor industry, specifically writing internal tools. I had to look at _many_ tools/scripts written in Tcl/SKILL/perl/awk/sed with some code being 30 years old and still being used.
Honestly, it wasn't the worse thing. There was usually very little abstraction, the code pretty much always told a story of what the writer was trying to achieve. Naming things was usually pretty bad and you'd get into some pretty gnarly regex/string matching, but at least it usually wasn't some over-complicated highly flexible but also restricting in-house framework.
In contrast, Python is easy to read, write and debug.
I quite like the small footprint of Lua as an embedded script language, which is easy to read, write and runs relatively fast.
I'd probably use Lua or a Scheme as an embedded scripting language if I had the need (generally, I try to not let my programs get so big that the require embedded scripting, but instead attempt to create orthogonal CLI commands that do one thing well and that can be scripted from the outside via POSIX sh).
Lua's match system is cool for what it is, and how small its implementation is. But it's really not powerful in the end and lua has almost no other string tools in the standard lib.
Tcl is built around string manipulation, it's the main and, for a long time, only data type. It has regex built in, a full set of string manipulation functions, and an interpolation primitive.
It's not about what's better. They are both full blown programming languages and you can implement any behavior you care to. In lua you have to care to implement string functions if you're going to do anything complex with strings. Tcl comes with them. That's all.
Not sure any of it would have carried thru till today even if the business hadn’t stumbled, but it very productive in the 1990s.
I see the response later that you’d rather use whatever the team already uses, but could you explain this very strong opinion?
I maintained a 200,000 loc TCL codebase running the front end for millions of lines of industrial C spread over hundreds of machines. It was glorious. After moving on, I’m still struggling to figure out why TCL is so unpopular outside that domain. Other comments seem to boil down to not understanding the language or how to apply it. So what’s your take?
I don’t know if all lisps suffer from that last one, and I don’t know about you, but there seems to be a clever solution hiding in every line. I think it would take discipline to have a codebase that remained cohesive and in a “single language”
- python
- nodejs
- bash
Of those, I personally would prefer bash as the “organizer of commands,” but that’s probably not surprising from someone who would learn tcl for fun.
Usually, the use case for these things is some combination of gluing programs together or gluing programs to the specifics of their OS. In that regard you might think, that picking a language specific for the problem at hand would be wise. But pragmatically, I believe you want devs who are mostly focused on feature work in their own programs to nevertheless be able to “bring their head above the water” and still have immediate intuition about how to swim. So even though nodejs might suck for this, it’s also extremely approachable for somebody who spends all their days in js.
I am really drawn to the idea of doing all of this in something like nix, but I tried it and didn’t have time to get past the learning curve. And that points to the other issue:
You’re probably going to need to glue rpm scripts, Dockerfiles, specific system tools which have different amounts of portability, etc. At the end of the day, if there’s a “lingua franca” for the organization, it usually makes sense to use it so that nobody needs to spend any time becoming accustomed to the idioms of some unfamiliar language while they are simultaneously dealing with the unfamiliar terrain of dealing with other systems that they have very incomplete knowledge of.
If there’s no lingua franca, I default to bash as a larger lingua franca
tcl is definitely much better, but equally uninspiring.
The idiomatic Python matching what the Tcl code does is:
#!/usr/bin/env python
import collections
import sys
for p in sys.argv[1:]:
try:
with open(p) as f:
c = collections.Counter(f)
for k, v in c.items():
if v > 1:
k = k.rstrip("\n")
print(f"{v}\t{k}")
except Exception as e:
print(f"dup: {e}")
continue
Except, even then there are three differences:1) the Tcl code accumulates the counts across all the files, the Python code resets each time. Though the Tcl code reports the counts after each file???
2) the Tcl code has a file descriptor leak (one of the big issues I have with Tcl is in dealing with resource management. I've used an upvar with a delete trace to emulate local scope, but that makes passing the result upstream tricky).
3) the Tcl code includes the empty string "" after the final newline in the count, which doesn't make sense.
To show the issue, I've replace the Tcl code so it outputs quotes around the line:
puts "$n\t'$line'"
and set the file descriptor limit to 10 before having it read a file containing only two lines, each containing "A": % printf "A\nA\n" > x
% limit descriptors 10
% tclsh dup.tclsh x x x x x x x x x x x
2 'A'
2 ''
4 'A'
3 ''
6 'A'
4 ''
8 'A'
5 ''
10 'A'
dup: couldn't open "x": too many open files
dup: couldn't open "x": too many open files
dup: couldn't open "x": too many open files
dup: couldn't open "x": too many open files
dup: couldn't open "x": too many open files
dup: couldn't open "x": too many open files
% python dup.py x x x x x x x x x x x
2 A
2 A
2 A
2 A
2 A
2 A
2 A
2 A
2 A
2 A
2 A# Echo command-line input
$args -join " "
(Yes, it's a single line.)# Count duplicates in a file
$lines = @{}
foreach ($path in $args) {
cat $path | % { $lines[$_] += 1 }
}
$lines.GetEnumerator() | % { if ($_.Value -gt 1) { "$($_.Value)`t$($_.Name)" } }
# GET request param($url)
irm $url
# Parallel GET requests $before = get-date
$args | % -parallel {
$res = iwr $_
$after = get-date
$time = ($after - $using:before).TotalSeconds
"${time}s`t$($res.RawContentLength)`t$_"
} -ThrottleLimit 12345
$after = get-date
$time = ($after - $before).TotalSeconds
"${time}s elapsed"
I'm showcasing a few different ways of doing things: using $args, using the param directive; using foreach, or % which is an alias of ForEach-Object; irm for Invoke-RestMethod, iwr for Invoke-WebRequest...For example, if I understand correctly, your "cat $path" won't actually work for all filenames, because "cat" is an alias [1] for "Get-Content -Path" which will expand wildcards in the filename [2][3], so you have to know to use "-LiteralPath" instead of you want it to work properly.
And then there's Microsoft's telemetry, which is enabled by default. [4]
[1] https://learn.microsoft.com/en-us/powershell/scripting/learn...
[2] https://stackoverflow.com/questions/33572502/unable-to-get-o...
[3] https://learn.microsoft.com/en-us/powershell/module/microsof...
[4] https://learn.microsoft.com/en-us/powershell/module/microsof...
Says everyone about every language.
Weird syntax? Weirder than TCL?
I frankly don't care about the telemetry.
Some of the lower-end SPARC machines could only have 16 or 32MB RAM and had perhaps 10 to 12 MIPS of computing power.
Expect, which became a very important extension of TCL, was released in 1990, when the SPARCstation 1+ with about 15 MIPS and max RAM of 128MB was introduced.
Thus any extension language like TCL had to be small and efficient; especially when you consider it was designed to be embedded inside another program.
1- It's easy to make an executable for Windows (as well as running it on Linux);
2- It's easy to write a GUI.
Worst way to make a scripted GUI: dtksh
I never liked TCL because the string quoting rules were almost as bad as Shell's. For small scripting I prefer Lua.
https://gist.github.com/antirez/6ca04dd191bdb82aad9fb241013e...
And of course Antirez has a soft-spot for TCL:
http://antirez.com/articoli/tclmisunderstood.html
Which inspired me to create a (trivial) TCL interpreter in golang. Not perfect, but almost as good as picol:
foreach line [split $file \n] {
incr count($line)
}
foreach {line n} [array get count] {
if {$n > 1} {
puts "$n\t$line"
}
}
......IMHO that's hard to grok. count shows up in the middle of a loop and it's not clear to me what it is. Could be a function, an integer, a map, or an array.
Same problem with this:
>foreach {line n} [array get count]
Not really obvious what any of this does. Guessing it's taking a map, converting it to key/value array, and iterating.
It is an array. 'incr' is a built in function that "increments" a variable by one.
Variable names followed by parenthesis () are Tcl arrays. So count(...) is "array syntax".
> Same problem with this:
> >foreach {line n} [array get count]
This is roughly equivalent to Python's "list comprehensions". "array get count" produces a list of values from the contents of the array named count (the list is "key1 value1 key2 value2 key3 value3 ..."
foreach is Tcl's "iterate over a list" function. In this case it is iterating two variables over the list at the same time taking two elements off the list each time. The first element is assigned to "line" the second element to "n".
Keep in mind that the default way C does for loops is the most un-intuitive way of doing a for loop. This Tcl example is less alien than the C for loop is for someone who has programmed in non-C like languages.
You learn it, and you get used to it.
for dish, price := range menu {
fmt.Println(dish, price)
}
for key, value in menu.items():
print(key, ' is ', value)
for (key, value) in hashmap {
println!("{} {}", key, value);
}
Really, it seems like that's a common syntax for languages with first-class dictionaries.Ousterhout saw that a number of design tools he was involved with all had ad hoc languages, so he developed Tcl as a single common solution to the 'scripting language problem'. This is partly why the language is designed to be as easily embeddable as it is.
Post this origin, Tcl developed a modicum of broader fame through Tk (one of the best/easiest early X11 toolkits) and expect (which is a tool for automating control of command line interactive tools).
The second one was through Maya, if you can call maya's language tcl.
It was fascinating to me back then, but nowadays I'm leaning more and more towards static typing.
Do you have any interesting stories from the porting process?
> Ousterhout's dichotomy claims that there are two general categories of programming languages:
>
> low-level, statically-typed systems languages
> high-level, dynamic scripting languages
I understand that this may have seemed somewhat more true at the time it was originally stated, but it's not really been true since the '90s, and it's certainly not true now. What an odd thing to lean into in this day and age.The only thing it really can't do is kernel code.
Even C++ is incorporating more and more high-level constructs.
Unless your team manages to never write YAML files.
My preferred one, which I have not found in other languages (besides LISP, obviously): treat data as code.
This way your configuration files are Tcl code using some proc's that you have defined, but they can also include arbitrary code. In order to read them, you simply fire up a slave (safe) interpreter and get the results.
If anyone knows how to do this easily in Python, I'm all ears!
Typically the reason for not using a full-blown programming language for config is to avoid Turing creep - the tendency to want to push complexity into config in order to make the application code simple. At some point the config gets to complex it needs its own config.
I run into this temptation using Lua a lot. Lua tables are perfect for config, but it would also sometimes be so convenient to just pop on a metatable and turn it into a class, maybe a proper API, when really all I need is a table that defines constants, maybe some conditionals, arrays, a map and that's it.
host thatmachine.example.org
port 9991
dbhost server.example.org
... which I can definitely parse myself. But it's also valid Tcl, if host/port/dbhost are defined proc's!So this allows people to be able to create config files like:
set mydomain example.org
host thatmachine.$mydomain
dbhost server.$mydomain
if {$mydomain == example.org} {
port 12345
} else {
port 9991
}
I don't expect people to create huge programs in their config files, but being able to give them a "full-blown programming language" for free is sometimes desirable.1. Of course "Hello world" takes more lines in Go than Python or Tcl. But that does not mean anything. The "overhead" here will become meaningful for larger scripts/projects, and Go or Python are far more capable of much more complex things than Tcl. This is simply a dumb example.
2. Other than certain circumstances, I never heard any one care about scripting performance on the ms level. Often not even at 1-2s (which rarely happens anyway). Putting a table out there and say "this is faster by <1ms" is meaningless to anyone other than the author.
> It was a good decision to invent tcltest rather than, say, "testtcl". Or, better yet, do not say the latter, at least in English.
1. a Package Manager (they sort of already have one called TEA, but its not active as far as I know)
2. an IDE or a good modes for the popular editors (VS Code, Emacs, Vim, Sublime Text)
For a UI language the absense of a good IDE or editor modes is very strange, you would expect that a Tcl/Tk IDE written in Tcl/Tk (one that also serve as UI framework or scaffolding for other apps) would be its killer app, but there isnt one
The IDE tools for Tcl have tended to be commercial and to not see much general adoption. Never understood why.
Otherwise please use something modern and sane like Deno, Python, Go or even Rust.
TCL has many large flaws that those languages don't share:
* No static typing so trivial errors are not caught early and code is difficult to understand and navigate.
* Everything is a string (TCL people try to claim otherwise but semantically it really is true). This leads to typing mistakes too.
* Quoting is a mess. It's simple in that the implementation is simple, but languages shouldn't optimise for that alone. Actually getting quoting right in TCL is ... well it's not as bad as Bash or even YAML but it's still something you have to think about, which isn't true for any of the other languages I listed.
If you need a language to embed in your program then while I'm no Lua fan, it is at least significantly better than TCL.
It's hard to be much more specific than that without talking detailed examples.
GPT-4 can easily generate trivial scripts in any language.
Any complex codebase can be maintained in a typical language.
Only a professional programmer is insulated from this obsolescence.
I am sympathetic to the fact that low-knowledge tech workers had their value erased over night. I don't have any solutions for this problem.
It can produce scripts in seconds that might have taken technical support people an entire day (stretched to a week or more of work) prior.
Of course, it can't maintain a codebase like a professional developer can. Thank God.
I would imagine if you just have a tiny bit harder problems, the Tcl code quickly becomes unwieldy and you end up digging through the archives to find some wizard's tome (or maybe there's a large community of Tcl users out there? IDK).
I can easily spin up a GUI in an hour tinkering with PyQt. Can you do the same in Tcl?
DSLs like this tend to have a very nice local maxima. But once you stray away from its original use case you start writing some really confusing code.
No, it would probably take 10 minutes with Tcl/Tk ;-)