Needed 1+1, built a functional programming language
hereticpleb.vercel.app
hereticpleb.vercel.app
> Any sufficiently complicated C or Fortran program contains an ad hoc, informally-specified, bug-ridden, slow implementation of half of Common Lisp.
More precisely, I had figured out how to parse an Excel formula (keeping track of braces and commas, I recall!). For the actual math symbols, all I had to do was rewrite them as Excel formulas!
3+6*2 -> 3+MUL(6,2) -> ADD(3,MUL(6,2))
In order of precedence. And then just evaluate the resulting formulas inside-out:
ADD(3,MUL(6,2)) -> ADD(3,12) -> 15
It somehow ended up being a thousand lines of code. Another guy doing the same assignment told me he did it in 50 with regular expressions.
In retrospect, what I did reminds me of hammering in a nail with a screwdriver.
I ended up revisiting the idea later (just the math parser, no ADD() syntax) and it was much nicer, like 30 lines of JS (no regex either!). (Well, you can do it in one line with eval, but yea :)
I did the same for my performant implementation of pure functional programming language BLC/BLC2, which in 400+ lines contains a graph reduction engine for combinatory logic, to which the lambda calculus programs are converted by Kiselyov's bracket abstraction algorithm.
Used to be a common refrain when we were whiteboarding new features. Usually as a check on somebody's overly ambitious thinking.
I counted 23 "just"s in the article. Bravo. Sometimes you have to _just_ plow ahead.
I only skimmed the linked article, but I do wonder whether the author ever realized that they needed to think about scope rules. I searched for the word “scope” but never found it. Closures seriously complicate language design and things get painful and counterintuitive unless you use lexical scope (or… you know… you like pain).
[1] https://en.wikipedia.org/wiki/Hash_table#Collision_resolutio...
> Closures seriously complicate language design and things get painful and counterintuitive unless you use lexical scope (or… you know… you like pain)
Well if you like both lexical scope and pain, there's always lisp!
You might end up founding one of the first e-commerce sites, sell it for a handsome payday, and then create the first startup accelerator and make insane amounts of money while transforming the industry.
Safer to stick to Blub.
Maybe in JavaScript... but there is a topic called datastructures, where arrays are arrays, and hashtables are hastables. If people say this without a flip of an eye I'm not surprised why people write stuff like
> I thought that hash tables were something that were basically impossible to make in C
It’s just not particularly common (or helpful) to view them that way.
I also like https://github.com/tidwall/hashmap.c for general use.
Why would you believe this in the first place?
No secondary data structures, super simple to implement, and it works great for small use cases.
Btw, this reminds me a bit of Cuckoo hashes. Never used them but seems like a nice idea.
The cool thing, however, is: if the size of the new hashtable is double the size of the old hashtable, your amortized insertion costs are still only O(1)!
(And you don't just have to take my word for it: take my original comment and paste it into the AI interface of your choice and have it create a concrete implementation. Ask it to add a remove operation, and an automatic doubling of the array size + rehashing when the table reaches a load factor of, say, 0.7 -- the resulting code should be very manageable, and then you can run your own tests and measure times!
This is maybe not the smartest way to do hashing, but its appeal lies in its simplicity and hence compactness of implementation. There are many cases where you don't even need a 'remove' operation, and where you never have to worry about growing the array because you know that you're only ever going to hash a certain number of elements at most.)
Which during my degree, the lab deliverables were 100% C code.
Here, one possible book:
Data Structures, Algorithms, and Software Principles in C (1994 edition)
(Almost by definition, anything doable in a higher level language is doable in a lower level language, but not necessarily vice versa. In fact, many higher level languages are themselves written in lower level languages, e.g. Python is written in C.)
But Python has got dicts built into it, which are nothing but hash tables, and they are almost certainly written in C.
Google for some videos by Raymond Hettinger about Python dictionaries.
Or look at the source code of the Python interpreter.
In CPython they are written in C (at the moment). Other Python implementations are written in other languages.
> (Almost by definition, anything doable in a higher level language is doable in a lower level language, but not necessarily vice versa. In fact, many higher level languages are themselves written in lower level languages, e.g. Python is written in C.)
It depends on what you mean by 'anything doable'. Eg Haskell compilers can in principle do lots of crazy optimisations that a C compiler would not be able to safely do, just because they don't have enough information. Even more so for Lean compilers, which can _know_ which of your loops are terminating, instead of making crude assumptions like C compilers.
Btw, higher level languages being implemented in lower level languages is mostly something for interpreters. Writing a C interpreter in Python is pretty much futile, if you care about speed. But writing a C compiler in Python is perfectly fine. And writing a Python compiler in Python is also fine. Many languages self-host (at least some of) their compilers.
Anyways and however, speaking in strict mathetical sense, the model of a tree actually breaks the classical mathematical model of operator precedence.
1 + 1 + 1 evaluates to:
(+)
/ \
(+) (1)
/ \
(1) (1)
The above will be correct, mathematically, but will break once you involve multiple mathematical operators in the statement. Because, mathematically, operators have precedence.The operators are essentially ordered based on their depth into the right of the question / statement but unless I am terribly wrong, this is not the case in mathematics and some operators have higher precedence regardless of their position in the statement.
Any comments on this?
1+2*3^(4+1)+2/3
We start the evaluation with 1+, the chain would be on (+) for now.
(+) <
/
(1)
The next operator is (*), and because this one is heavier it falls down, bringing 2 with it. (+)
/ \
(1) (*) <
/
(2)
Now checking the next operator (^), once again heavier, goes down the three. (+)
/ \
(1) (*)
/ \
(2) (^) <
/
(3)
The next operation is between parenthesis, that takes precedence and goes down, but the pointer comes up after the operation has taken place. (+)
/ \
(1) (*)
/ \
(2) (^) <
/ \
(3) (+)
/ \
(4) (1)
Next operation is (+), this one floats, and as it's the same weight at the root it doesn't matter its relationship with it, so we put it higher. (+) <
/ \
(+) 2
/ \
(1) (*)
/ \
(2) (^)
/ \
(3) (+)
/ \
(4) (1)
Finally we got the division, which is heavier again. (+)
/ \
(+) (/)
/ \ / \
(1) (*) (2) (3)
/ \
(2) (^)
/ \
(3) (+)
/ \
(4) (1)
Did this on the fly, so it might have edge cases, but it works well as a starting point.Turns you have a fairly, if not a very complicated tree for a simple problem. But, fair, enough, you can always get AI to write the code for this and probably, the code will be reused over and over.
But, is it just me or does someone else think that having to rebalance or re-order the tree might be a good breaking point for the camels back?
:-)
(+)
/ | \
(1) (*) (/)
/ | | \
(2) (^) (2) (3)
/ \
(3) (+)
/ \
(4) (1)
If you transform the mathematical operation to lisp terms you will have the tree explicitly written.The only way that I can think of to make it simpler is to go the array language route of going right to left and ignore precedence, or some similar way of working.
The code to create the tree should not complex, so yes, AI could be used, but most competent coders should be able to create a basic version and test it in an afternoon.
If you are parsing left to right it can be quite easy to balance things, you should not need rebalancing at all. When you are in an operation node and you have to add an operation of the same level of priority, you always get the existing operation and subtree, put it on the left of the new operation, and continue from there.
With this if you try to do - next to a - or +, or a / to a / or *, you'll be preserving the order of the operations, and the calculation will be correct. Try it with 2*3/4.
(*)
/ \
(2) (3)
(/)
/ \
(*) (4)
/ \
(2) (3)
If you add some more operations to the right (*2/5*12/7) you keep growing the tree. It will not be balanced, but it doesn't need to be.I generally tend to find the idea of 'competent' coders misleading. For the most part, I have been developing or writing code in a certain language for while, then it turns out that I am competent coder? Because, yeah, I've been using C or Python for a while and can easily solve a lot of "complex" problems in C but with all due respect, I am obviously not a competent coder. It's like a situation where you frequent a certain part of town that you're very familiar with it and the people living there but then at the same time you don't live there - lol.
Anyways, for this problem I would parse this statement but definitely not into a tree. I'd assign the integrals to objects. I would then parse or go through the statement again executing the operands.
Well...
And in a way you are solving it the same way, if I understood you correctly.
Assuming you mean that you'd have classes, and create objects with the operations, it's the same as a tree.
1+2*3^(4+1)+2/3 -> plus( plus(1, mul(2, power(3, plus(4,1)))), div (2/3)) would be the object hierarchy created.
OTOH if you are parsing it and doing the operations that can be done because all the operands are known:
1+2*3^(4+1)+2/3 -> 1+2*3^5+0.66 -> 1+2*243+0.66 -> 1+486+0.66 -> 487.66
You'd be, once again, doing the tree but instead of having it as an structure you'd be directly parsing the leafs than can be operated and act on them. This way would maybe be faster for simpler expressions (no need to construct the tree), but probably be more expensive than tree construction and resolution for more complex ones.
Trees being incompatible with precedence is like parentheses being incompatible with precedence.
IMO, trees are not incompatible with precedence but I just prefer a more barebones approach to this fairly simple problem.
(+)
/ \
(1) (*)
/ \
(2) (1)
The tree representation is unambiguous and once you get that there's no need to think about precedence.There's many ways to do this parsing, e.g. https://matklad.github.io/2020/04/13/simple-but-powerful-pra...
I had a look at the link. The BNF looked good:
Expr =
Expr '+' Expr
...
Then the author fixed the left-recursion and precedence (which also looks good,) but then complains about the fixed version - "the “shape” of expressions feels completely lost in this new formulation." : Expr =
Factor
| Expr '+' Factor
...
Then the author takes us through Pratt parsing and ends up at: fn expr_bp(lexer: &mut Lexer, min_bp: u8) -> S {
let mut lhs = match lexer.next() {
Token::Atom(it) => S::Atom(it),
t => panic!("bad token: {:?}", t),
};
loop {
let op = match lexer.peek() {
Token::Eof => break,
Token::Op(op) => op,
t => panic!("bad token: {:?}", t),
};
...
Yikes! I think he criticised the wrong code. I'll take the '{expression} is a {factor} or an {expression plus a factor}' formulation over the 'mut-loop-peek-panic-lexer-next' approach any day!Here is an example in Haskell for operator parsing that uses nothing no more than the standard library: https://hackage-content.haskell.org/package/parser-combinato... Excluding comments it’s less than 50 lines of code, and it handles arbitrary precedence, prefix and postfix, ternary operators, infix operators with all three kinds of associativity.
Store indices into the arena array? You could probably even use 4-byte indices and cut down the memory usage...
> Fib(40) literally took 12+ GIGABYTES before hitting an OOM and crashing. Why? Because it spawns approximately 1.3 Billion nodes.
Okay, maybe you can keep 8-byte indices.
> The mark-and-sweep garbage collector we just completed is a stop-the-world garbage collector. And the algorithm we’re running is inherently exponential.
How about a copying collector then? The recursive Fibonacci generates a lot of garbage but IIRC its live set is actually pretty small at any single point of time. If you need a benchmark for GC when your function actually has a huge live set, then something like
def garbage(n):
if n == 0:
return None
return (garbage(n-1), garbage(n-1))
should do the trick; if you don't have proper data structure you can simulate it with closures pretty trivially.> And we can do something about how we’re evaluating fib itself
You mean "switch from recursively walking AST" or "write a non-exponential Fibonacci"?
In the 1%, though, you'll have needs that malloc/free don't fit. Complex object graphs don't really have a single point of ownership or requiring that you free something in all the places that it might be released is too much to handle. In this case you can reach for a garbage collector (including, for example, implementing reference counting). Other times, you may need to make a lot of allocations in a short time where they can all be freed at once. Request processing in a network server is a common example of this: once the request is complete, everything allocated can be dropped, and you generally want minimal latency.
All of this comes with tradeoffs, though. With a GC, you lose predictability and performance changes; sometimes for the better and sometimes for the worse. Depending on the GC, you may not be able to have stable pointers, and you may lose the ability to finalize objects. With an arena, you can't free or reallocate, so you need to scope the arena to a small region of execution (this is where the author went wrong, for example).
Finally, regardless of which approach you use, you're going to need to thread the allocator through the application; probably implement your own datastructures, etc. Depending on how complex your memory management model is, you may need more than one allocator at any given point (e.g., a GC for the persistent data and an arena per connection and per request). If you're implementing your own allocator, you'll also likely have bugs, and allocator bugs tend to be insidious and obnoxious to debug.
If you can avoid going down that route, I recommend it. Sometimes, though, you have enough constraints that you need to brave the jungle.
It also reminds me of the Emily programming language: https://github.com/mcclure/emily/blob/stable/doc/tutorial.md
Another relatively well known language in this space for its use as an alternative to especially Lua in embedding situations is Io: https://iolanguage.org/
> Perl Contains the Lambda Calculus
> (How to write a 163 line program to compute 1+1)
> Length: 90 minutes
Prerequisites: None.
This is also a fun introduction to building up lambda calculus. And it gets up to fizzbuzz!
When I was in my early teens, I experimented with created programming languages, I distinctly remember one I made on https://esolangs.org
Fun stuff. I should get back into it.