Loda-lang – language, computational model, and OEIS miner
loda-lang.org
loda-lang.org
seq $0,6005 ; The odd prime numbers together with 1.
max $0,2
and if we go down the rabbit hole, we find that the program for A006005 uses the instruction seq $0,98090 ; Numbers k such that 2k-3 is prime.
where the A098090 program in turn uses seq $3,10051 ; Characteristic function of primes: 1 if n is prime, else 0.
and this pattern continues with 4 programs containing instructions seq $1,80339 ; Characteristic function of {1} union {primes}: 1 if n is 1 or a prime, else 0.
seq $0,38548 ; Number of divisors of n that are at most sqrt(n).
seq $0,94820 ; Partial sums of A038548.
seq $1,94820 ; Partial sums of A038548.
seq $0,132106 ; a(n) = 1 + floor(sqrt(n)) + Sum_{i=1..n} floor(n/i).
Where at last the program for A132106 doesn't use the seq instruction.There also seems to be a danger here of creating cyclic definitions.
OEIS has a number of references from other sequences. If it's a popular sequence then the number of references is high. Usually it's humans that specify the references.
Unlike OEIS, the LODA programs references are mined. The most popular LODA program (A132106) has only a few references in OEIS, maybe it's overlooked in OEIS.
A300402 (program): Smallest integer i such that TREE(i) >= n.
which is claimed to correspond to LODA program
min $0,4
div $0,2
add $0,1
[1] https://oeis.org/A300402These are the current requirements: - the OEIS sequence must have at least 10 terms - the first 100 terms must match (if it has so many terms) - if it has more than 100 terms, we check up to 2000 terms. If the program yields a different term, it is discarded. It is accepted if the program can't be evaluated for n>100 because the maximum number of steps exceeded or the number get too big.
How many of the solution programs were originally hand-written vs machine generated from other known solutions or actually "mined"? Reading through https://github.com/loda-lang/loda-cpp/blob/main/miners.defau..., it seems like the "template" for each generator could have been minted one way or another. I'm a little curious now...
It's kinda fun to see a program minimizer in ~200LoC++, but I wonder what other optimization tricks would be worth implementing for the purposes of integer sequence search? I wonder if it would be worth wiring up a proper compiler framework? Or are there other projects that already do this kind of thing faster?
Instead of just "matching" it might be cool to sort these solutions by similarity and then compare that to the similarity of the sequences themselves.
Finally, given that the OEIS databases exists with "FORMULA" (on every entry?), wouldn't it be faster to just parse that? Maybe you'd find some bugs?
We don't want to parse descriptions or formulas from OEIS. The goal is to do mining just based on the numbers. It is easy to translate a formula into a program. But that's not the goal. The goal is to find NEW or BETTER formulas / algorithms.
1) How can I see the programs that I mine? I ran mine but it's unclear if I've found any?
2) How can I start from the beginning? I'd like to see how long it takes to mine A01, for example
- You will see ALERT logging messages like "First program for..." - The found programs will be written to $HOME/loda/programs/local - In client mode, the files will be also sent to a central server and integrated into the main repository after some days. - We also use a Slack workspace where you can see all new programs. Let me know if you want to join it. https://loda-lang.slack.com/