- Is it guaranteed that the ruleset can produce any sequence of collisions?
- Seems like the vast majority of possible states would represent words with concurrent phonemes, which is pretty much undefined behavior for the human mouth. I guess you don't want the ideograms to be too informationally-dense, but that's a pretty sparse encoding.
- Example table only does CV syllables and it's not immediately obvious to me how to extend it to more syllable types. (especially crazy consonant-dense clusters like "twelfths") You could have ∅ entries in each vector, but then you start to lose some of the natural "timing" analogy if some collisions don't represent syllables.
- Not a lot of visual distinction between different words (e.g. "like" and "site" are just 1 step vertically transposed from each other).
- Each word has many many possible ideograms, not just because each collision can be generated multiple ways, but also since you can create non-colliding "noise" cells in unused areas.
I guess you could leverage the expressiveness somehow, like tending to use certain collision types or cell states implies something about the context / pragmatics.
I do really like the out-of-the-box thinking that inspired this, but I also like seeing it taken to the rational extreme. How can you get the dynamic nature of the cellular automata shine in your system? What distinguishes it from using the same table, but just putting numbers in the cells for when each combination occurs?