Learning to Communicate with Deep Multi-Agent Reinforcement Learning
github.com
github.com
When I was a teen I read Steven Levy's book "Artifical Life". It captured my imagination in a way that the biomorph program in a Blind Watchmaker also had. Ever since then I've played with toy alife models. Some of the attraction is just aesthetic - it's fun to set these things up and design the rules, kinda like a higher-order god-game. Another is the enticing fantasy that something like this will one day lead to AGI. A third is the occasionally fulfilled desire to see the fabled emergent behavior of evolutionary computations, where the agents learn to do something you never expected, or exploit some funky loophole to their advantage.
When I was an undergrad I created an alife system as an honors project. It was an ambitious project, with sexually or asexually reproducing agents controlled by recurrent neural nets. The agents had perception and the ability to move, reproduce, attack, co-operate, or interact with objects. It had an entity framework to make it possible to design things like programmable Skinner boxes, traps, feeders, it had feeding schedules, it had live visualization and a slick Qt interface to edit the entities inline or take control of specific agents and visualize their brains.
I was so proud of it.
My assigned supervisor, who was an austere Georgian category theorist, had zero interest. He compared it unfavorably to a virus. Anyway, it was so dispiriting an experience that I kind of swore off Alife after that. But the fun I had made me realize I didn't want to be a mathematician, so I suppose it played a useful role in my life after all.
With latest developments in ML/AI space this is a very interesting period to be living in. A few decades later they will probably say: "remember when there was no artificial agents and the computers were those dumb boxes?" I hope you find your way back and join the fun!
This type of approach is what led me to the current data model of ibGib. Each ibGib address is a node and will inevitably compete for attention "resources". The interesting aspect is when you think of any agent's snapshot in time as being the source of a fork which could itself be an agent. So each point in time of an agent configuration is itself its own possible branching agent configuration in thr future.
It is monotonically increasing because you never know which agent's state will come in handy to get out of a local max/min. Also, it reuses agent dna (pointers are cheap) so overall this is not as expensive as it sounds.
And worst case is that if the overall data store is strained for storage resources, any individual ibgib has its dependencies. So you can do a dependency graph projection, essentially creating a slimmer copy of the original.
Btw Your username reminds me of a book series I read as a kid...Taliesin was King Arthur's father maybe?