We’re in a constant cycle of “sure the current stuff doesn’t live up to the hype but the next release that’s just on the horizon, that one will blow you away!”
This leak is clearly targeted at creating the buzz necessary to raise the money to keep the pipe dream flowing for the time being. But just thinking through the steps outlined here, it should be clear none of this makes sense. Any “correct” synthetic training data is either going to be badly biased, or be limited to such banal “logical” output that the model it trains will be entirely unable to process natural language input and you’ll need to specify your questions in something more precise. At which point, we’re back to traditional programming languages, only with several layers of unnecessary and expensive processing going on in between.