You know, there is at least a dozen chatbot providers who can handle these nonstructured dialogues with multiple entry points, ESPECIALLY with pizza.
In fact, the pizza order is the No. 1 scenario looked at by the chatbot providers.
In fact, it was exactly the scenario my old startup took as a case study and the first application we built with it. It could handle the different toppings, the sizes, and more. You could submit all your requests in one move, it would be parsed and sorted into its little slots.
The problem? There is only a handful of scenarios similar to the pizza. In most cases in the real world, you have to select from an external database, look at proprietary product names, and more. Another staple of chatbot demos, plane ticketing, only works well when limited to North America (in the English word). Good luck asking for a flight to Kinshasa, Kuala Lumpur, or even Wagga Wagga in Australia.
I am not even talking about the switchboards for multiple domains, like in Alexa. These ones only work with "leaky abstraction" (making the user learn magic keywords).
Another problem is really stupid. It's the availability of the datasets. The funny thing is, ye olde style semantic frameworks fare better than the machine learning ones, because there is not enough data for the machine learning chatbot frameworks, and without it, their mighty capabilities are pretty much the proverbial spherical cow in vacuum. But because the semantic paradigm is not kosher/kewl anymore, very few enterprises agree to deploy it.
None of that matters though because the users never liked typing a lot. Back in 1980s - 1990s the adventure type computer games switched from (mostly working) command line interface to point-and-click, and very few users objected.
My take is, the key is a conversational UI with strong visual feedback. For the pizza scenario above, I would draw icons of cheese and numbers, so that the user can be sure it worked.