Pretty interesting stuff. I'm amazed that there are only 14 nodes in the cluster with 1TB SSD each; message size must be fairly small on disk (less than 295 bytes, unreplicated). I'd also recommend looking at maybe tuning the shard_request_cache; there are some possible improvements to be made there, if you're running that many indices. Finally, are all of the indices of approximately the same size? Are they time-boxed?
edit: Put together a logging cluster consisting of 14 nodes, 12 data + 2 indexer/search API, with ~40TB of consumer-grade SSDs. Ran ~14k indices based on log type & timestamp, with a whole raft of custom field configuration to handle aggregations and different tokenizations. In sum, Elasticsearch looks easy to configure and tune but is amazingly hard to do well - but incredibly rewarding.