On a more serious note, I'm sure something like this exists already. I wasn't even the first one to write a Dota picker, I got the idea from someone else's tool. I even saw one that could analyse the display and recognise the hero images to enter the existing picks automatically!
Even my little utility merely scratched the surface. It didn't handle the first few picks well, because the choices are not independent. Depending on the game mode chosen, each team can take turns picking. In the fully competitive variant of the game, players can even exclude ("ban") heroes from that individual match.
I only optimised mine for picking the last one or two heroes out of the ten players, because I either played solo or with one other friend. We'd pick last based on my tool's recommendations.
That alone was hard enough to program! Theoretically, it's a trivial probability problem. Just build a 10-d matrix of win rates, and then picking is trivial. Unfortunately, this is a huge amount of data, and impossible to train well. Reducing it to, say, the 3-d case of ally+ally+oppose and computing that over every combination in the current picks doesn't cover things like the support/lead roles well, so it over-recommends some heroes. Some heroes do well with short games, some with long. Some can reduce the game length, others can drag it out. Some are "greedy", consuming shared team resources, others are frugal. I ended up with a complex heuristic model that included all of these high-level traits in its combined recommendation plan.
Fundamentally, there's nowhere near enough data available for a full model, and the labels are super noisy, because even a good combo may only shift the win rate by 10 or 20 percent at best. You might see thousands of samples for some popular combination of strong heroes, but then you'd get a long tail of individually unique games.
I found a theoretical mathematical paper that covered this exact training scenario, and it basically concluded that this is an entirely new branch of probability theory that has been essentially unexplored. In other words: "Good luck with that, we couldn't solve it either!"
It's certainly a fun problem space to play with. There's tons of data, and you can immediately test the output yourself, personally. There's no end to the depth of it either! You can go to the n-th degree and even start recommending ideal hero builds (i.e.: which items to buy in-game), where each hero should be positioned, etc...
And I think even a human would struggle winning the character select mini game just with tabular data related to the players' win rates with those characters in previous matches. Learning about actual play styles of each playerand customising their play style with a given character would be a huge factor in them gaining an edge.
For example, let's say there is a player picking a slow but tanky character, but I happen to know the player who picked it often attempts a certain cheeky move to gain an advantage (like jungling early in the case of LoL, I really don't know much about Dota or LoL so you'll have to work with the metaphor here) and THAT cheeky technique is what increases their win rate. But then I know of a move that can be done by a certain character only in that specific cheeky scenario that will tip the scales and let me get a kill on them early, tipping the entire balance of the game early. As a pro player, that could be why I win with my selection, but the cheeky strategy also why the other guy often wins with his, even against other players who pick mine.
There's essentially a list of known play styles between different characters and optimal strategies to use with them against certain other characters. I think when you filtered to the top 10% of players, you essentiallly made sure a larger percent of those optimal play styles were the ones that generated the outcome data.
But still, if the player doesn't know what those in game strategies are, or their opponent doesnt know the ones they're "supposed" to be using, it'll throw off the value of making it the best selection.
In retrospect, a neural net would have likely been the best approach, but the (maximal!) noise in the training data would have made it difficult to train. I wanted a billion or more game outcome records to enable NN training, but the game stats API is very throttled and it would have taken months to get that much data. In that timeframe the game is often patched with new rules, which would invalidate the older data.
An interesting effect I noticed was that the pub games have all sorts of weird second-order effects.
E.g.: Some characters are unpopular because pub players play them ineffectively, and/or they're hard to play well because they're so niche, so few people get enough practice with them in the right kind of scenario.
Dota's Jakiro is a great example. It's a lumbering support hero with very slow spells and a glacial turn rate. It's like trying to do acrobatics with a jumbo jet. He's frustrating to play and when most people pick him, the effect on the overall win rate is about -15%, which is insanely bad for 1 out of 10 players in a game! My win rate with him was something like 25%, which is absolutely atrocious. That's "throwing the game" bad.
My picker app often recommended Jakiro when the enemy team had over-represented heroes that typically commit to a fight, such as Legion Commander or Axe. Those heroes are nailed to the ground in a fight and can't escape Jakiro's devastating-but-slow attacks.
I ended up playing Jakiro a lot, and eventually I got about a 55% overall win rate with him, which is an amazing swing if you think about it. It took practice though, figuring out how to best utilise him. If I had played him in random games, I never would have gotten anywhere. The picker app however made this possible, by providing these hints of when and how to use the hero.