The problem isn't that P=.05 is an arbitrary measure of significance. The problem is that only publishing significant results is a bias against the null hypothesis.
Let's say you're doing a study of flipping coins. The null hypothesis is that the coin is evenly weighted. If the null hypothesis is true, when you flip a coin once, it will come up heads with P=.5. If you flip the coin twice, the null hypothesis is that both flips come up heads with P=.25. The probability of all coins coming up heads is P=0.125 for 3 flips, P=0.0625 for 4 flips, and P=.03125 for 5 flips. So if we flip a coin 5 times, and get heads all 5 times, we can conclude that the coin is weighted in some way with P=0.03125.
Let's say all the major journals of coin flipping only publish results with the high significance of P<0.05. Alice flips a quarter 5 times and gets 3 heads and 2 tails, and nobody will publish her study because it has P=.3125 (see [1] for an intuitive explanation of how this was calculated). Bob flips a quarter 5 times and gets 2 heads and 3 tails--again, no journal will publish him. Catherine, David, Ellen, Frank, and Geri all perform the same experiment, most of them getting groupings 3:2 result ratios, some getting 4:1 result ratios, but nobody getting all heads or all tails, just as one would expect given the null hypothesis. And the journal editors tirelessly send out rejection letters to all their studies.
Now somewhere down the line, Robert flips a coin 5 times and gets 5 heads. This result has P=.03125, which meets the requirement of P<.05! He sends it to the American Journal of Coin Flipping Studies (AJCFS), and they are very excited to publish his results! Nature and Science magazines do front page pieces with headlines "Quarters Found Heavy-headed" and "Washington Shows His Face" respectively. A casino hires Robert as a consultant for the design of their coin-flipping games. His quarter-flipping study is cited in the abstracts of two dime-flipping studies and a half-dollar flipping study. During the trial of a murderer who placed quarters tails-side up on his victims, Robert is called as an expert witness to say that the coins were placed there, not flipped there.
A week later Sally flips a quarter 5 times and gets 2 heads and 3 tails. She sends the results of her study to the AJCFS noting a failure to reproduce Robert's result, but her study is rejected because it has P=.3125.
Now, if you survey the AJCFS and all the other academic journals on coin flipping, you'd conclude that quarters are significantly weighted towards heads. But in fact, the null hypothesis is true: quarters are pretty evenly weighted. Robert's low-P result is exactly what you'd expect to happen eventually if you have enough people perform the 5-quarter-flip experiment--in fact, if a lot of people are studying coin flips, the P of getting a low-P result approaches P=1. But because the AJCFS has a P=.05 requirement, they've created a bias against the null hypothesis, which deceives the public into thinking that flipped quarters are more likely to come up heads than tails.
This is likely the reason why so many fields, most notably psychology[2], are having a replication crisis[3] and a similar effect can be used in P-hacking[4] to bolster results that are essentially fake.
Unlike coin-flipping, fields with replication crises like psychology and medicine have real affects on real people's lives. It's irresponsible and unethical for journals to publish with a bias against the null hypothesis, and adjusting the P-value requirements to another significance requirement, even a less arbitrary one, doesn't fix the issue.
The solution, I think, is for journals to commit to publish studies before the study has been performed, based on the methodology, previous studies on the subject, and qualifications of the researcher. This would mean that many, many studies would be published with null results, and that would be a good thing.
[1] There are 32 possible outcomes for flipping a coin 5 times. If we group them by how many heads and tails, we can calculate a probability for each outcome:
HHHHH 1 result of 5 heads -> P = 1/32 = .03125
HHHHT
HHHTH
HHTHH 5 results of 4 heads, 1 tails -> P = 5/32 = .15625
HTHHH
THHHH
HHHTT
HHTHT
HHTTH
HTHHT
HTHTH
HTTHH 10 results of 3 heads, 2 tails -> P = 10/32 = .3125
THHHT
THHTH
THTHH
TTHHH
HHTTT
HTHTT
HTTHT
HTTTH
THHTT 10 results of 2 heads, 3 tails -> P = 10/32 = .3125
THTHT
THTTH
TTHHT
TTHTH
TTTHH
HTTTT
THTTT
TTHTT 5 results of 1 heads, 4 tails -> P = 5/32 = .15625
TTTHT
TTTTH
TTTTT 1 result of 5 tails -> P = 1/32 = .03125
[2] https://thepsychologist.bps.org.uk/what-crisis-reproducibili...[3] https://en.wikipedia.org/wiki/Replication_crisis
[4] https://journals.plos.org/plosbiology/article?id=10.1371/jou...
EDIT: Also see roenxi's excellent post on how "significant" means different things in statistics and colloquial English: https://news.ycombinator.com/item?id=20895893