I generate them using LLMs, but optimize them by manually removing chunks or rearranging the order of the instructions. It works well.
Start from scratch, do some test runs, find the bugs, add the minimal possible text to avoid the bug, iterate
You can get 95% of my impl workflow skill by just telling Claude "split the work into slices and use ephemeral subagents" and the other 5% takes like 10x as much text to achieve