As someone who has dabbled in AI generated crosswords I found that providing samples of "good crossword clues" (which I curated from historical NYT monday puzzles) as part of the LLM context helped tremendously in generating better clues.
Part of the deep satisfaction in solving a crossword puzzle is the specificity of the answer. It's far more gratifying to answer a question with something like "Hawking" then to answer with "scientist", or answering with "mandelbrot" versus "shape".
So ideally, you want to lean towards "specificity" wherever possible, and use "generics" as filler.
Producing a valid layout is a different situation. Ultimately depends on how strict you are in terms of acceptable layouts. Getting a NYT compatible crossword puzzle is going to be more difficult by virtue of its symmetrical requirements than if you're okay with just laying out sequences of perpendicular intersecting words.