Very much agreed. I briefly tried to use poetry, tox, and pytest, and even using the support packages that are supposed to make poetry and tox play nicely together, it was a massive failure. tox would claim to be running tests with python3.8, say, but actually it was still using the default system python3.10, etc.
I now refuse to use any packaging tools for anything other than generating a requirements.txt, and refuse to use any version matrix tools like tox. I don't have a preference for/against pyenv for sourcing python interpreters (I'll use an AUR or scoop sometimes), or poetry or pipenv or anything else for specifying dependencies, but whatever method is used, I will always only use them to export a requirements.txt, and I will manually create my own venv for each interpreter, and run anything I want inside each venv via oneliners like:
bash -c '. "$0/bin/activate" && exec "$@"' venv38 pip install --upgrade pip setuptools wheel
bash -c '. "$0/bin/activate" && exec "$@"' venv38 pip install -r requirements.txt
bash -c '. "$0/bin/activate" && exec "$@"' venv38 pytest --doctest-modules -v
Usually I'll wrap that `bash -c` preamble into a simple with_venv.sh script, and make a similar with_venv.ps1 for Windows. Create a Makefile with targets for those commands for each interpreter you want to test, and you can run everything in knowledge that you'll always know what's happening, instead of relying on stacked magic.