Full threadlogotype·Try it out and see how fast inference can be for agentic workflows. Works on CUDA and ROCm. Feedback appreciated!View on HN