Show HN: I reduced LLM inference GPU calls by 94% using semantic routingicomnewtechnologies.com·2 pts·kanacki·1