ParentFull threadsylware·Something does not add up there: inference on interesting models is already very slow by itself on consumer systems. "Sending to a GPU" is literaly "zero time".View on HN