very cool! can we also just oultine trace / approx as hints on input image (or feed it two images, say raw and markup image with whatever visual hints use?
the model does come with bounding box capability but we did not optimize for it in this release. In the next releases we intend to tackle that along with video capabilities and some other features.