
Tutorial
DINOv3 is Meta's newest vision foundation model — it looks at an image and turns every patch into a 384-dimensional "semantic feature." No labels, no training, no captions. It just knows what things are. What it does TOP in → TOP out. Wire any image/video TOP into it.
