CoCa: Contrastive Captioners Are Image-Text Foundation Models (2022)
This is a 2022 paper from Google. For me, the key idea is captured in the following passage and two diagrams. Training combines annotated images with noisy image–text pairs. The class labels in
Oct 9, 20263 min read6
