Okebiz Video Search



Title:DINO: Emerging Properties in Self-Supervised Vision Transformers
Duration:52:32
Viewed:2,436
Published:07-05-2021
Source:Youtube

Presenter: Michael Zhang
Affiliation: Stanford University

Article's title: DINO: Emerging Properties in Self-Supervised Vision Transformers
Authors: Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, Armand Joulin
Institutions: Facebook AI Research, Inria, Sorbonne University
Paper: https://arxiv.org/abs/2104.14294

Article's abstract:
"In this paper, we question if self-supervised learning provides new properties to Vision Transformer (ViT) [18] that stand out compared to convolutional networks (convnets). Beyond the fact that adapting self-supervised methods to this architecture works particularly well, we make the following observations: first, self-supervised ViT features contain explicit information about the semantic segmentation of an image, which does not emerge as clearly with supervised ViTs, nor with convnets. Second, these features are also excellent k-NN classifiers, reaching 78.3% top-1 on ImageNet with a small ViT. Our study also underlines the importance of momentum encoder [31], multi-crop training [10], and the use of small patches with ViTs. We implement our findings into a simple self-supervised method, called DINO, which we interpret as a form of self-distillation with no labels. We show the synergy between DINO and ViTs by achieving 80.1% top-1 on ImageNet in linear evaluation with ViT-Base."

SHARE TO YOUR FRIENDS


Download Server 1


DOWNLOAD MP4

Download Server 2


DOWNLOAD MP4

Alternative Download :



SPONSORED
Loading...
RELATED VIDEOS
DINO: Emerging Properties in Self-Supervised Vision Transformers (Facebook AI Research Explained) DINO: Emerging Properties in Self-Supervised Vi...
39:13 | 48,967
DINO: Emerging Properties in Self-Supervised Vision Transformers | Paper Explained! DINO: Emerging Properties in Self-Supervised Vi...
31:54 | 459
Yann LeCun: Yann LeCun: "Energy-Based Self-Supervised Learn...
00:10 | 24,175
Yann LeCun - Self-Supervised Learning: The Dark Matter of Intelligence (FAIR Blog Post Explained) Yann LeCun - Self-Supervised Learning: The Dark...
58:37 | 53,378
DETR - End to end object detection with transformers (ECCV2020) DETR - End to end object detection with transfo...
09:35 | 950
In the Age of AI (full film) | FRONTLINE In the Age of AI (full film) | FRONTLINE
54:17 | 335,516
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (Paper Explained) An Image is Worth 16x16 Words: Transformers for...
29:56 | 129,589
Multi-task attention-based semi-supervised learning for medical image segmentation Multi-task attention-based semi-supervised lear...
28:01 | 91

shopee ads

coinpayu