DINO: Self-Supervised Vision Transformers

DINO: Emerging Properties in Self-Supervised Vision Transformers | Paper Explained!

Yann LeCun: Deep Learning, ConvNets, and Self-Supervised Learning | Lex Fridman Podcast #36

WHO DO I LOVE MOST?

УСИК против Дерзкого РУССКОГО! Этот Бой Невозможно Забыть!

МНЕ НУЖЕН ЕЩЕ 1 ПОДПИСЧИК! #cat #funny #pets #funnycats #animals #memes #cute

DINO: Emerging Properties in Self-Supervised Vision Transformers

Stanford Contrastive & SS Learning Group

Переглядів 5 505

Додати в
- Мій плейлист
- Переглянути пізніше
Поділитися

Поділитися

Вставка

Розмір відео:

Показувати елементи керування програвачем

Автоматичне відтворення

Автоповтор

Опубліковано 15 чер 2024
Presenter: Michael Zhang
Affiliation: Stanford University
Article's title: DINO: Emerging Properties in Self-Supervised Vision Transformers
Authors: Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, Armand Joulin
Institutions: Facebook AI Research, Inria, Sorbonne University
Paper: arxiv.org/abs/2104.14294
Article's abstract:
"In this paper, we question if self-supervised learning provides new properties to Vision Transformer (ViT) [18] that stand out compared to convolutional networks (convnets). Beyond the fact that adapting self-supervised methods to this architecture works particularly well, we make the following observations: first, self-supervised ViT features contain explicit information about the semantic segmentation of an image, which does not emerge as clearly with supervised ViTs, nor with convnets. Second, these features are also excellent k-NN classifiers, reaching 78.3% top-1 on ImageNet with a small ViT. Our study also underlines the importance of momentum encoder [31], multi-crop training [10], and the use of small patches with ViTs. We implement our findings into a simple self-supervised method, called DINO, which we interpret as a form of self-distillation with no labels. We show the synergy between DINO and ViTs by achieving 80.1% top-1 on ImageNet in linear evaluation with ViT-Base."
Розваги

КОМЕНТАРІ • 16

@adizhol 3 роки тому ⁺⁶
The attention maps visualization is from the output CLS token
@mathildecaron1821 3 роки тому ⁺⁴
Nice video! Minor remark on the last question: we do show comparison with other self-supervised losses for Jaccard distance with deit-S 16x16 in Appendix :). Our conclusion is that the segmented heat maps appear for all the SSL works we experimented with!
@stanfordcontrastivesslearn3141 3 роки тому
Thank you for the comment, that clarifies it! And keep up the good work. Reviewing the paper was very nice!
@AbcDef-xm9rp 3 роки тому ⁺²
It's good to see that you guys are explaining latest SOTA techniques. Keep up the good work guys!
@stanfordcontrastivesslearn3141 3 роки тому
Thanks. Happy to hear that it also helps you guys online!
@AbcDef-xm9rp 3 роки тому ⁺¹
@@stanfordcontrastivesslearn3141 My pleasure!
@piku1920 3 роки тому ⁺²
Hi- For visualisation of masks, it is mentioned in the paper that the mask is obtained by thresholding the self attention maps to keep 60% of the mass. What does the mass represent here? Can you please explain this thresholding technique a bit. Thank you
@kartiksachdev8807 3 роки тому ⁺¹
Great explanation guys!! Is there a slack or discord channel where I could connect with you and contribute in the future?
@stanfordcontrastivesslearn3141 3 роки тому
Hi Kartik, thank you! Very nice to hear that you want to contribute too! Would you like to only participate in the discussion or also present a paper yourself?
@kartiksachdev8807 3 роки тому
@@stanfordcontrastivesslearn3141 thank you for the reply! I would like to present a paper. If that's possible?
@stanfordcontrastivesslearn3141 3 роки тому
@@kartiksachdev8807 Do you already know what paper you would like to present?
@kartiksachdev8807 3 роки тому
@@stanfordcontrastivesslearn3141 yes, I have one paper in mind.
@stanfordcontrastivesslearn3141 3 роки тому
@@kartiksachdev8807 Ok nice! You can write us at stanfordcontrastivelearning [at] gmail.com, send us a bio, let us know what article you would like to present, and we will give you the instructions.
@prof_shixo 3 роки тому ⁺¹
Is this group open to scientific comments or not?!!!! I put a critic for the ViT method and it has been deleted, really weird behaviour!
@stanfordcontrastivesslearn3141 3 роки тому ⁺⁵
Hi Sherif, yes this channel is very open to scientific comments and feedback from the community, thank you very much for participating. I am not sure what happened with your comment. I am only able to see the beginning of your comment in the channel notifications. Do you mind trying to post it again? I suspect it may have been automatically deleted for some reason. I see that we got the notification about your comment twice, so my best guess right now would be that you submitted it twice by accident and that it was detected as spam. But that's just a wild guess. If you are still having issues, just send your comment to stanfordcontrastivelearning [at] gmail.com and we will repost it with quotation marks. We do not want to censor anybody!
@noamzilo6730 Рік тому
This is, like, impossible to, like, listen to, right?

Наступне

Автоматичне відтворення

DINO: Self-Supervised Vision Transformers

DINO: Self-Supervised Vision Transformers

DINO: Emerging Properties in Self-Supervised Vision Transformers | Paper Explained!

DINO: Emerging Properties in Self-Supervised Vision Transformers | Paper Explained!

Yann LeCun: Deep Learning, ConvNets, and Self-Supervised Learning | Lex Fridman Podcast #36

Yann LeCun: Deep Learning, ConvNets, and Self-Supervised Learning | Lex Fridman Podcast #36

WHO DO I LOVE MOST?

WHO DO I LOVE MOST?

УСИК против Дерзкого РУССКОГО! Этот Бой Невозможно Забыть!

УСИК против Дерзкого РУССКОГО! Этот Бой Невозможно Забыть!

МНЕ НУЖЕН ЕЩЕ 1 ПОДПИСЧИК! #cat #funny #pets #funnycats #animals #memes #cute

МНЕ НУЖЕН ЕЩЕ 1 ПОДПИСЧИК! #cat #funny #pets #funnycats #animals #memes #cute

🤯 Прохожу тест на итальянца🍕@nastyawhere

🤯 Прохожу тест на итальянца🍕@nastyawhere

SimCLR: A Simple Framework for Contrastive Learning of Visual Representations

SimCLR: A Simple Framework for Contrastive Learning of Visual Representations

Vision Transformer Quick Guide - Theory and Code in (almost) 15 min

Vision Transformer Quick Guide - Theory and Code in (almost) 15 min

#55 Dr. ISHAN MISRA - Self-Supervised Vision Models

#55 Dr. ISHAN MISRA - Self-Supervised Vision Models

CAP6412 2022: Lecture 25 - Emerging Properties in Self-Supervised Vision Transformers

CAP6412 2022: Lecture 25 - Emerging Properties in Self-Supervised Vision Transformers

Vision Transformer (ViT) - An image is worth 16x16 words | Paper Explained

Vision Transformer (ViT) - An image is worth 16x16 words | Paper Explained

MAMBA from Scratch: Neural Nets Better and Faster than Transformers

MAMBA from Scratch: Neural Nets Better and Faster than Transformers

DINO in PyTorch

DINO in PyTorch

Yann LeCun - Self-Supervised Learning: The Dark Matter of Intelligence (FAIR Blog Post Explained)

Yann LeCun - Self-Supervised Learning: The Dark Matter of Intelligence (FAIR Blog Post Explained)

What is RAG? (Retrieval Augmented Generation)

What is RAG? (Retrieval Augmented Generation)

Самый милый тренд 🥰 #истории #история #новость #новости #shorts

Самый милый тренд 🥰 #истории #история #новость #новости #shorts

😳 МНЕ НУЖЕН ЕЩЕ 1 ПОДПИСЧИК !

😳 МНЕ НУЖЕН ЕЩЕ 1 ПОДПИСЧИК !

Мама помогла Папе 🥹❤️ #shorts #фильмы

Мама помогла Папе 🥹❤️ #shorts #фильмы

Ученые создали игру, которая умеет проникать в глубины подсознания…🤯 #фильм #кино

Ученые создали игру, которая умеет проникать в глубины подсознания…🤯 #фильм #кино

БАТЯ ПЛАКИ-ПЛАКИ

БАТЯ ПЛАКИ-ПЛАКИ

На Какой Заправке Лучшие Хот-доги ?!🌭 #рекомендации

На Какой Заправке Лучшие Хот-доги ?!🌭 #рекомендации

Малыш Борется За Свою Жизнь 😱

Малыш Борется За Свою Жизнь 😱

Неожиданная концовка. Знакомство с оптимусом праймом? #shorts #опрос #прикол #rec #fyp

Неожиданная концовка. Знакомство с оптимусом праймом? #shorts #опрос #прикол #rec #fyp