# ML Club Video: Vision Transformers
Date: 2024-02-20
Tags: ViT

Transformers have proven to be kings of Natural Language Processing, but can they be kings in Computer Vision too? A set of engineers at Google set out to answer this question, and they came up with [Vision Transformers](https://arxiv.org/pdf/2010.11929.pdf)! But how does a transformer (which normally takes in words) take in an image as input? Do Vision Transformers provide better performance compared to Convolution Neural Networks? Watch the video to find out!

[Interactive: ML Club Video: Vision Transformers](https://www.youtube.com/embed/U0Hb8nCCOIY?feature=oembed)
