CvT
This is an official implementation of CvT: Introducing Convolutions to Vision Transformers.
About
This is an official implementation of CvT: Introducing Convolutions to Vision Transformers. We present a new architecture, named Convolutional vision Transformers (CvT), that improves Vision Transformers (ViT) in performance and efficienty by introducing convolutions into ViT to yield the best of both designs. This is accomplished through two primary modifications: a hierarchy of Transformers containing a new convolutional token embedding, and a convolutional Transformer block leveraging a convolutional projection. These changes introduce desirable properties of convolutional neural networks (CNNs) to the ViT architecture (e.g. shift, scale, and distortion invariance) while maintaining the merits of Transformers (e.g. dynamic attention, global context, and better generalization). We validate CvT by conducting extensive experiments, showing that this approach achieves state-of-the-art performance over other Vision Transformers and ResNets on ImageNet-1k, with fewer parameters and lower FLOPs. In addition, performance gains are maintained when pretrained on larger dataset (e.g. ImageNet-22k) and fine-tuned to downstream tasks.
Open Source Health
- Stars
- 605
- Forks
- 130
- License
- MIT
- Last commit
- 3 years ago
Resources & Links
Related Categories
Vendor
Microsoft
Empower every person and organization to achieve more.
Quick Links
Open Source
More by Microsoft
Related Products
Rclip
Semantic photo search for the command line
Top category match
ChineseNER
A neural network model for Chinese named entity recognition
Top category match
Head Pose Estimation
Realtime human head pose estimation with ONNXRuntime and OpenCV.
Top category match
Video Subtitle Extractor
视频硬字幕提取,生成srt文件。无需申请第三方API,本地实现文本识别。基于深度学习的视频字幕提取框架,包含字幕区域检测、字幕内容提取。A GUI tool for extracting hard-coded subtitle (hardsub) from videos and generating srt files.
Top category match
U 2 Net
The code for our newly accepted paper in Pattern Recognition 2020: "U^2-Net: Going Deeper with Nested U-Structure for Salient Object Detection."
Top category match