MIT HAN Lab
Efficient AI Computing. PI: Song Han
Consolidated Product Portfolio
StreamingVLM: Real-Time Understanding for Infinite Video Streams
[ICML 2024] Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
[MICRO'23, MLSys'22] TorchSparse: Efficient Training and Inference Framework for Sparse Convolution on GPUs.
[MLSys'25] QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving; [MLSys'25] LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
[HPCA'21] SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
[CVPR 2020] GAN Compression: Efficient Architectures for Interactive Conditional GANs
[NeurIPS 2025] Radial Attention: O(nlogn) Sparse Attention with Energy Decay for Long Video Generation
[ACL'20] HAT: Hardware-Aware Transformers for Efficient Natural Language Processing
[IJCV] FastComposer: Tuning-Free Multi-Subject Image Generation with Localized Attention
[NeurIPS 2020] Differentiable Augmentation for Data-Efficient GAN Training
[CVPR 2021] Anycost GANs for Interactive Image Synthesis and Editing
[ICLR 2019] ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware
TinyChatEngine: On-Device LLM Inference Library
[ICLR 2020] Once for All: Train One Network and Specialize it for Efficient Deployment
[ICML 2023] SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
[MLSys 2024 Best Paper Award] AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
[ICLR 2024] Efficient Streaming Language Models with Attention Sinks
Company Profile & Strategy
Effizientes KI-Computing. PI: Song Han. Das MIT HAN Lab verwaltet 16 Open-Source-Projekte auf GitHub, darunter Streaming LLM, LLM Awq und Once For All. Primäre Sprachen: Python, Cuda, C++.
Open-Source Footprint
Related vendors
Save & manage the vendor's products in Productapp.
Save and manage this vendor's products in Productapp.