Building efficient machine intelligence.
Modern AI is powerful but expensive. Our mission is to close the gap
between what models can do and what real devices can afford — designing
algorithms so intelligence becomes fast, compact, and deployable everywhere.
We're looking for talented students
The Efficient Machine Intelligence Lab is recruiting motivated Ph.D. students, M.S. students, and interns interested in efficient deep learning. Strong programming and math backgrounds are welcome from all disciplines.
Efficient Machine Intelligence Lab은 효율적인 딥러닝에 관심 있는 열정적인 박사과정·석사과정 학생과 인턴을 모집합니다. 탄탄한 프로그래밍 및 수학 실력을 갖춘 분이라면 전공에 관계없이 환영합니다.
See how to apply · 지원 방법 보기Research Directions
Efficient LLMs & Multimodal AI
Token pruning, efficient attention, and serving techniques that make large language and vision-language models cheaper to run at scale.
Scalable Agentic AI
Cutting the token and compute cost of scaling agentic and reasoning systems, so long-horizon LLM agents stay affordable at inference time.
Brain-Inspired & Neuromorphic
Spiking neural networks and event-driven temporal computation that exploit sparsity for energy-efficient inference on neuromorphic hardware.
Continual & Adaptive Learning
Prompt-based continual learning and test-time adaptation so models keep improving in open-world, shifting environments.
Hardware-Aware Co-Design
Algorithm–hardware co-design and in-memory computing that map efficient models onto real accelerators end to end.
Model Compression
Quantization, pruning, and neural architecture search that shrink networks while preserving accuracy for edge and on-device deployment.
Recent News
- 2026 · Apr Our paper "Real-Time Visual Attribution Streaming in Thinking Model" is accepted to ICML 2026 (Spotlight).
- 2026 · Feb Our papers "ZOO-Prune: Training-Free Token Pruning via Zeroth-Order Gradient Estimation in Vision-Language Models" and "VisRef: Visual Refocusing while Thinking Improves Test-Time Scaling in Multi-Modal Large Reasoning Models" are accepted to CVPR 2026.
- 2025 · Dec Our paper "Backpropagation-Free Test-Time Adaptation via Probabilistic Gaussian Alignment" is accepted to NeurIPS 2025.
Research Highlights
A few recent papers from the lab and its collaborators.
Real-Time Visual Attribution Streaming in Thinking Model
Streams visual attribution in real time as a reasoning model thinks, revealing which image regions drive each step of its thought.
VisRef: Visual Refocusing while Thinking Improves Test-Time Scaling
A test-time strategy that lets multimodal reasoning models revisit visual evidence mid-thought to scale reasoning quality.
ZOO-Prune: Training-Free Token Pruning via Zeroth-Order Gradients
Prunes redundant visual tokens in vision-language models with no training, using zeroth-order gradient estimation.
Task Vector Quantization for Memory-Efficient Model Merging
Quantizes task vectors to merge many fine-tuned models into one with a fraction of the memory footprint.