Building efficient machine intelligence.

Modern AI is powerful but expensive. Our mission is to close the gap
between what models can do and what real devices can afford — designing
algorithms so intelligence becomes fast, compact, and deployable everywhere.

We're looking for talented students

The Efficient Machine Intelligence Lab is recruiting motivated Ph.D. students, M.S. students, and interns interested in efficient deep learning. Strong programming and math backgrounds are welcome from all disciplines.

Efficient Machine Intelligence Lab은 효율적인 딥러닝에 관심 있는 열정적인 박사과정·석사과정 학생과 인턴을 모집합니다. 탄탄한 프로그래밍 및 수학 실력을 갖춘 분이라면 전공에 관계없이 환영합니다.

See how to apply · 지원 방법 보기
What we do

Research Directions

Efficient LLMs & Multimodal AI

Token pruning, efficient attention, and serving techniques that make large language and vision-language models cheaper to run at scale.

Scalable Agentic AI

Cutting the token and compute cost of scaling agentic and reasoning systems, so long-horizon LLM agents stay affordable at inference time.

Brain-Inspired & Neuromorphic

Spiking neural networks and event-driven temporal computation that exploit sparsity for energy-efficient inference on neuromorphic hardware.

Continual & Adaptive Learning

Prompt-based continual learning and test-time adaptation so models keep improving in open-world, shifting environments.

Hardware-Aware Co-Design

Algorithm–hardware co-design and in-memory computing that map efficient models onto real accelerators end to end.

Model Compression

Quantization, pruning, and neural architecture search that shrink networks while preserving accuracy for edge and on-device deployment.

Recent News

  • 2026 · Apr Our paper "Real-Time Visual Attribution Streaming in Thinking Model" is accepted to ICML 2026 (Spotlight).
  • 2026 · Feb Our papers "ZOO-Prune: Training-Free Token Pruning via Zeroth-Order Gradient Estimation in Vision-Language Models" and "VisRef: Visual Refocusing while Thinking Improves Test-Time Scaling in Multi-Modal Large Reasoning Models" are accepted to CVPR 2026.
  • 2025 · Dec Our paper "Backpropagation-Free Test-Time Adaptation via Probabilistic Gaussian Alignment" is accepted to NeurIPS 2025.
Selected work

Research Highlights

A few recent papers from the lab and its collaborators.

Real-Time Visual Attribution Streaming
ICML 2026 · Spotlight

Real-Time Visual Attribution Streaming in Thinking Model

Streams visual attribution in real time as a reasoning model thinks, revealing which image regions drive each step of its thought.

VisRef
CVPR 2026

VisRef: Visual Refocusing while Thinking Improves Test-Time Scaling

A test-time strategy that lets multimodal reasoning models revisit visual evidence mid-thought to scale reasoning quality.

ZOO-Prune
CVPR 2026

ZOO-Prune: Training-Free Token Pruning via Zeroth-Order Gradients

Prunes redundant visual tokens in vision-language models with no training, using zeroth-order gradient estimation.

Task Vector Quantization
ICCV 2025

Task Vector Quantization for Memory-Efficient Model Merging

Quantizes task vectors to merge many fine-tuned models into one with a fraction of the memory footprint.

Our collaborators are at…

Google Amazon Meta ByteDance Physion Labs Physion Labs Sungkyunkwan University USC Thaki Cloud Thaki Cloud Oak Ridge National Laboratory