Making AI smaller, faster, and more power-efficient — we pursue technologies that run advanced deep learning even under tight computational constraints.
Deep learning has achieved remarkable progress in fields such as image recognition and natural language processing, but its high accuracy comes at the cost of enormous computation and memory consumption. Large models typically assume powerful cloud GPUs, so real-time inference on edge devices — smartphones, drones, sensors — remains a persistent challenge.
Real-world constraints such as power consumption, heat, communication latency, and privacy further hinder the social deployment of AI. We return to our lab's founding focus — high-performance computing (HPC) and computer architecture — to tackle this challenge from both software and hardware angles.
Combining pruning, quantization, and knowledge distillation to substantially reduce model size and computation while preserving accuracy. We automatically analyze per-layer redundancy and develop optimization methods (e.g. IHSOpti) that maximize hardware parallelism.

Researching "token pruning" that dynamically removes low-importance tokens from today's mainstream ViTs. We focus on the attention-reliability problem in shallow layers, enabling depth-adaptive token selection and independent pruning of Mask2Former's multi-scale features.

Designing custom accelerators for FPGA / FPGA-DPU / RISC-V using quantization-friendly attention and SIMD instructions. Our DPU (Deep Learning Processor Unit) is a programmable engine for CNNs, made even more compact by pruning existing models for DPU acceleration. We also use High-Level Synthesis (HLS) to generate FPGA circuits directly from C/C++ for deep learning hardware design.
Processor architecture is a foundational technology in computer science, demanding ever-greater speed as its applications expand. We have proposed value- and branch-prediction methods to address the data and control dependencies that hinder speedup, along with several schemes to mitigate misprediction penalties.

To address the communication bottleneck when splitting ViT processing between edge and cloud, we propose a method that reduces transferred data volume through cosine-similarity-based token clustering.