Year 2026
Journal
-
Ryuto ISHIBASHI, Hayata KANEKO, and Lin MENG, "Rethinking Attention Reliability for Token Pruning in Vision Transformers," [Download] SCI/SCIE IF:6.5 JCR Q1 CAS Q2Vision transformers; Token pruning; Token importance estimation; Static layer-wise routing; Efficient inference概要を表示
Vision Transformers (ViTs) achieve strong performance in visual recognition but incur quadratic computational cost with respect to the number of tokens, motivating extensive research on token pruning and reduction. Most existing pruning methods estimate token importance directly from attention weights, implicitly assuming that attention magnitude provides a reliable proxy for semantic relevance across all layers. Our analysis shows that the validity of this assumption varies substantially with transformer depth and model scale, and can also depend on the training paradigm. Through a systematic analysis of attention selectivity using multiple concentration and stability measures, attention distributions in shallow layers tend to be highly diffuse and weakly discriminative, making attention-based scoring unreliable for early-stage pruning. In contrast, attention becomes increasingly informative in deeper layers as token representations mature. Motivated by this observation, a broad range of non-attention importance scores derived from token embeddings, including statistics- and similarity-based criteria, is examined. Across ViT variants and diverse training settings, these non-attention scores exhibit more stable pruning behavior in shallow layers, whereas attention-based scoring becomes effective only after sufficient representational discrimination is achieved. Importantly, the depth at which this transition occurs is model-dependent and not strictly monotonic, indicating that uniform attention-based pruning is fundamentally mismatched to the representational dynamics of ViTs. Based on these findings, token pruning is formulated as a layer-wise selection problem governed by the reliability of attention, and lightweight static routing configurations are investigated without retraining or dynamic inference control. For equivalent FLOPs, the resulting pruning patterns achieve a trade-off between accuracy and efficiency that is comparable to or superior to that of representative token reduction methods. Overall, these results establish token importance estimation in ViTs as an inherently layer-dependent problem shaped by representation maturity, model characteristics, and training paradigm rather than uniform attention magnitude.
-
ZhenLing SU, YeXin ZHANG, and Lin MENG, "Artificial Intelligence on the Wing: Fully On-board Visual Servoing for Object Tracking with Autonomous Nano Unmanned Aerial Vehicles," [Download] SCI/SCIE IF:8.0 JCR Q1 CAS Q1 topNano unmanned aerial vehicle; Fully on-board; Autonomous; Artificial intelligence; Visual servoing; Object tracking概要を表示
Nano Unmanned Aerial Vehicle (UAV) platforms are well-suited for tasks in confined spaces, such as indoor single-person tracking. However, they are constrained by payload, and milliwatt (mW) level computing budgets. To address these limitations, we present a fully on-board tracking system featuring Feather You Only Look Once (FeatherYOLO), an ultra-lightweight detector tailored to the embedded processor, and a custom image-based visual servoing controller deployed on a 29 gram Crazyflie 2.1 nano UAV. FeatherYOLO utilizes a depthwise-separable backbone with a decoupled, anchor-free head, requiring 20 thousand parameters, 1.94 million multiply-accumulate operations, and 224 kilobytes of memory. On our self-collected indoor human detection dataset (six participants across five sites) under a cross-subject and cross-environment held-out protocol, the model achieved 99.5% mean Average Precision (mAP) at an Intersection Over Union threshold of 0.50 and 75.0% mAP over thresholds from 0.50 to 0.95. On-board profiling reveals that pure inference consumes 46.5 mW at 150 Frames Per Second (FPS), accounting for less than 1% of the total flight power, and outperforms a recent nano UAV baseline consuming 225.7 mW at 43 FPS. The proposed visual servoing strategy was refined through flight trials and task-driven tuning, transforming detector outputs into stable, bounded velocity commands with hysteresis and filtering for closed-loop indoor tracking. Real-world flight tests validated the tracking performance with an average tracking success rate of 90.0%, succeeding in 36 out of 40 experimental runs. The primary flight challenges identified include collisions and target loss.
-
Zenghui Wang, Zhihong Man, Lin Meng, Shijian Cang, and Yanxia Sun, "AI-driven digital twin and delay-aware surrogate MPC framework for biogas production," [Download] SCI/SCIE IF:4.0 JCR Q1 CAS Q2Anaerobic digestion; Biogas; Digital twin; Machine learning; Model predictive control; Bayesian optimization概要を表示
Anaerobic digesters exhibit nonlinear dynamics, long input–output delays, irregular sampling, and operational constraints that complicate biogas prediction and control. This study develops a delay-aware digital-twin MPC benchmarking framework in which Anaerobic Digestion Model No. 1 (ADM1) serves as a mechanistic reference plant, while established machine-learning surrogates (Random Forest, KNN, SVR, XGBoost, LSTM, and TabPFN) provide fast one-step predictions under irregular measurements. A unified workflow integrates time-stamp alignment, sliding-window reconstruction, and Bayesian hyperparameter optimization. The surrogates are evaluated on an industrial dataset and an ADM1-based simulator incorporating a 7-day actuator delay, seasonal variability, noise, and missing data. The trained models are embedded in a constrained MPC layer, where multi-day inputs are optimized using Bayesian Optimization or Particle Swarm Optimization under hard bounds and daily ramp-rate limits. Both open-loop replay and closed-loop digital-twin MPC are investigated. Results show that PSO–MPC with inexpensive surrogates achieves the largest methane gains (up to approximately 25%), whereas BO–MPC is preferable for computationally expensive surrogates due to superior sample efficiency. Closed-loop simulations demonstrate that steady-state performance is preserved through feedback correction despite surrogate mismatch. The primary contribution is a reproducible digital-twin MPC scaffold enabling systematic integration and benchmarking of surrogate–optimizer combinations. The framework provides a reusable evaluation testbed for data-driven control of slow, delay-dominated biochemical processes, with potential extension to other chemical and energy systems subject to long delays and irregular monitoring.
-
Zhaowei Sun, Jintao Chen, Cong Lin, Xuebin Yue, Kuozhan Wang, and Lin Meng, "YOLO-CF: An Object Detection Model that Improves Feature Expression at Both Coarse-Grained and Fine-Grained Levels for Industrial Surface Image Defect Detection," [Download] SCI/SCIE IF:1.21object detection pooling operation dimension reduction multi-scale feature fusion概要を表示
In industrial surface defect detection, enhancing the model’s capability to express features at both coarse-grained and fine-grained levels is crucial. Accordingly, this paper proposes YOLO Coarse and Fine (YOLO-CF), a novel object detection model that significantly enhances feature expression by integrating innovative feature fusion strategies with an improved network architecture. YOLO-CF incorporates Res2Net and Res2Net block modules, notably improving the expression of fine-grained features without increasing computational complexity. Additionally, this model introduces a multi-scale feature fusion module, which seamlessly integrates coarse and fine-grained information by combining top-down and bottom-up pathways. This enhancement effectively expands the perceptual range and significantly improves the model’s generalization capability, making YOLO-CF a powerful tool for detecting diverse defects in complex industrial images. A hybrid downsampling module is introduced, combining max pooling, average pooling, and convolution operations with a stride of 2 to provide richer feature representations. In the GC10-DET dataset, YOLO-CF achieved a mean average precision (mAP) of 59.09%, surpassing the second-ranked RetinaNet by 3.41 percentage points. On the PCB, crack, NEU-DET and Dish-20 public datasets, YOLO-CF achieved mAPs of 97.38%, 84.72%, 73.35%, and 99.54%, at IoU =0.5, respectively. The experimental results indicate that by integrating feature extraction at both coarse-grained and fine-grained levels, YOLO-CF effectively enhances the model’s ability to detect objects of various sizes in complex scenes, demonstrating significant performance improvements. The code is available at http://www.ihpc.se.ritsumei.ac.jp/obidataset.html.
-
Takashi Higashi, Lin Meng, and Ryuto Ishibashi, "Enhancing computational efficiency in video action recognition via temporal token merging in ViViT," [Download]video recognition action recognition ViViT token merging
概要を表示
This study proposes a temporal token merging (TTM) method for ViViT-based action recognition to improve computational efficiency while maintaining accuracy. The method merges temporally similar tokens at the same spatial locations across frames, reducing redundancy without discarding information. Experimental results show that the proposed method achieves up to 39.2% reduction in FLOPs while limiting accuracy degradation to less than 1%. Compared with frame pruning, which reduces FLOPs but significantly degrades accuracy, TTM preserves temporal coherence and achieves a better trade-off between efficiency and performance. Furthermore, comparison with semantic-aware temporal accumulation (STA) demonstrates that TTM maintains higher accuracy despite slightly lower computational reduction. These results indicate that TTM is effective for efficient and reliable video recognition.
-
Zhenling Su, Qi Li, Song Wang, and Lin Meng, "AI model design and the challenging of CPU acceleration by ChatGPT," [Download]
-
Hayata KANEKO and Lin MENG, "An Embedded Vision Transformer with Quantization-Friendly Simple Attention, Optimized Buffers, and SIMD Acceleration," [Download] SCI/SCIE IF:2.7 JCR Q1 CAS Q3
-
Hayata KANEKO, Ryuto ISHIBASHI, and Lin MENG, "SIMD-CP: SIMD with Redundant Bits Compression and Mixed-Precision Packing for Quantized DNNs," [Download] SCI/SCIE IF:2.8 JCR Q2 CAS Q3
-
Yexin Zhang, Yuhao Yan, Zhenling Su, Takehito Nakamura, Daiki Nishimura, Qi Li, and Lin MENG, "An Engineering-Oriented Data-to-Deployment Framework for Reliable Surface Defect Detection of Annular Industrial Parts," [Download] SCI/SCIE IF:9.0 CAS Q1 JCR Q1Surface defect detection Deep learning Data leakage mitigation Industrial inspection Annular industrial parts概要を表示
Reliable surface defect detection of annular industrial parts remains challenging in practical manufacturing environments due to curved geometries, reflective metallic surfaces, and severe data redundancy caused by continuous rotational scanning. Since adjacent frames contain highly overlapping visual information, conventional random splitting may cause rotational data leakage and overestimated performance. To address these issues, we propose an engineering-oriented data-to-deployment framework that applies lightweight deep learning to automated surface defect inspection. The framework integrates physics-guided data engineering, leak-aware evaluation, geometry-constrained annotation, lightweight detector design, and practical system deployment. First, a physics-based Sector Split strategy reduces the risk of rotational data leakage and aligns evaluation with the physical acquisition process. A geometry-constrained annotation tool further improves labeling consistency across consecutive rotational frames. Second, a Spatial-background suppression, Group-shuffle fusion, and Strip-attention enhanced You Only Look Once detector (YOLO-SGS) is developed for defect localization on reflective metallic surfaces. YOLO-SGS achieves 84.6% mean average precision at 50% intersection over union (mAP@50) with a model-only inference latency of 6.73 ms and 23.3% fewer parameters than the baseline detector. Five-fold physical-part-level cross-validation yields
81.2
±
1.9
%
mAP@50. On three unseen physical parts, the Sector Split model achieves 81.9% mAP@50, compared with 69.0% for the Random Split model. Finally, the AnnularInspector software system is implemented using an asynchronous multi-threaded architecture and achieves a mean end-to-end latency of 494.3 ms with no observed frame loss. External validation and industrial-interference tests further support cross-part generalization within the investigated setting and practical applicability under representative industrial disturbances. -
Xiangheng WANG, Hengyi LI, and Lin MENG, "A Lightweight Web Platform for Real-Time Kuzushiji Recognition via One-Shot Pruning," [Download] SCI/SCIE IF:2.6 JCR Q1
-
Long Zhang, Yu Chen, Xueqing Shi, Taifeng Zhou, Jie Han, Lin Meng, Na Li, Matthew B. Greenblatt, Zemin Ling, Peiqiang Su, Fuxin Wei, and Ren Xu, "Discovery of a peripheral Myh11-expressing nucleus pulposus cell population demontrating therapeutic potential for disc degeneration," [Download] SCI/SCIE IF:18.1 CAS Q1 JCR Q1
-
Xiangheng Wang, and Lin Meng, "A Novel Spatiotemporal Database for Recognition and Spatial Analysis of Chinese Inscriptional Rubbings," [Download]
-
Xin Wang,Jiale Ren,Jiasheng Yang,XinZhe Yue,Qiu Xie, Jie Han, Ren Xu,Lin Meng,Jie Han, "Applications of Single-Cell and Spatial Transcriptomics in Osteoarthritis,"
-
Jiaze CAI, Bang LI, Hengyi LI, and Lin MENG, "A Review of Artificial Intelligence Techniques in Oracle Bone Inscriptions," [Download] SCI/SCIE IF:3.1 CAS Q3 JCR Q1
-
Hengyi LI, Qibo XU, Aihui WANG, and Lin MENG, "Exploring Prediction Confidence Capability of Deep Learning Architectures," SCI/SCIE IF:1.21
International Conference
-
Haruhiro TAKAHASHI, Ryuto ISHIBASHI, and Lin MENG, "Aware Spatial Recalibration for Training-Free Image Classification,"
-
Haruhiro TAKAHASHI, Ryuto ISHIBASHI, and Lin MENG, "Refining Vision Transformers via Attention Map Enhancement with Two-Branch Supervision,"
-
Shengda Gao, Zhenling Su, Mingcong Deng, and Lin Meng, "A Lightweight Visual Detection Framework for Cherry Tomato Grasping with a Robotic Manipulator,"
-
Yifan XU, Mengtao WANG, Zhizhi ZHOU, and Lin MENG, "Quantum Evolutionary Algorithm-Based Feature Optimization for Breast Cancer Detection Based on Quantum-Inspired Evolutionary Algorithms,"
-
Mikito SAITO and Lin MENG, "A Multimodal-Based Framework for the Reorganization of Ancient Japanese Manuscripts,"
-
Hengyi LI and Lin MENG, "Bayesian SNR-Driven Structured Pruning for Convolutional Neural Networks,"
-
Chaojie HUANG, Qi LI, and Lin MENG, "YOLO-Based Vehicle Detection for Night-time Highway Surveillance,"
-
Kento ICHIHARA and Lin MENG, "Physics-Guided Safety Gating: Balancing Structure Preservation and Raindrop Removal,"
-
Muhammad Hamza Mehdi and Lin MENG, "HighPhytoSparseNet: A Lightweight Multi-Head Object Detection Model for Edge-Based Agricultural Applications,"
-
XinZhe YUE and Lin MENG, "Weighted Bidirectional LSTM with Dynamic Attention for Logographic News Text Classification,"
-
Xinzhe Yue, Runqian Zhang, Zenghui Wang and Lin Meng, "Real-Time Visual Overflow Monitoring in Biogas Digesters via P2-Enhanced Instance Segmentation,"
-
Ryuto Tanigawa, Yingrui Geng, Hayata Kaneko, Ryuto Ishibashi, Qi Li, and Lin Meng, "PHOENIX: A Physics-Integrated Mixture-of-Experts Framework for High-Fidelity Battery SOH Estimation,"
-
Kento Ichihara, and Lin Meng, "From Frequency Filtering to Coherence Supervision for Raindrop Segmentation in Uncontrolled Environments,"
-
Zhen GONG, Yexin ZHANG, Yuhao YAN, Qi LI, and Lin MENG, "Automatic Deep Learning Model Generation for Medium-Sized Fruit Detection in Harvesting Application,"
-
Zhen GONG, Yexin ZHANG, Yuhao YAN, Qi LI, and Lin MENG, "A Review of UAV-Based Precision Pollination,"
-
Haruto Hashizume, Ryuto Ishibashi, Ryuto Tanigawa, and Lin Meng, "Label-Free Highway Traffic Anomaly Detection via Type-Specialized Isolation Forests,"
-
Ryuto Tanigawa, Yingrui Geng, Ryuto Ishibashi, and Lin Meng, "Rethinking Physics Equations in Physics-Informed Battery Forecasting with Batch- Level Statistical Validation,"
-
Mikito Saito, and Lin Meng, "WCIM: Local-Pair-Guided Reading Order Estimation for Warichu in Japanese Historical Documents,"
-
Sota Ono, and Lin Meng, "Session-Disjoint Evaluation of Species-Agnostic Plant Segmentation Methods in Field Images,"
-
Takumi Yamamoto and Lin Meng, "Automatic Deep Learning Model Generation for Medium-Sized Fruit Detection in Harvesting Application,"
-
Sota Ono, and Lin Meng, "Species-Agnostic Weed-Candidate Extraction by Differencing Plant and Crop Masks,"
-
Haruki Asakawa and Lin Meng, "Adaptive Sub-Slot Attention for Image Exposure Correction,"
-
Ryuichi Nozaki, and Lin Meng, "Feature Importance in Household Power Forecasting Changes with the Time Resolution of the Data,"
-
Kento Ichihara, and Lin Meng, "Spectral Entropy-Guided Evidential Learning for Real-Time Reliable Perception in Adverse Weather,"
-
Mikihisa Ishino, Takumi Yamamoto, Runqian Zhang, Shengda Gao, and Lin Meng, "Score-Based Harvesting Order Planning and False Detection Handling for Robotic Tomato Harvesting,"
-
Ren Doyama, Ryuto Ishibashi, and Lin Meng, "Systematic Evaluation of Object Detection Model Modifications for Small-Dataset Apple Surface Defect Detection ,"
-
Yuichi Imoto, and Lin Meng, "Mask-Guided Restoration of Damaged Kuzushiji: A Benchmark Study with Modern Architectures,"
-
Takumi Yamamoto, Ryuto Ishibashi, Mikihisa Ishino, Ryuto Tanigawa, and Lin Meng, "Object Detection Optimization for Small and Medium Targets via Neck Simplification and Channel Redistribution,"
-
Daisuke Fujita, Ryuto Ishibashi, Ryuto Tanigawa, Kohei Yamaguchi, and Lin Meng, "Sensitivity-Guided Mixed-Precision Post-Training Quantization for MambaVision ,"
-
Rui Jing, Lin Meng, and Mingcong Deng, "A Lightweight Strawberry Keypoint Detection Network for Edge Deployment in Harvesting Robots ,"
National Conference
-
Qi LI, 孟 林, "マルチスケール特徴融合を用いた軽量ステレオマッチング法,"
-
横山 拓哉, 孟 林, "推論レイテンシを考慮した軽量Vision TransformerのJetson Nanoへの実装と評価,"
-
佐々木 優希, 孟 林, "物体検出と分類モデルを融合した日本古典籍文字認識手法の実装と評価,"
-
楊 嘉晟, 孟 林, 謝 秋, 許 Ren, "Feature Dropoutを用いたPC利用時の行動状態の認識,"
-
王 超, 李 祺, 孟 林, "高分解能画像に対応したVision Transformer の計算効率化に関する研究,"
-
石橋 龍人, 孟 林, "画像認識モデルに対するJPEG符号化余剰の分析と量子化テーブルの解析的導出,"
-
金子 隼大, 孟 林, "組込みAI向け量子化手法の整理とFPGAにおける行列演算ユニットの効率化,"