Research Projects & Industry Collaboration
Industry Collaboration & R&D
Huawei / HiSilicon Quantization-Based AI Model Acceleration for Edge Platforms
The collaboration addresses accuracy degradation, memory constraints and quantization tuning efficiency on resource-constrained edge platforms. It studies model conversion, preprocessing and postprocessing, inference implementation, and on-device accuracy and performance evaluation across vision and selected speech and text tasks. As project leader, I lead the algorithm and hardware co-design roadmap, organize core quantization and inference development, and drive on-device testing, optimization and delivery reviews. The collaboration is underway and aims to develop reusable edge inference and quantization optimization approaches, lower model adaptation costs on resource-constrained chips, and improve edge AI application development efficiency. Evaluation focuses on quantization bit width, accuracy loss relative to the original model, memory use, inference latency and postprocessing latency.
MindPIPE: Automated Large-Model Compression and Quantization
The collaboration addresses differing compression sensitivities in multimodal encoders, complex mixture-of-experts optimization and compression-induced changes to cross-modal alignment. It integrates quantization, pruning, distillation, recovery fine-tuning and automated evaluation, with deployment adaptation for GPU and Huawei Ascend environments. As project leader, I lead the technical roadmap, core compression algorithm design and R&D execution. I drive the integration of quantization, pruning, distillation, evaluation and hardware adaptation into an end-to-end compression framework. The work aims to provide compression and inference optimization solutions for multimodal large models, reduce repetitive tuning and testing, lower memory and computing requirements, and improve deployment efficiency across computing platforms. Evaluation focuses on compression ratio, inference speedup, peak memory and task-accuracy loss.
Huawei Automotive BU: Low-Bit Quantization of Multimodal Large Models
The collaboration focuses on low-bit quantization of automotive multimodal large models, addressing inference bandwidth and deployment efficiency through weight and activation quantization and calibration strategies, including PTQ and QAT in PEFT settings. As project leader, I lead the model quantization collaboration and coordinate technical development. Evaluation focuses on weight and activation bit widths, accuracy loss and inference bandwidth.
- April 2022 - April 2023
ByteDance Text-to-Image and Text-to-Video Model Acceleration, Compression and Quantization
The collaboration focuses on efficient inference for text-to-image and text-to-video generation, optimizing diffusion sampling, model structures and redundant computation to improve generation speed. As project leader, I lead the technical roadmap and core algorithm R&D in model compression, low-bit weight and activation quantization, and efficient generation inference. The work reduces computing and memory requirements while preserving image quality, prompt alignment and temporal consistency in video generation, supporting efficient deployment of generative models. Evaluation focuses on generation quality, inference latency, throughput, peak memory and low-bit quantization accuracy.
Alibaba / Tmall Model Compression and Acceleration
The collaboration studies compiler-optimization-driven model design for frequently invoked, latency-sensitive cloud services with complex deployment pipelines. It connects architecture selection with compilation, evaluation and deployment optimization to address the mismatch between theoretical computation cost and actual runtime efficiency. As project leader, I lead model compression and architecture search development, connect model evaluation with compiler and runtime constraints, and develop the key algorithms for cloud-side model optimization. The work provides optimization approaches and technical solutions for candidate evaluation and deployment, supporting cloud model systems and helping reduce repeated optimization during evaluation and rollout. Evaluation focuses on model size, compiled inference latency, candidate evaluation cost and task accuracy.
HONOR Smartphone Image Super-Resolution and Enhancement
The collaboration develops smartphone image super-resolution and enhancement algorithms under mobile NPU and ISP computing and storage constraints. It combines neural architecture search, dynamic-distribution pruning and hardware-aware compression to explore lightweight networks suited to smartphone execution. As project leader, I lead lightweight network, architecture search and compression development, addressing the trade-off between reconstruction quality and execution efficiency on mobile hardware. The work provides lightweight model solutions for smartphone imaging, supports real-time execution and deployment optimization, and helps balance computing resources, storage and response time. Evaluation focuses on computation, parameter count, image quality and mobile execution latency. Compared with representative contemporary methods, computation is reduced 12.5-fold and parameter count is reduced 24.8-fold.
vivo Model Compression for Moiré and Shadow Removal
The collaboration studies how to reduce model size, storage and computation for smartphone moiré and shadow removal while preserving image detail and restoration quality. It combines compression, architecture search and modeling of mobile hardware constraints. As project leader, I lead compression strategy and hardware-aware optimization development, addressing the key trade-offs between restoration quality, model size and mobile execution cost. The work provides technical approaches for mobile imaging optimization and improves the feasibility of model execution and engineering deployment in resource-constrained environments. Evaluation focuses on restoration quality, model size, memory use and on-device inference latency.
OPPO Mobile Vision Algorithms and Lightweight Deployment R&D
The work addresses mobile imaging under computing, storage and latency constraints through lightweight visual algorithms, model compression and hardware adaptation. It connects algorithm design with actual device execution conditions to improve practical runtime efficiency. As project leader, I lead model architecture optimization, compression and deployment development around mobile computing, memory and latency constraints. The work provides model optimization and deployment approaches for mobile hardware, supports the transition from algorithm validation to device adaptation, and advances lightweight mobile imaging development. Evaluation focuses on visual-task quality, model size, memory use and mobile inference latency.
Huawei Large-Model Compression: Automatic Combination of Multiple Compression Strategies
The collaboration develops automatic combinations of model compression strategies, including large-model weight sparsification and joint sparse quantization, to improve efficient inference and deployment. I contribute project management and original algorithm development. The project received Huawei's Product-Line-Level Outstanding Collaborative Project award; its results were applied to related products and enhanced Ascend product competitiveness. Evaluation focuses on model compression, inference speedup and accuracy retention. 10B-class LLMs achieve 8-fold compression and 4-fold acceleration; 70B-class LLMs achieve 10-fold compression and 4.5-fold acceleration.
Tencent Youtu Smart Security and Consumer Vision R&D
This work concerns smart security and consumer vision products, addressing age-related recognition challenges, large-scale image and video retrieval and localization, and tracking stability and computing cost in real-time interaction. As a core algorithm developer, I develop recognition, retrieval and tracking methods and connect efficient training, visual representations and fast retrieval with the requirements of practical vision systems. The work supports smart security and consumer vision applications and helps translate visual methods from research validation into business systems. Evaluation focuses on recognition and retrieval quality, retrieval latency, tracking stability and real-time execution cost.
DiDi Traffic-Scene Visual Perception and Fast Localization R&D
The work studies localization using compressed visual fingerprints for traffic-scene analysis and visual search. It combines fine-grained image matching through local descriptors with large-scale retrieval through global visual fingerprints, while controlling feature storage and search costs. As a core algorithm developer, I develop compressed visual fingerprints, scene-matching algorithms and fast localization methods, addressing retrieval efficiency and storage constraints. The work provides approaches for traffic-scene visual retrieval and localization, supporting scene understanding and map data updates with attention to the efficiency and accuracy of large-scale map information processing. Evaluation focuses on localization accuracy, retrieval latency, feature storage and large-scale matching efficiency.
Research Funding & Institutional Support
January 2026 - December 2029Neural Architecture Search and Compression Strategies for Efficient Deployment of Large Models
Research on automated architecture search and compression strategies for efficient large-model deployment.
Awarded 2022Key Technologies for Neural Architecture Search
Postdoctoral research on architecture evaluation, efficient search and task adaptability.
Awarded 2022Neural Architecture Search with Weak Supervision
Research on candidate architecture evaluation and selection with limited supervision.
Platform task participationPeng Cheng Cloud Brain Platform Construction and Urgent R&D Tasks
Participation in Peng Cheng Cloud Brain platform construction and urgent R&D tasks.
Research supportXiamen University Xiaomi Young Scholar Program
Support for research development and team building in automated model design and efficient large-model deployment.








