中文主页

👋 Profile

I am Guanzhou Ke (柯冠舟), an Embodied AI Researcher at Avant Robotics in Shenzhen. I study active embodied intelligence under partial observability, with a focus on active perception, world models, and self-improving agents. My current work is grounded in UAV autonomy and inspection.

My research asks how embodied agents can actively acquire task-relevant evidence, learn from difficult simulated and real-world scenarios, and improve through an evaluation–data–training loop. I received my Ph.D. from Beijing Jiaotong University and was a CSC visiting Ph.D. researcher at Singapore Management University. My earlier work on multi-view representation learning and missing-modality completion provides the foundation for reasoning and action under incomplete observations.

Links: Google Scholar · ORCID · GitHub · ResearchGate · Email

🎯 Research Agenda

  1. Active Perception for UAV Inspection — Ongoing research. How should a drone coordinate wide-field cameras, a high-resolution gimbal, and body motion to decide what, where, and when to observe under viewpoint, bandwidth, latency, and safety constraints?
  2. Self-evolving Simulation and Data Engines. Build evaluation-driven loops that identify failure modes, generate targeted interaction data, and improve navigation and action models.
  3. Reliable Multimodal Intelligence under Partial Observability. Learn and reason when observations are incomplete, missing, uncertain, or viewpoint-dependent, connecting prior multimodal work with embodied reasoning and action.

🚁 Current Systems and Research

Simulation-driven embodied learning loop

As an Embodied AI Researcher, I work on the simulation, evaluation, and data-engine pipeline that connects hard-case discovery, targeted interaction collection, and model improvement across VLN/VLA, object search, exploration, navigation, and obstacle avoidance.

The resulting team system produces 39 million valid simulated interaction steps per month, improves sampling throughput by 3× on a single RTX 5090, and supports a 0.8B world/action model reporting 74% navigation-and-avoidance success. These figures describe the integrated team system rather than an individual result.

Real-world-grounded scenario construction

I work on geospatially aligned scenario construction using real-world coordinates and aerial observations, emphasizing task executability, repeatability, evaluation, and the return of failure cases into the data loop. Specific scene sources and named locations remain private.

Active evidence acquisition for UAV inspection

This is an ongoing research agenda, not a completed benchmark or public system. The goal is to jointly reason over viewpoint, sensing resolution, gimbal control, body motion, latency, bandwidth, and safety, so a UAV can determine what evidence is missing and when enough evidence has been gathered.

📣 News

  • [05/2026] One paper accepted at IEEE T-PAMI.
  • [05/2026] Recognized as an ICML 2026 Gold Reviewer.
  • [05/2026] One paper accepted at ICML 2026.
  • [03/2026] One paper accepted to the CVPR 2026 Findings track.
  • [06/2025] One paper accepted at ICCV 2025.
  • [02/2025] Knowledge Bridger accepted at CVPR 2025.
  • [12/2024] Two papers accepted at AAAI 2025.
  • [10/2024] Started a one-year CSC visiting Ph.D. appointment at Singapore Management University.
  • [02/2024] Joined Microsoft Research Asia as a research intern.
  • [02/2024] MRDD accepted at CVPR 2024.
  • [10/2023] One paper accepted in Information Fusion.
  • [07/2023] DMRIB accepted at ACM MM 2023.
  • [10/2022] One paper accepted at the ICDM 2022 Workshop.
  • [09/2022] Started Ph.D. studies at Beijing Jiaotong University.
  • [12/2021] One paper accepted at IEEE BigData 2021.

💼 Experience

  • 12/2025 – Present: Embodied AI Researcher
    • Avant Robotics, Shenzhen, China.
    • Active embodied intelligence, simulation-based evaluation, data engines, world/action models, and UAV autonomy.
    • Mentor: Zhenguo Li.
  • 02/2024 – 10/2024: Research Intern
    • Microsoft Research Asia, Shanghai AI/ML Group.
    • Multimodal medical report generation and hallucination mitigation.
    • Mentor: Xinyang Jiang.
  • 06/2023 – 12/2023: Research Intern
    • Institute of Automation, Chinese Academy of Sciences.
    • Multimodal deepfake detection across visual, textual, and audio signals.
    • Mentor: Bo Wang.

🎓 Education

  • Ph.D., Management Science and Engineering, Beijing Jiaotong University, 2022–2026.
  • CSC Visiting Ph.D. Researcher, Computer Science, Singapore Management University, 2024–2025. Advisor: Prof. Shengfeng He.
  • M.S., Systems Engineering, Wuyi University, 2019–2022. Outstanding Thesis Award.
  • B.Eng., Communication Engineering, Wuyi University, 2017–2019. Outstanding Graduate.

📄 Selected Publications

Diagram for How Far Are We from Generating Missing Modalities with Foundation Models?
How Far Are We from Generating Missing Modalities with Foundation Models?
Guanzhou Ke, Bo Wang, Guoqing Chao, Weiming Hu, Shengfeng He
IEEE Transactions on Pattern Analysis and Machine Intelligence (T-PAMI)
Rank: CCF A, MISC: [PDF] [CODE]
Diagram for Knowledge Bridger: Towards Training-free Missing Multi-modality Completion Venue banner for Knowledge Bridger: Towards Training-free Missing Multi-modality Completion
Knowledge Bridger: Towards Training-free Missing Multi-modality Completion
Guanzhou Ke, Shengfeng He, Xiao-Li Wang, Bo Wang, Guoqing Chao, Yuanyang Zhang, Xie Yi and HeXing Su
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
Rank: CCF A, MISC: [PDF] [CODE]
Diagram for Rethinking Multi-view Representation Learning via Distilled Disentangling
Rethinking Multi-view Representation Learning via Distilled Disentangling
Guanzhou Ke, Bo Wang, Xiaoli Wang, and Shengfeng He
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024
Rank: CCF A, MISC: [PDF] [CODE]
Diagram for Disentangling Multi-view Representations Beyond Inductive Bias
Disentangling Multi-view Representations Beyond Inductive Bias
Guanzhou Ke, Yang Yu, Guoqing Chao, Xiaoli Wang, Chenyang Xu, and Shengfeng He
The 31st ACM International Conference on Multimedia (ACM MM 2023)
Rank: CCF A, MISC: [PDF] [CODE]

📚 Full Publications

For the latest citation record, see Google Scholar.

2026

Diagram for How Far Are We from Generating Missing Modalities with Foundation Models?
How Far Are We from Generating Missing Modalities with Foundation Models?
Guanzhou Ke, Bo Wang, Guoqing Chao, Weiming Hu, Shengfeng He
IEEE Transactions on Pattern Analysis and Machine Intelligence (T-PAMI)
Rank: CCF A, MISC: [PDF] [CODE]
Diagram for Reliable Neighborhood-Aware Multi-View Outlier Detection Venue banner for Reliable Neighborhood-Aware Multi-View Outlier Detection
Reliable Neighborhood-Aware Multi-View Outlier Detection
Huijie Ma, Haoyuan Xin, Lei Meng, Guanzhou Ke, Yongyong Chen, Guoqing Chao
International Conference on Machine Learning (ICML), 2026
Rank: CCF A, MISC: [PDF] [CODE]
Diagram for OKGraph: Online Knowledge Graph Probing for Open-vocabulary Recognition Venue banner for OKGraph: Online Knowledge Graph Probing for Open-vocabulary Recognition
OKGraph: Online Knowledge Graph Probing for Open-vocabulary Recognition
Junhui Yin, Zhizhen Cai, Puze Wang, Guanzhou Ke, Jianhua Yang, Man Zhang, Qiang Zhang, and Shengfeng He
IEEE/CVF Conference on Computer Vision and Pattern Recognition Findings Track (CVPR Findings), 2026
Rank: CCF A, MISC: [PDF] [CODE]

2025

Diagram for LightBSR: Towards Lightweight Blind Super-Resolution via Discriminative Implicit Degradation Representation Learning Venue banner for LightBSR: Towards Lightweight Blind Super-Resolution via Discriminative Implicit Degradation Representation Learning
LightBSR: Towards Lightweight Blind Super-Resolution via Discriminative Implicit Degradation Representation Learning
Jiang Yuan, JI Ma, Bo Wang, Guanzhou Ke, Weiming Hu
International Conference on Computer Vision (ICCV), 2025
Rank: CCF A, MISC: [PDF] [CODE]
Diagram for Knowledge Bridger: Towards Training-free Missing Multi-modality Completion Venue banner for Knowledge Bridger: Towards Training-free Missing Multi-modality Completion
Knowledge Bridger: Towards Training-free Missing Multi-modality Completion
Guanzhou Ke, Shengfeng He, Xiao-Li Wang, Bo Wang, Guoqing Chao, Yuanyang Zhang, Xie Yi and HeXing Su
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025
Rank: CCF A, MISC: [PDF] [CODE]
Diagram for Global-Semantic Alignment Distillation for Partial Multi-view Classification Venue banner for Global-Semantic Alignment Distillation for Partial Multi-view Classification
Global-Semantic Alignment Distillation for Partial Multi-view Classification
Xiao-Li Wang, Anqi Huang, Yongli Wang, Guanzhou Ke, Xiaobin Hong, and Jun Liu
The 39th Annual AAAI Conference on Artificial Intelligence (AAAI)
Rank: CCF A, MISC: [PDF] [CODE]
Diagram for Incomplete Multi-view Clustering via Diffusion Contrastive Generation Venue banner for Incomplete Multi-view Clustering via Diffusion Contrastive Generation
Incomplete Multi-view Clustering via Diffusion Contrastive Generation
Yuanyang Zhang, Weiqing Yan, Yijie Lin, Li Yao, Xinhang Wan, Guangyuan Li, Chao Zhang, Guanzhou Ke, and Jie Xu
The 39th Annual AAAI Conference on Artificial Intelligence (AAAI)
Rank: CCF A, MISC: [PDF] [CODE]

2024

Diagram for Rethinking Multi-view Representation Learning via Distilled Disentangling
Rethinking Multi-view Representation Learning via Distilled Disentangling
Guanzhou Ke, Bo Wang, Xiaoli Wang, and Shengfeng He
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024
Rank: CCF A, MISC: [PDF] [CODE]

2023

Diagram for Knowledge distillation-driven semi-supervised multi-view classification
Knowledge distillation-driven semi-supervised multi-view classification
Xiaoli Wang, Yongli Wang, Guanzhou Ke, Yupeng Wang, and Xiaobin Hong
Information Fusion
Rank: SCI Q1, MISC: [PDF] [CODE]
Diagram for A Clustering-guided Contrastive Fusion for Multi-view Representation Learning
A Clustering-guided Contrastive Fusion for Multi-view Representation Learning
Guanzhou Ke, Guoqing Chao, Xiaoli Wang, Chenyang Xu, Yongqi Zhu, and Yang Yu
IEEE Transactions on Circuits and Systems for Video Technology (TCSVT)
Rank: CCF B, MISC: [PDF] [CODE]
Diagram for Disentangling Multi-view Representations Beyond Inductive Bias
Disentangling Multi-view Representations Beyond Inductive Bias
Guanzhou Ke, Yang Yu, Guoqing Chao, Xiaoli Wang, Chenyang Xu, and Shengfeng He
The 31st ACM International Conference on Multimedia (ACM MM 2023)
Rank: CCF A, MISC: [PDF] [CODE]

2022

Diagram for MORI-RAN: Multi-view Robust Representation Learning via Hybrid Contrastive Fusion
MORI-RAN: Multi-view Robust Representation Learning via Hybrid Contrastive Fusion
Guanzhou Ke, Yongqi Zhu, and Yang Yu
ICDM workshop
Rank: CCF B, MISC: [PDF] [CODE]
Diagram for Efficient Multi-view Clustering Networks
Efficient Multi-view Clustering Networks
Guanzhou Ke, Zhiyong Hong, Wenhua Yu, Xin Zhang, and Zeyi Liu
Applied Intelligence Springer
Rank: CCF C, MISC: [PDF] [CODE]

2021

Diagram for CONAN: Contrastive Fusion Networks for Multi-view Clustering
CONAN: Contrastive Fusion Networks for Multi-view Clustering
Guanzhou Ke, Zhiyong Hong, Zhiqiang Zeng, Zeyi Liu, Yangjie Sun, and Yannan Xie
IEEE International Conference on Big Data (Big Data)
Rank: CCF C, MISC: [PDF] [CODE]

🏆 Awards

  • Second Prize, “Huawei Cup” National Graduate Mathematical Modeling Competition, 2020, 2021, and 2022.
  • Second Prize, National Finals, Blue Bridge Cup Information Competition Group B, 2018.
  • National Scholarship of China, 2015.

🤝 Academic Service

  • Journals: IEEE Transactions on Multimedia, IEEE Transactions on Circuits and Systems for Video Technology, IEEE Transactions on Neural Networks and Learning Systems, Neurocomputing, and others.
  • Conferences: NeurIPS, CVPR, ICML (Gold Reviewer), AAAI, ACM Multimedia, and others.

📬 CV and Contact

The downloadable English and Chinese CV files are being refreshed to synchronize the August 2026 graduation status, current title, and research positioning. Until then, this homepage is the current public profile.

Email: guanzhouk@gmail.com · Google Scholar · ORCID · GitHub