Inclusive computer vision · Human motion 包容性计算机视觉 · 人体运动
Heming Du 杜赫铭
Computer vision for every body, not a single standard body.
让计算机视觉理解每一种身体,而非一种“标准身体”。
I develop 2D and 3D perception that adapts to people with limb differences. I also study transparent action reasoning for embodied AI.
我的研究面向肢体差异人群,探索能够适应不同身体结构的二维与三维感知,并研究具身智能中可解释的行动推理。
Postdoctoral Research Fellow · The University of Queensland
Research Scientist · FOME.ai · Brisbane, Australia
博士后研究员 · 昆士兰大学
研究科学家 · FOME.ai · 澳大利亚布里斯班
1297 citations · h-index 18 (Google Scholar, updated weekly) 引用 1297 次 · h 指数 18 (Google Scholar,每周自动更新)
01 · Research agenda01 · 研究议程
Beyond the standard body超越“标准身体”
Most vision systems inherit a fixed idea of human anatomy. My current work asks a different question: what if body structure is inferred for each person instead of imposed by a template?
大多数视觉系统继承了一套固定的人体结构假设。我的研究提出另一个问题:能否为每个人推断身体结构,而不是把统一模板强加给所有人?
- 01 · Published已发表 2D pose二维姿态 Residual limb keypoints残肢关键点
- 02 · Published已发表 Video视频理解 Motion and consistency运动与一致性
- 03 · Ongoing进行中 Dense parsing稠密解析 Regions informed by anatomy基于解剖结构的区域划分
- 04 · Published已发表 3D mesh三维网格 Adaptive topology自适应拓扑
- 05 · Next下一步 Fair multimodal reasoning公平的多模态推理 Representation and reasoning表征与推理
Current focus当前重点
Inclusive human-centric vision包容性人体感知
Pose, video understanding, dense parsing, and 3D reconstruction that represent morphological diversity explicitly.
在姿态、视频理解、稠密解析和三维重建中显式建模人体形态多样性。
Research foundation研究基础
Embodied AI具身智能
From relation graphs and visual transformers to transparent action reasoning with language models.
从关系图与视觉 Transformer,延伸到基于大语言模型的可解释行动推理。
Research translation成果转化
Sport and Paralympic technology体育与残奥科技
Pose-based athlete talent identification and video analysis that supports Paralympic classification.
将姿态估计用于运动员人才识别,并为残奥分级提供视频分析、数据集与模型支持。
Earlier foundations早期研究基础
A continuous path from navigation to inclusive perception 从视觉导航到包容性感知的连续研究路径
- 2020 · ECCV Object relation graphs ↗
- 2021 · ICLR Visual Transformer Network ↗
- 2023 · CVPR History-aware object-goal navigation (HiNL) ↗
02 · Selected work02 · 代表性研究
Four works, one line of inquiry一条主线,四项工作
These four works, newest first, share one line of inquiry: a position paper that frames the agenda, and the benchmarks, metrics, and adaptive 3D reconstruction that ground it.
四项工作按时间倒序排列,属于同一条研究主线:一篇提出研究议程的立场论文,以及支撑它的数据集、评估指标与拓扑自适应三维重建。
A Position on Topological Generalization in Human-Centric Vision
My role: Joint first author我的贡献:共同第一作者
A research agenda that moves beyond fixed skeletal templates and treats morphological diversity as a central generalization problem.
提出超越固定骨架模板的研究议程,将形态多样性视为视觉模型泛化能力的核心问题。
LDPose: Inclusive Human Pose Estimation in the Wild
My role: Joint first author我的贡献:共同第一作者
The first unconstrained pose estimation benchmark for people with limb differences, with residual limb endpoints, dedicated losses, and evaluation metrics.
首个面向肢体差异人群的非受限场景姿态估计基准,引入残肢端点、专用损失函数与评估指标。
03 · From research to impact03 · 从研究到影响
Research that broadens who technology can serve. 让技术服务于更广泛的人群。
Athlete talent identification运动员人才发掘
Athlete talent identification using pose estimation to support the pathway toward the Brisbane 2032 Olympic and Paralympic Games.
基于姿态估计的运动员人才识别,为布里斯班 2032 奥运会与残奥会人才路径提供技术支持。
UQ researcher profile昆士兰大学研究主页 ↗Paralympic classification残奥分级
AI video analysis, datasets, and pose models for Paralympic classification, developed with classification experts.
与残奥分级专家合作,研发 AI 视频分析、专用数据集与姿态估计模型。
Project details项目详情 ↗04 · News04 · 动态
Recent news近期动态
| Jun 01, 2026 | Our position paper on topological generalization was accepted to ICML 2026. 🎉拓扑泛化立场论文被 ICML 2026 接收。🎉 |
|---|---|
| May 15, 2026 | ResiHMR, a method for adaptive 3D human mesh recovery for people with limb differences, was accepted to CVPR 2026.ResiHMR(残肢感知三维人体重建)被 CVPR 2026 接收。 |
| Mar 20, 2026 | InclusiveVidPose, a video pose estimation benchmark for people with limb differences, was accepted to ICLR 2026.面向肢体差异人群的视频姿态估计基准 InclusiveVidPose 被 ICLR 2026 接收。 |
| Oct 19, 2025 | LDPose, the first in-the-wild pose estimation benchmark for people with limb differences, was presented at ICCV 2025.首个面向肢体差异人群的姿态估计基准 LDPose 在 ICCV 2025 发表。 |
05 · Collaborate05 · 合作
Let’s build vision systems that recognize diverse bodies and motion. 共同构建能够理解多样身体与运动的视觉系统。
I welcome research collaboration and conversations with prospective PhD candidates working on inclusive vision, human motion, or embodied AI.
欢迎围绕包容性视觉、人体运动与具身智能开展研究合作,也欢迎有意攻读博士的同学联系交流。