joedong.ai

Xu "Joe" Dong · 董旭

General models will win robotics. Building them is a systems problem.

I build the infrastructure physical AI runs on — data, evaluation, hardware, simulation. Cofounder & CTO of Chestnut Robotics (previously TetherIA); tech lead for the planning model behind XPeng's production end-to-end driving; planning foundation models at Waymo. Physicist by training, computer vision by trade. Loves thinking, writing, and blogging — author of Physical AI Deep Dives.

Building

The infrastructure a general robot model needs — and the hands to prove it on.

I'm building the layer under the model: capture that turns human demonstrations into robot skill, evaluation that says whether it worked, hands cheap enough to run it at scale, and simulation to close the loop. At Chestnut that is Aero Hand, Aero UMI and GAIN3D — and a robotic foundation model on top.

Selected publications

Physics → computer vision → autonomous driving → physical AI

Selected work from fields deep learning remade — medical-image reconstruction and radar-camera perception, and now the dexterous hands physical AI will learn on; the full record is on Google Scholar.

Aero Hand Open: A Simulation-Ready Tendon-Driven Hand for Dexterous Manipulation Learning

N. Wang, M. Yadav, J. Wulff, A. Rosenbaum, K. Chen, Y. Sharma, X. Dong*, Y. Tao*

ICRA submission · 2026 · Dexterous manipulation · * corresponding author

arXiv · project · code

Radar-Camera Fusion via Representation Learning in Autonomous Driving

X. Dong, B. Zhuang, Y. Mao, L. Liu

CVPR-W · 2021 · Autonomous driving

DOI · arXiv

Probabilistic Oriented Object Detection in Automotive Radar

X. Dong, P. Wang, P. Zhang, L. Liu

CVPR-W · 2020 · Autonomous driving

DOI · arXiv

Sinogram Interpolation for Sparse-View Micro-CT with Deep Learning Neural Network

X. Dong, S. Vekhande, G. Cao

SPIE MI · 2019 · Medical imaging

DOI · arXiv

A Sparse-View CT Reconstruction Method Based on Combination of DenseNet and Deconvolution

Z. Zhang, X. Liang, X. Dong, Y. Xie, G. Cao

IEEE TMI · 2018 · Medical imaging

DOI

See all publications on Google Scholar →

Writing — Physical AI Deep Dives

I think in public about how physical AI actually scales.

The most useful split in robot learning isn't VLA vs. world model — it's system thinking vs. model thinking.

About

A researcher who builds in production.

I trained as a physicist (USTC, 2010–14), then moved into computer vision for my PhD, working on medical image analysis (Virginia Tech, 2014–19). From 2019 to 2025 I carried vision into autonomous driving through three of the field's paradigm shifts — monocular 3D detection, learning-based planning, then end-to-end foundation models, the last shipped to production at XPeng. In 2025 I cofounded Chestnut Robotics to do the same for robots. Every move has followed one read — a field AI had just cracked open. I think robotics is next: a general model will run the physical world, and I'm building the infrastructure it needs — data, evaluation, hardware, simulation, and a platform that scales.

Focus — dexterous manipulation; physical-AI infrastructure: data, evaluation, hardware, simulation, a platform that scales.
Belief — physical AI follows the LLM path; scaling wins.
Method — system thinking over model thinking; goal-driven, data-first.