Di Yang 杨迪

Ph.D. Researcher | Generative Models, World Models & Intelligent Agents
3D/4D Generation · Robotic Planning · Agent Learning · Visual Reasoning
🏛️ The College of William & Mary, Williamsburg, VA
👨‍🏫 Advisor: Dr. Yanhai Xiong
🤝 Also working with: Dr. Haipeng Chen, Dr. Yufei Wang
📢 Seeking Internship
Actively seeking Summer 2027 Research Internship opportunities in Generative Models, World Models, Embodied AI / Robotics, and Intelligent Agents. Feel free to reach out!

I am a Ph.D. student in Data Science at The College of William & Mary, supervised by Dr. Yanhai Xiong, and also work with Dr. Haipeng Chen. Currently, doing research with SparcAI Inc. (mentored by Dr. Yufei Wang). Previously, I received my M.Sc. in Electrical Engineering from National University of Singapore (NUS) (advised by Dr. John S. Ho) and B.Eng. in Microelectronic Science and Engineering from University of Electronic Science and Technology of China (UESTC). Prior to my Ph.D., I worked as a Digital Logic Engineer at Cambricon Technologies on large-scale AI accelerator hardware verification.

My research interests lie at the intersection of Generative Modeling, World Models, and Intelligent Agents. My research spans 3D/4D generation, structured world representations, and learning-based planning, with growing interests in agent learning, multimodal reasoning, and decision-making.

Di Yang

News & Updates

2026.10 📝 Published a new research blog post: "Does a Failed Agent’s Workspace Help the Next One? A Controlled Pilot on Failure-Conditioned Recovery Between Coding Agents", featuring full open-source code and data on GitHub!
2026.09 🎉🎉 Double acceptance! Both "VTV-FM" and "HyCO" have been Accepted by NeurIPS 2026!
2026.06 Awarded the Summer Research Fellowship at The College of William & Mary.
2026.05 Developed Channel Match for high-resolution 3D distillation on TRELLIS.2.
2026.02 🎉 Our paper "LENS: Learning to Navigate with Active Search for Partially Observable MAPF in Unknown Environments" was Accepted by Transactions on Machine Learning Research (TMLR)!
2026.01 Preprint on HyCO released on arXiv:2609.07990.
2025.08 🎉 Our paper "UP-Bench: A Benchmark for Underwater Path Planning Algorithms" was accepted for an Oral Presentation at KDD 2025 (Benchmark Track)!
2025.06 Awarded the Summer Research Fellowship at The College of William & Mary.
2024.09 Joined The College of William & Mary as a Ph.D. student in Data Science.

Research Interests

3D/4D Generative Modeling & World Models

Explicit 4D spatiotemporal representations (Sparc4D), few-step 3D distillation (Channel Match), continuous-time flow matching (VTV-FM), and physical world action models (WAM).

Robotic Planning & Embodied Intelligence

Learning-based spatial planning under partial observability (LENS, TMLR), realistic 3D simulation platforms (UP-Bench, KDD Oral), and closed-loop robot manipulation policies (WAM).

Agent Learning & Multimodal Reasoning

Investigating multi-agent coordination, failure-conditioned recovery in coding agents, and visual-spatial reasoning in multimodal foundation models for complex decision-making.

Selected Publications

Preprints & Under Review
Di Yang, Zhihao Li, Yanhai Xiong, Yufei Wang
arXiv:2610.01229, 2026 Preprint
A feed-forward autoencoder encoding monocular videos into a compact, sparse 4D scene state by sharing static features and compressing dynamic motion into spatially anchored temporal slots. Decodes directly into 2D Gaussian surfels, achieving competitive novel-view synthesis (21.51 dB on MultiCamVideo) at ~0.6M state size (~12× smaller than single-frame baselines) with zero-shot transfer to real handheld DyCheck videos.
Note: The [PDF] link above has been updated to our latest refined manuscript (featuring robot manipulation downstream results on StackCube). The updated version will be synced to arXiv soon!
@article{yang2026sparc4d, title={Sparc4D: A Compact Explicit 4D Representation for Dynamic Scenes}, author={Yang, Di and Li, Zhihao and Xiong, Yanhai and Wang, Yufei}, journal={arXiv preprint arXiv:2610.01229}, year={2026} }
Di Yang, Z. Li, Y. Xiong, Y. Wang
Still under review Under Review
Distills a native 3D generator into a two-step student via continuous self-consistency training with a custom JVP FlashAttention kernel. Channel Match calibrates sampling-induced latent distribution drift, achieving 4.5× faster geometry sampling on TRELLIS.2 (1024³ resolution) without texture artifacts.
@article{yang2026channelmatch, title={Channel Match: Calibrating Shape Latents for Few-Step 3D Generation}, author={Yang, Di and Li, Z. and Xiong, Y. and Wang, Y.}, journal={Under Review}, year={2026} }
Peer-Reviewed
Y. Li, Di Yang*, H. Chen, Y. Xiong (*Co-first author)
Annual Conference on Neural Information Processing Systems (NeurIPS), 2026 NeurIPS 2026
Designs an RL-to-diffusion handover using policy entropy and model disagreement to combine sequential construction with generative solution completion, reducing TSP-1000 optimality gap from 15.4% to 4.6%.
@inproceedings{li2026hyco, title={HyCO: A Hybrid Neural Solver for Combinatorial Optimization}, author={Li, Y. and Yang, Di and Chen, H. and Xiong, Y.}, booktitle={Advances in Neural Information Processing Systems (NeurIPS)}, year={2026} }
H. Jiang, Y. Li, Di Yang*, Y. Xiong, H. Chen, Y. He (*Co-first author)
Annual Conference on Neural Information Processing Systems (NeurIPS), 2026 NeurIPS 2026
Co-developed phase-space flow matching with cubic Hermite trajectories and a closed-form minimum-acceleration variational closure, achieving FID 2.99 on CIFAR-10 versus 3.45-3.73 for baselines.
@inproceedings{jiang2026vtvfm, title={VTV-FM: Flow Matching through Variational Terminal-Velocity Closure}, author={Jiang, H. and Li, Y. and Yang, Di and Xiong, Y. and Chen, H. and He, Y.}, booktitle={Advances in Neural Information Processing Systems (NeurIPS)}, year={2026} }
Di Yang et al.
Transactions on Machine Learning Research (TMLR), 2026 Accepted
A hybrid Multi-Agent Pathfinding architecture decoupling macro-topological spatial reasoning (via U-Nets) from micro-level collision avoidance, achieving generalization comparable to 85M foundation models with <0.2% training data.
@article{yang2026lens, title={LENS: Learning to Navigate with Active Search for Partially Observable MAPF in Unknown Environments}, author={Yang, Di and others}, journal={Transactions on Machine Learning Research}, year={2026} }
Di Yang, Yanhai Xiong
ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), 2025 Benchmark Track Oral
A configurable open-source 3D simulation benchmark platform incorporating realistic fluid dynamics, partial observability, and dynamic obstacles, benchmarking Deep RL (SAC) and classical planning algorithms.
@inproceedings{yang2025upbench, title={UP-Bench: A Benchmark for Underwater Path Planning Algorithms}, author={Yang, Di and Xiong, Yanhai}, booktitle={Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining}, year={2025} }

Selected Research & Industry Experience

Research Collaboration & Research Intern
SparcAI Inc. (Remote) • Advisor: Dr. Yufei Wang
  • VLA Policy Manipulation: Integrating robot-centric point maps into a 5B VLA policy for metric spatial manipulation; built OmniGibson/BEHAVIOR pipelines with failure-case data augmentation.
  • Channel Match (Few-Step 3D Distillation): Distilled a native 3D generator to 2 steps with continuous self-consistency (4.5× faster shape sampling); implemented a custom Triton JVP FlashAttention kernel.
  • Explicit 4D Representation (Sparc4D): Developed Sparc4D, a compact explicit 4D autoencoder for dynamic scenes compressing time-varying features into temporal slots with 2D Gaussian surfels (~12× smaller state than single-frame baselines). Successfully matched predictive planning performance on the StackCube benchmark (vs. Structured 4D Latent Predictive Model, Li et al.) using drastically fewer tokens (~1,170 tokens by decoupling persistent static context from sparse dynamic motion); ongoing work actively explores its downstream applications in robot manipulation and predictive planning.
  • Next-Gen World Models (Ongoing Exploration): Actively exploring the potential of extending dynamic 4D representations towards next-generation World Models and unified physical world state tokenizers.
Ph.D. Researcher
The College of William & Mary • Advisor: Dr. Yanhai Xiong
  • HyCO: Designed RL-to-diffusion handover for combinatorial optimization; reduced TSP-1000 optimality gap from 15.4% to 4.6% (NeurIPS 2026, arXiv:2609.07990).
  • LENS & UP-Bench: Proposed LENS (TMLR) for partially observable MAPF; designed UP-Bench (KDD 2025 Oral) 3D underwater simulation benchmark.
  • VTV-FM: Co-developed second-order phase-space flow matching with variational terminal-velocity closure (NeurIPS 2026).
Digital Logic Engineer
Cambricon Technologies Co., Ltd. (Beijing, China)
Built functional verification pipelines for large-scale AI accelerator chips; owned system-level performance verification for low-latency data flow.
Graduate Researcher
National University of Singapore (NUS) • Advisor: Dr. John S. Ho
Conducted graduate research in wireless micro-systems and intelligent embedded hardware.

Education

The College of William & Mary
Ph.D. Student in Data Science • Advisor: Dr. Yanhai Xiong
National University of Singapore (NUS)
M.Sc. in Electrical Engineering • Advisor: Dr. John S. Ho
University of Electronic Science and Technology of China (UESTC)
B.Eng. in Microelectronic Science and Engineering • Chengdu, China

Teaching & Academic Services

Teaching Experience
DATA 305: Problem Solving with Generative AI
Teaching Assistant • The College of William & Mary
DATA 303: Introduction to Data Visualization
Teaching Assistant • The College of William & Mary
Academic Service
Conference Reviewer
AAAI (2025, 2026)

Visitor Analytics & Global Reach

Tracking academic engagement and global visitor distribution across institutions and regions:

👀 Total Pageviews: ... 🌍 Unique Visitors: ...
🗺️ Interactive visitor map powered by MapMyVisitors (Click to view full analytics)