Weihang Guo

I am a second-year Ph.D. student at Rice University, advised by Prof. Lydia Kavraki. I also have the pleasure of collaborating with Prof. Zak Kingston.

My research interests include high-performance robotic systems, agentic robotics for long-horizon planning, and video generation. I aim to build hierarchical robotic systems that combine multimodal agents for and with for execution.

I am a developer of the Open Motion Planning Library (OMPL) OMPL GitHub stars , a cornerstone of the motion planning community with nearly two decades of development.

Weihang Guo

Publications

* Equal contribution · † Mentoring

paper thumbnail
Foundation-Model-Guided Topology-Aware Semantic Risk Fields for Manipulation
Giung Lee, Weihang Guo†, Lydia E. Kavraki

Abstract: Robot motion planning in everyday environments must satisfy hard geometric constraints while accounting for context-dependent semantic risk. We present a foundation-model-guided, topology-aware semantic risk field that extends manipulation safety beyond collision avoidance. For each manipulated-object/scene-object pair, a foundation model provides six directional risk weights and a pair-specific spatial decay scale. The method combines these priors with voxelized 3D scene geometry using topology-aware shielding and geodesic spatial decay. A GPU-parallel backend batches object-level distance and risk computations to construct a dense 3D field that serves as a modular cost for downstream motion planning. We evaluate the field's shielding behavior under full and partial barriers and compare its 3D workspace representation with a pixel-wise semantic-prior baseline. Across three household simulation scenarios, trajectories optimized with the proposed field have lower semantic exposure than collision-only trajectories under the same geometric constraints. We also evaluate the computational practicality and reliability of the supporting pipeline. Together, these results support the proposed field as a practical topology-aware semantic cost representation for manipulation planning beyond collision avoidance.

paper thumbnail
Scaling Video Generation for Reasoning: At What Cost?

Abstract: We study whether scaling video generation enables models to reason about hidden information from the past frames, and at what computational cost. Our controlled benchmark requires predicting nine prescribed moves of an initially solved 2x2x2 Rubik's Cube from a fixed view of three faces. Correct predictions require inferring how actions change hidden states, and the simulator provides exact ground truth for evaluation. Models learn plausible cube geometry early, while correct sticker configurations require substantially more training. Although validation MSE follows approximate power-law scaling, lower MSE loss does not reliably indicate downstream reasoning capabilities. Smaller autoregressive models achieve higher state accuracy with limited compute, while larger models reach higher accuracy after more training. At roughly 0.1 PF-days, the 70M-parameter model correctly predicts the visible sticker configuration in 44.6% of post-action frames, compared with 0.3% for the 1B model, which reaches 83.7% at 3.14 PF-days. Symbolic state supervision raises the 20M model's frame accuracy from 31.1% to 67.3% at the same training-data budget, suggesting that learning representations of state changes can complement scaling.

paper thumbnail
Compress to Remember: Learning Compact Memory via On-Policy Distillation for Long Video Generation

Abstract: Standard video generators do not natively compact historical context into reusable memory tokens. As generation continues, the growing history makes it increasingly difficult to retain information from earlier frames due to long-context degradation. Key-frame-based approaches address this challenge by retaining selected past frames, but can discard information needed for future generation. Rather than relying on frame selection alone, we study whether a frozen video generator can supply the supervision needed to learn a compact representation of the history. We propose Prediction-Aligned Context Compaction (PACC), which uses a learned compressor to aggregate information across past frames into compact memory tokens. We train the compressor through on-policy distillation, using the same frozen generator both as a student when conditioned on compressed memory and as a teacher when conditioned on the full history. The student generates continuations, while the teacher provides targets for the same noisy inputs at each denoising step. Only the compressor is updated to align the student's predictions with these targets. We evaluate PACC on MBench, which jointly measures memory-event coverage and consistency. PACC outperforms the strongest baseline by 6.63 points on Causal-rCM and 3.19 points on Causal Forcing. Evaluation on VBench-Long using MovieGen prompts further shows that PACC produces minute-long videos with generation quality competitive with baselines. Together, these results show that learning to compact historical context can improve long-video memory without modifying the underlying generator.

paper thumbnail
Zero-Shot Reactive Obstacle Avoidance for Generative Robot Policies
Weihang Guo, Lydia E. Kavraki

Abstract: We propose NUDGE (Nudge Update via Differentiable GEometry), a training-free obstacle-avoidance procedure that can be incorporated in any robot policy based on diffusion or flow matching, including diffusion policies and vision-language-action models. Our work injects gradients from a signed distance field, a function returning each point's distance to the nearest obstacle, into the policy at inference time to steer it away from obstacles. It supports any common action parameterization, from absolute or relative joint poses to end-effector poses, through a differentiable joint-trajectory decoder. Experiments show that NUDGE preserves the policy's task distribution and runs reactively in real time.

paper thumbnail
Using VLM Reasoning to Constrain Task and Motion Planning

Abstract: In task and motion planning, high-level task planning is done over an abstraction of the world to enable efficient search in long-horizon robotics problems. However, the feasibility of these task-level plans relies on the downward refinability of the abstraction into continuous motion. When a domain's refinability is poor, task-level plans that appear valid may ultimately fail during motion planning, requiring replanning and resulting in slower overall performance. Prior works mitigate this by encoding refinement issues as constraints to prune infeasible task plans. However, these approaches only add constraints upon refinement failure, expending significant search effort on infeasible branches. We propose VIZ-COAST, a method of leveraging the common-sense spatial reasoning of large pretrained Vision-Language Models to identify issues with downward refinement a priori, bypassing the need to fix these failures during planning. Experiments on two challenging TAMP domains show that our approach is able to extract plausible constraints from images and domain descriptions, drastically reducing planning times and, in some cases, eliminating downward refinement failures altogether, generalizing to a diverse range of instances from the broader domain.

@misc{yan2025using,
  title={Using VLM Reasoning to Constrain Task and Motion Planning}, 
  author={Muyang Yan and Miras Mengdibayev and Ardon Floros and Weihang Guo and Lydia E. Kavraki and Zachary Kingston},
  year={2025},
  eprint={2510.25548},
  archivePrefix={arXiv},
  primaryClass={cs.RO},
  url={https://arxiv.org/abs/2510.25548}, 
}
paper thumbnail
The Open Motion Planning Library 2.0
Weihang Guo, Theodoros Tyrovouzis, Emiliano Flores, Clayton W. Ramsey, Zachary Kingston, Ioan A. Şucan, Mark Moll, Lydia E. Kavraki

Abstract: The Open Motion Planning Library (OMPL), first released in 2008, has become a cornerstone of the motion planning community, providing implementations of a wide range of state-of-the-art sampling-based algorithms. Over almost two decades of continuous development, we have steadily expanded the library with new planners, state spaces, and problem formulations. These additions range from asymptotically optimal and lazy planners to constrained motion planning and planning with temporal-logic goals. Building on this foundation, we introduce OMPL 2.0, a major evolution of the library that targets real-time motion planning through hardware acceleration and integrates seamlessly with modern AI research workflows. We also reflect on how OMPL and the field of motion planning have grown together over the years, and discuss the library's broader impact on the research community.

@misc{guo2026openmotionplanninglibrary,
  title={The Open Motion Planning Library 2.0},
  author={Weihang Guo and Theodoros Tyrovouzis and Emiliano Flores and Clayton W. Ramsey and Zachary K. Kingston and Ioan A. Şucan and Mark Moll and Lydia E. Kavraki},
  year={2026},
  eprint={2605.29301},
  archivePrefix={arXiv},
  primaryClass={cs.RO},
  url={https://arxiv.org/abs/2605.29301},
}
paper thumbnail
Python Bindings for a Large C++ Robotics Library: The Case of OMPL

Abstract: Python bindings are a critical bridge between high-performance C++ libraries and the flexibility of Python, enabling rapid prototyping, reproducible experiments, and integration with simulation and learning frameworks in robotics research. Yet, generating bindings for large codebases is a tedious process that creates a heavy burden for a small group of maintainers. In this work, we investigate the use of Large Language Models (LLMs) to assist in generating nanobind wrappers, with human experts kept in the loop. Our workflow mirrors the structure of the C++ codebase, scaffolds empty wrapper files, and employs LLMs to fill in binding definitions. Experts then review and refine the generated code to ensure correctness, compatibility, and performance. Through a case study on a large C++ motion planning library, we document common failure modes, including mismanaging shared pointers, overloads, and trampolines, and show how in-context examples and careful prompt design improve reliability. Experiments demonstrate that the resulting bindings achieve runtime performance comparable to legacy solutions. Beyond this case study, our results provide general lessons for applying LLMs to binding generation in large-scale C++ projects.

@inproceedings{guo2026ompl-python,
    Author = {Weihang Guo and Theodoros Tyrovouzis and Lydia E. Kavraki},
    Title = {Python Bindings for a Large C++ Robotics Library: The Case of OMPL},
    Booktitle = {IEEE International Conference on Robotics and Automation (ICRA)},
    Year = {2026}
}
paper thumbnail
Efficient Multi-Robot Motion Planning for Manifold-Constrained Manipulators by Randomized Scheduling and Informed Path Generation

Abstract: Multi-robot motion planning for high degree-of-freedom manipulators in shared, constrained, and narrow spaces is a complex problem and essential for many scenarios such as construction, surgery, and more. Traditional coupled methods plan directly in the composite configuration space, which scales poorly; decoupled methods, on the other hand, plan separately for each robot but lack completeness. Hybrid methods that obtain paths from individual robots together require the enumeration of many paths before they can find valid composite solutions. This paper introduces Scheduling to Avoid Collisions (StAC), a hybrid approach that more effectively composes paths from individual robots by scheduling (adding stops and coordination motion along all paths) and generates paths that are likely to be feasible by using bidirectional feedback between the scheduler and motion planner for informed sampling. StAC uses 10 to 100 times fewer paths from the low-level planner than state-of-the-art hybrid baselines on challenging problems in manipulator cases.

@ARTICLE{guo2024efficient,
  author={Guo, Weihang and Kingston, Zachary and Hang, Kaiyu and Kavraki, Lydia E.},
  journal={IEEE Robotics and Automation Letters}, 
  title={Efficient Multi-Robot Motion Planning for Manifold-Constrained Manipulators by Randomized Scheduling and Informed Path Generation}, 
  year={2026},
  volume={11},
  number={4},
  pages={4385-4392},
  keywords={Robot kinematics;Collision avoidance;Schedules;Manipulators;Manifolds;System recovery;Multi-robot systems;End effectors;Trajectory;Timing;Multi-robot systems;constrained motion planning;motion and path planning;collision avoidance},
  doi={10.1109/LRA.2026.3662639}
}
paper thumbnail
CaStL: Constraints as Specifications Through LLM Translation for Long-Horizon Task and Motion Planning

Abstract: Large Language Models (LLMs) have demonstrated remarkable ability in long-horizon Task and Motion Planning (TAMP) by translating clear and straightforward natural language problems into formal specifications such as the Planning Domain Definition Language (PDDL). However, real-world problems are often ambiguous and involve many complex constraints. In this paper, we introduce Constraints as Specifications through LLMs (CaStL), a framework that identifies constraints such as goal conditions, action ordering, and action blocking from natural language in multiple stages. CaStL translates these constraints into PDDL and Python scripts, which are solved using an custom PDDL solver. Tested across three PDDL domains, CaStL significantly improves constraint handling and planning success rates from natural language specification in complex scenarios.

@INPROCEEDINGS{guo2025castl,
  author={Guo, Weihang and Kingston, Zachary and Kavraki, Lydia E.},
  booktitle={2025 IEEE International Conference on Robotics and Automation (ICRA)}, 
  title={CaStL: Constraints as Specifications Through LLM Translation for Long-Horizon Task and Motion Planning}, 
  year={2025},
  volume={},
  number={},
  pages={11957-11964},
  keywords={Constraint handling;Translation;Uncertainty;Large language models;Planning;Formal specifications;Robotics and automation},
  doi={10.1109/ICRA55743.2025.11127555}
}

Honors and Awards

Invited Talks

Organizing

Thank you to all my co-organizers for making this happen!

Reviewing