Qirui Wu

qiruiw at sfu dot ca

self.jpeg

I am a Ph.D student in 3DLG Lab at Simon Fraser University, supervised by Prof. Angel X. Chang. I also closely collaborate with Prof. Manolis Savva and Prof. Daniel Ritchie. My research aims at designing systems and developing algorithms for understanding and generating interactable and dynamic 3D world from realistic noisy signals. More specifically, I’m interested in

  • 4D foundation world models that understand dynamic events, possess persistent memory and evolve alongside streaming visual content.
  • Compositional, interactable and efficient 3D content generation ranging from object parts to large scenes.
  • 3D vision applications in edge mobile devices, robotics and human-centric AI.

Google Scholar / Twitter / Github / LinkedIn

Selected Research

2026

  1. gator.png
    GATOR: Generative and Agentic 3D Object Reconstruction From Casual Images
    arXiv preprint arXiv:2610.11215, 2026
  2. jrm.png
    JRM: Joint Reconstruction Model for Multiple Objects without Alignment
    In Proc. of the Conference on Computer Vision and Pattern Recognition (CVPR), 2026
  3. shaper.png
    ShapeR: Robust Conditional 3D Shape Generation from Casual Captures
    In Proc. of the Conference on Computer Vision and Pattern Recognition (CVPR), 2026
  4. artiverse.png
    Artiverse: A Diverse and Physically Grounded Dataset for Articulated Objects
    In Proc. of the Conference on Computer Vision and Pattern Recognition (CVPR), 2026

2025

  1. diorama.png
    Diorama: Unleashing Zero-shot Single-view 3D Scene Modeling
    In Proc. of the International Conference on Computer Vision (ICCV), 2025
    Highlight

2024

  1. r3ds.png
    R3DS: Reality-linked 3D Scenes for Panoramic Scene Understanding
    In Proc. of the European Conference on Computer Vision (ECCV), 2024
  2. generalizing_shape_retrieval.png
    Generalizing Single-View 3D Shape Retrieval to Occlusions and Unseen Objects
    In Proc. of the International Conference on 3D Vision (3DV), 2024

2022

  1. d3net.png
    D^3Net: A Unified Speaker-Listener Architecture for 3D Dense Captioning and Visual Grounding
    In Proc. of the European Conference on Computer Vision (ECCV), 2022

2021

  1. plan2scene.png
    Plan2Scene: Converting Floorplans to 3D Scenes
    In Proc. of the Conference on Computer Vision and Pattern Recognition (CVPR), 2021