Latest 3D Vision & NeRFs Research Papers
The newest 3D Vision & NeRFs papers from across the field — arXiv, NeurIPS, CVPR, Nature, and more — refreshed daily and ranked by relevance. Distill AI tracks 3D Vision & NeRFs so you don’t have to: get the standout work delivered to your inbox every morning, with 2-sentence summaries and the option to chat with any paper.
Get the latest 3D Vision & NeRFs papers in your inbox — free →Recent papers
- Inter-Shot Motion Correction of Segmented 3D-GRASE ASL Perfusion Imaging With Self-Navigation and CAIPI.Minhao Hu, Frederik J Lange, Peter Jezzard, Joseph G. Woods et al. · Open Access CRIS of the Uni... · Oct 1, 2026
Purpose Segmented 3D Gradient and Spin Echo (GRASE) is commonly used in Arterial Spin Labeling (ASL) perfusion imaging. However, it is vulnerable to inter-shot motion, leading to subtraction errors that cannot be corrected. We developed a r…
- 3D Point Splatting for mmWave Radar Novel View SynthesisAdnan Armouti, Yixuan Gao, Rajalakshmi Nandakumar · arXiv · Sep 10, 2026
Solving novel view synthesis (NVS) for millimeter-wave (mmWave) radar requires a renderer that is physically faithful, complex-valued, and multi-viewpoint-tractable. No prior method achieves these three properties simultaneously. Differenti…
- Guided Super-Resolution of Digital Elevation Models with Diffusion-Based Image GeneratorsArmand Mihai Nicolicioiu, Dominik Narnhofer, Nando Metzger, Daniel Panangian et al. · arXiv · Sep 10, 2026
High-resolution digital surface models (DSMs) play an important role in urban analysis, 3D building reconstruction, and infrastructure monitoring, yet their availability remains limited due to the high cost and complexity of data acquisitio…
- Revisiting Avatar-As-Image: High-Fidelity Registration is All You NeedMargaret Kostyrko, Yuxuan Xue, Garvita Tiwari, Gerard Pons-Moll · arXiv · Sep 10, 2026
The representation of 3D clothed humans as standardized 2D UV texture and displacement maps over an underlying body model has long been studied. This compact representation is enticing as it enables pretrained image networks to process, gen…
- Single-Stream Multi-Feature Fusion with Temporal Robustness for Gait Emotion RecognitionShirong Lyu, Silu Quan, Yixuan Ding, Chengpeng Wang · arXiv · Sep 10, 2026
3D skeleton-based gait emotion recognition faces high annotation costs, data scarcity, and poor generalization on heterogeneous data. This paper proposes SV-GCN, a single-stream multi-feature fusion framework with temporal invariance. We in…
- Learn the Solid, Not the File: Canonical Inputs for Neural Networks on CAD Boundary RepresentationsHeinrich Jiang, Hager Yasser Mohamed, Alexander Hitt, Valeriia Lomakina et al. · arXiv · Sep 10, 2026
Boundary representation (B-rep) is the standard format used by modern CAD systems for parametric 3D models. It turns out, the exact same solid can be represented by different B-reps: for example, two engineers using different operations, a …
- UBone3D: Physics-Rectified Conditional Flow Matching for Anatomical 3D Shape Completion from UltrasoundWeiying Chen, Yuchong Gao, Siyuan Li, Marek Reformat et al. · arXiv · Sep 10, 2026
Three-dimensional ultrasound (US) is a safe, radiation-free complementary modality to CT and X-rays for longitudinal monitoring, yet its segmentation-derived partial point clouds are extremely artifact-laden. Consequently, it is challenging…
- Recursive Code World Models: Building Complex Worlds through Recursive Scene ProgramsZhiqi Li, Yuxuan Liao, Bo Zhu · arXiv · Sep 10, 2026
Code world models represent worlds as executable programs, but this representation alone does not determine how to construct a complex world. We introduce Recursive Code World Models (RCWM), a framework for reconstructing complex 3D worlds …
- BridgeMatch: Conditional Transport Bridges in Matching Matrix Space for 3D Deformable RegistrationQianliang Wu, Haobo Jiang, Guangwei Gao, Shuo Chen et al. · arXiv · Sep 10, 2026
Reliable non-rigid point cloud correspondences are important for deformable anatomical registration, embodied perception and manipulation, and dynamic 3D reconstruction. Coarse-to-fine methods reduce computational cost by selecting the top-…
- Multi-Modal Controlled Coherent Motion GenerationYifei Liu, Qiong Cao, Hongwei Yi, Huaiguang Jiang et al. · arXiv · Sep 10, 2026
It is natural for humans to walk and talk simultaneously. This paper tackles the challenge of replicating such natural behaviors in 3D avatar motion generation driven by concurrent multimodal inputs, such as a text description of a man walk…
- Hologram Representation via Quadratic Phase Gaussian SplattingHaolong Wang, Yicheng Zhan, Kaan Akşit, Simeng Qiu · arXiv · Sep 10, 2026
We introduce Complex-Valued Quadratic Phase Gaussian (CVQPG), a novel hologram representation method that replaces standard 2D Gaussian representations used in 2D Gaussian Splatting with 2D quadratic phase functions. CVQPG incorporates addi…
- R4Tun: LLM-guided adaptive segmental tunnel lining segmentation in point cloudsXinghui Tao, Zehao Ye, Guangming Wang, Jelena Ninić et al. · arXiv · Sep 10, 2026
Automated inspection of segmental tunnel linings requires adaptive segmentation from 3D point clouds, yet expert-tuned pipelines often degrade when tunnel conditions vary. This paper presents R4Tun, a large language model (LLM)-driven adapt…
- SAMV-DUSt3R: Instance-Centric 3D Scene Decoupling from Sparse Multi-ViewsLangxu Zhao, Zuan Gu, Yingdan Zhang, Pengfei Zhao et al. · arXiv · Sep 10, 2026
With the rising demand to decouple objects from 3D scenes, we propose SAMV-DUSt3R, an end-to-end model that injects SAM2 2D masks into MV-DUSt3R reconstruction. A Cross Flow Mask Block uses these masks to steer the network toward the target…
- Fast and Accurate Monomodal 3D High Resolution Deep Registration of Drosophila Larval Brain VolumesDaniel Reisenbüchler, Yousef Sadegheih, Michael Dittrich, Pratibha Kumari et al. · arXiv · Sep 10, 2026
The larval stage of Drosophila melanogaster is a compact model system for neuroscience whose genetic toolkit allows fluorescent markers to be expressed in defined neural populations, and comparing the resulting expression patterns across an…
- SCINTILLA-SNN: A Spiking Multi-Scale Selective Aggregation Network for Perineural Invasion PredictionYoungung Han, Yului Jeong, Kyeonghun Kim, Dohyun Kweon et al. · arXiv · Sep 10, 2026
Preoperative prediction of perineural invasion (PNI) in cholangiocarcinoma (CCA) is clinically valuable but remains challenging because PNI-related cues on magnetic resonance imaging (MRI) are subtle, sparse, and spatially localized around …
- Tri-DehazeGS: Scene--Medium Decoupled Gaussian Splatting with Transmittance-Aware OptimizationKui Jiang, Yang Gu, Jiacheng Liu, Shiyu Liu et al. · arXiv · Sep 10, 2026
Recovering clean 3D scenes from hazy multi-view images is challenging because haze attenuates scene radiance and introduces atmospheric scattering. Recent scattering-aware Gaussian Splatting methods introduce physical haze models into recon…
- Guiding Image-to-3D Generation with Test-Time Partial ObservationsJerred Chen, Simon Weber, Ronald Clark · arXiv · Sep 9, 2026
Image-to-3D models can generate visually compelling 3D assets from a single RGB image, but their geometry is often only loosely constrained by the available observations, limiting their use in applications that require geometric fidelity. I…
- Field Converter: Geometry-Initialized Temporal Residual Refinement for World-Grounded Player Pose Estimation from Soccer BroadcastsSimon Khan, Laurent Gajny, Jennyfer Lecompte, Sébastien Laporte · arXiv · Sep 9, 2026
Recovering 3D human pose from monocular sports broadcasts remains challenging when players must be localized in a shared metric world coordinate system rather than only reconstructed relative to their own body. We introduce Field Converter,…
- Shape-guided Gaussian Splatting for Sparse-View X-ray 3D ReconstructionPranav Poudel, Florence Dell'Aniello Picard, Nairouz Shehata, Frédéric Lavoie et al. · arXiv · Sep 9, 2026
Sparse-view X-ray 3D reconstruction is essential for reducing radiation exposure, but recovering a density field from a handful of X-ray projections is severely ill-posed. Recently, 3D Gaussian Splatting has achieved state-of-the-art perfor…
- SceneHI: High-Resolution 3D-Consistent Scene Texturing with Controllable IlluminationAthanasios Tragakis, Marco Aversa, Daniela Ivanova, Chaitanya Kaul et al. · arXiv · Sep 9, 2026
SceneHI is a framework that lifts high-resolution, illumination-aware priors from 2D diffusion models to perform 3D texture synthesis. It is the first to demonstrate that high-resolution textures, previously limited to 2D synthesis, can be …
- Geometry Without Coordinates: LiDAR Diffusion as a 3D Feature BridgeSamed Doğan, Nico Leuze, Alfred Schöttl · arXiv · Sep 9, 2026
Transferring the rich priors of large 2D foundation models to sparse 3D LiDAR remains challenging, as training native 3D foundation models at comparable scale is limited by data and annotation scarcity. We introduce a LiDAR-conditioned diff…
- View-Structured Conformal Prediction for 3D Gaussian SplattingJunzheng Chu, Bin Pan, Zhenwei Shi · arXiv · Sep 9, 2026
3D Gaussian Splatting (3DGS) renders novel views in real time, but an uncertainty heatmap does not certify that a rendered view meets a certain prediction coverage. We treat novel-view synthesis as structured regression and ask that, with p…
- LinearMask-GS: Stable-Mask Importance Pruning for Compact 3D Gaussian SplattingDonghun Ryu, Minhyeok Lee · arXiv · Sep 9, 2026
3D Gaussian Splatting (3DGS) enables real-time novel view synthesis but produces millions of primitives through adaptive densification, leading to significant storage overhead. Learned-mask pruning methods such as LP-3DGS address this by as…
- RouteBridge: Reliability-Routed Bidirectional Distillation Between Neural Radiance Fields and 3D Gaussian SplattingYuanHang Wang, Xin Cao · arXiv · Sep 9, 2026
Neural radiance fields (NeRFs) and 3D Gaussian Splatting (3DGS) encode a scene with complementary inductive biases, but existing cross-representation distillation typically fixes one representation as teacher for the entire scene. A globall…
- AnimalLift: Reconstructing Animatable 3D Animals from a Single Image by Learning Canonical Shape, Texture, and Fur MapsChunyi Sun, Ruyi Zha, Weijian Deng, Junlin Han et al. · SIGGRAPH Asia 2026 · Sep 8, 2026
Reconstructing a fully animatable 3D animal from a single image remains challenging because animation-ready assets require not only plausible geometry, but also a unified topology, editable appearance, and fur representations compatible wit…
- RoMa-$Ω$: What Feed-Forward 3D Models Know About Image MatchingDavid Nordström, Xinyue Zhang, Thibaut Loiseau, Vincent Lepetit et al. · arXiv · Sep 8, 2026
Learned image matching has experienced significant progress in recent years, culminating in robust and accurate matchers such as RoMa, whose robustness is often attributed to its use of frozen DINO features. In a parallel development, feed-…
- Learning Global Camera Poses from Noisy View-Graphs for Structure from MotionFadi Khatib, Meirav Galun, Ronen Basri · arXiv · Sep 8, 2026
Camera pose estimation is a key step in 3D reconstruction and view-synthesis pipelines. We present a deep, global Structure-from-Motion framework based on learned view-graph aggregation. Our method employs a permutation-equivariant, edge-co…
- LeCor: Learning to Be Corrected by Meta-Learned Test-Time Training for Interactive 3D Lung-Tumour SegmentationYi Luo, Yike Guo, Wenxuan Li, Zongwei Zhou et al. · arXiv · Sep 8, 2026
Delineating lung tumours on computed tomography (CT) takes a considerable share of the time spent on radiotherapy planning, and a contour proposed by a model can be refined interactively by the clinician. Promptable foundation models such a…
- OmniPoint: Universal Monocular Metric Pointcloud from Any CameraBotao Ye, Marc Pollefeys, Ming-Hsuan Yang, Abhijit Kundu · arXiv · Sep 8, 2026
Recovering metric 3D geometry from monocular images is a fundamental computer vision task, yet current methods remain heavily fragmented by fixed camera model assumptions and inflexible input schemes. We present OmniPoint, a unified framewo…
- Point4D: Long-range 4D Motion ReconstructionMinsik Jeon, Jay Karhade, Deva Ramanan, Shubham Tulsiani · arXiv · Sep 8, 2026
We introduce Point4D, a feed-forward model for 4D reconstruction of long-range video sequences. Point4D is able to reliably infer dense per-point 3D trajectories across multi-hundred-frame videos, unlike existing 4D methods that are limited…