3D retrieval

MVTN: Learning Multi-view Transformations for 3D Understanding

Multi-view projection techniques have shown themselves to be highly effective in achieving top-performing results in the recognition of 3D shapes. These methods involve learning how to combine information from multiple view-points. However, the …

TrackNeRF: Bundle Adjusting NeRF from Sparse and Noisy Views via Feature Tracks

Neural radiance fields (NeRFs) generally require many images with accurate poses for accurate novel view synthesis, which does not reflect realistic setups where views can be sparse and poses can be noisy. Previous solutions for learning NeRFs with …

Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors

We present `Magic123`, a two-stage coarse-to-fine solution for high-quality, textured 3D meshes generation from a single unposed image in the wild using both 2D and 3D priors. In the first stage, we optimize a neural radiance field to produce a …

EgoLoc: Revisiting 3D Object Localization from Egocentric Videos with Visual Queries

With the recent advances in video and 3D understanding, novel 4D spatio-temporal methods fusing both concepts have emerged. Towards this direction, the Ego4D Episodic Memory Benchmark proposed a task for Visual Queries with 3D Localization (VQ3D). …

SPARF: Large-Scale Learning of 3D Sparse Radiance Fields from Few Input Images

Recent advances in Neural Radiance Fields (NeRFs) treat the problem of novel view synthesis as Sparse Radiance Field (SRF) optimization using sparse voxels for efficient and fast rendering (Plenoxels,InstantNGP). In order to leverage machine learning …

MVTN: Learning Multi-view Transformations for 3D Understanding

TrackNeRF: Bundle Adjusting NeRF from Sparse and Noisy Views via Feature Tracks

Magic123: One Image to High-Quality 3D Object Generation Using Both 2D and 3D Diffusion Priors

EgoLoc: Revisiting 3D Object Localization from Egocentric Videos with Visual Queries

SPARF: Large-Scale Learning of 3D Sparse Radiance Fields from Few Input Images

VARS: Video Assistant Referee System for Automated Soccer Decision Making From Multiple Views

Voint Cloud: Multi-View Point Cloud Representation for 3D Understanding

MVTN: Multi-View Transformation Network for 3D Shape Recognition