VLM-Grounder
[CoRL 2024] VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
About
We present VLM-Grounder, a novel framework using vision-language models (VLMs) for zero-shot 3D visual grounding based solely on 2D images. VLM-Grounder dynamically stitches image sequences, employs a grounding and feedback scheme to find the target object, and uses a multi-view ensemble projection to accurately estimate 3D bounding boxes. Experiments on ScanRefer and Nr3D datasets show VLM-Grounder outperforms previous zero-shot methods, achieving 51.6% Acc@0.25 on ScanRefer and 48.0% Acc on Nr3D, without relying on 3D geometry or object priors.
Open Source Health
- Stars
- 133
- Forks
- 4
- License
- Not stated
- Last commit
- 1 years ago
Resources & Links
Related Categories
Vendor
Intern Robotics
Building inclusive infrastructure for Embodied AI, from Shanghai AI Lab
Quick Links
Open Source
More by Intern Robotics
Related Products
The Algorithm
Source code for the X Recommendation Algorithm
Shared categories
Recommenders
TensorFlow Recommenders is a library for building recommender system models using TensorFlow.
Shared categories
RPG_KDD2025
This repository provides the code for implementing RPG described in our KDD'25 paper "Generating Long Semantic IDs in Parallel for Recommendation".
Shared categories
Pairec
A Go web framework for quickly building recommendation online services based on JSON configuration.
Shared categories
Textrank
TextRank implementation for Python 3.
Shared categories