This work moves beyond closed-set segmentation (Mask2Former) to open-set detection using SAM and Grounding DINO.This work moves beyond closed-set segmentation (Mask2Former) to open-set detection using SAM and Grounding DINO.

Foundation Models for 3D Scenes: DINOv2 vs. CLIP for Instance Differentiation

2025/12/11 02:00

Abstract and 1 Introduction

  1. Related Works

    2.1. Vision-and-Language Navigation

    2.2. Semantic Scene Understanding and Instance Segmentation

    2.3. 3D Scene Reconstruction

  2. Methodology

    3.1. Data Collection

    3.2. Open-set Semantic Information from Images

    3.3. Creating the Open-set 3D Representation

    3.4. Language-Guided Navigation

  3. Experiments

    4.1. Quantitative Evaluation

    4.2. Qualitative Results

  4. Conclusion and Future Work, Disclosure statement, and References

2.2. Semantic Scene Understanding and Instance Segmentation

f 3D scenes. This domain has been thoroughly explored using closed-set vocabulary methods, including our prior work [1], which utilizes Mask2Former [7] for image segmentation. Various studies [18, 19, 20] have adopted a similar approach to achieve object segmentation, resulting in a closed-set framework. While these methods are effective, they are constrained by the limitation of predefined object categories. Our approach employs SAM [21] to acquire segmentation masks for open-set detection. Moreover, our methodology, distinct from many existing techniques that depend heavily on extensive pre-training or fine-tuning, integrates these models to forge a more comprehensive and adaptable 3D scene representation. This emphasizes enhanced semantic understanding and spatial awareness.

\ To improve the semantic understanding of the objects detected within our images, we harness detailed feature representations using two foundational models: CLIP [9] and DINOv2 [10]. DINOv2, a Vision Transformer trained through self-supervision, recognises pixel-level correspondences between images and captures spatial hierarchies. Compared to CLIP, DINOv2 more effectively distinguishes between two distinct instances of the same object type, which poses challenges for CLIP.

\ It’s crucial to differentiate individual instances following the semantic identification of objects. Early methods employed a Region Proposal Network (RPN) to predict bounding boxes for these instances [22]. Alternatively, some strategies suggest a generalized architecture for managing panoptic segmentation [23]. In our preceding approach, we utilized the segmentation model Mask2Former [7], which employs an attention mechanism to isolate object-centric features. Recent research also tackles semantic scene understanding using open vocabularies [24], utilizing multi-view fusion and 3D convolutions to derive dense features from an open-vocabulary embedding space for precise semantic segmentation. Our current pipeline leverages Grounding DINO [25] to generate bounding boxes, which are then input into the Segment Anything Model (SAM) [21] to produce individual object masks, thus enabling instance segmentation within the scene.

\

:::info Authors:

(1) Laksh Nanwani, International Institute of Information Technology, Hyderabad, India; this author contributed equally to this work;

(2) Kumaraditya Gupta, International Institute of Information Technology, Hyderabad, India;

(3) Aditya Mathur, International Institute of Information Technology, Hyderabad, India; this author contributed equally to this work.

(4) Swayam Agrawal, International Institute of Information Technology, Hyderabad, India;

(5) A.H. Abdul Hafez, Hasan Kalyoncu University, Sahinbey, Gaziantep, Turkey;

(6) K. Madhava Krishna, International Institute of Information Technology, Hyderabad, India.

:::


:::info This paper is available on arxiv under CC by-SA 4.0 Deed (Attribution-Sharealike 4.0 International) license.

:::

\

Disclaimer: The articles reposted on this site are sourced from public platforms and are provided for informational purposes only. They do not necessarily reflect the views of MEXC. All rights remain with the original authors. If you believe any content infringes on third-party rights, please contact [email protected] for removal. MEXC makes no guarantees regarding the accuracy, completeness, or timeliness of the content and is not responsible for any actions taken based on the information provided. The content does not constitute financial, legal, or other professional advice, nor should it be considered a recommendation or endorsement by MEXC.

You May Also Like

a16z Opens First Asia Office: Park From Naver and Monad to Lead

a16z Opens First Asia Office: Park From Naver and Monad to Lead

The post a16z Opens First Asia Office: Park From Naver and Monad to Lead appeared on BitcoinEthereumNews.com. a16z crypto, the crypto-focused venture arm of Andreessen Horowitz, has officially entered the Asian market with the opening of its first regional office in Seoul, South Korea. The Silicon Valley-based venture fund appointed Sungmo Park as Head of APAC go-to-market to lead the Seoul operations. Park brings extensive regional expertise from his previous roles at Monad Foundation and Polygon Labs. Sponsored Sponsored Asia Emerges as Global Crypto Powerhouse Chief Operating Officer Anthony Albanese made the announcement. The decision to establish a physical presence in Asia reflects the region’s growing dominance in global crypto adoption. Chainalysis reports that Asia-Pacific accounted for $2.36 trillion in on-chain value over the 12 months to June 2025. This figure represents a 69% increase from $1.4 trillion in the previous year. South Korea stands as the world’s second-largest crypto market, with nearly one in three adults holding digital assets—a rate that surpasses stock ownership. Japan has seen on-chain activity surge 120% over the past year. Singapore has one of the highest crypto ownership rates in the world. About 40% of Gen Z and Millennials in the country invest in digital assets. India leads the Chainalysis Global Crypto Adoption Index, driven by mobile-first technology adoption and limited access to traditional banking. Notably, 11 of the top 20 countries in Chainalysis’s Global Crypto Adoption Index are located in Asia. Excited to announce that @a16zcrypto is expanding into Asia and opening our first office in Seoul, South Korea. As part of this, we’re thrilled to have @sungmo_apac16z join our team as Head of APAC go-to-market to lead the Seoul office and start building our presence in the… pic.twitter.com/KBljioBCqx — Anthony Albanese (@AAlbaneseNY) December 10, 2025 The Seoul launch follows other leading venture and crypto firms boosting their Asian presence. Competition for deals, talent, and growth is intensifying as the…
Share
BitcoinEthereumNews2025/12/11 10:34
The Crucial Proposal Arriving This Month

The Crucial Proposal Arriving This Month

The post The Crucial Proposal Arriving This Month appeared on BitcoinEthereumNews.com. South Korean Stablecoin Regulation: The Crucial Proposal Arriving This Month Skip to content Home Crypto News South Korean Stablecoin Regulation: The Crucial Proposal Arriving This Month Source: https://bitcoinworld.co.in/south-korean-stablecoin-regulation-proposal/
Share
BitcoinEthereumNews2025/12/11 09:52