The Language-Guided Navigation module leverages an LLM (like ChatGPT) and the open-set O3D-SIM.The Language-Guided Navigation module leverages an LLM (like ChatGPT) and the open-set O3D-SIM.

VLN: LLM and CLIP for Instance-Specific Navigation on 3D Maps

2025/12/16 10:28
3 min read
For feedback or concerns regarding this content, please contact us at [email protected]

Abstract and 1 Introduction

  1. Related Works

    2.1. Vision-and-Language Navigation

    2.2. Semantic Scene Understanding and Instance Segmentation

    2.3. 3D Scene Reconstruction

  2. Methodology

    3.1. Data Collection

    3.2. Open-set Semantic Information from Images

    3.3. Creating the Open-set 3D Representation

    3.4. Language-Guided Navigation

  3. Experiments

    4.1. Quantitative Evaluation

    4.2. Qualitative Results

  4. Conclusion and Future Work, Disclosure statement, and References

3.4. Language-Guided Navigation

In this section, we leverage the LLM-based approach from [1], which uses ChatGPT [35] to understand and map language commands to pre-defined function primitives that the robot can understand and execute. However, there are a few differences between our current approach and the approach in [1] regarding the use case of the LLM and the implementation of our function primitives. The previous approach used the LLM’s ability to bring in an open-set understanding by mapping general queries to the already-known closed-set class labels obtained via Mask2Former [7].

\ However, given the open-set nature of our new representation, O3D-SIM, the LLM does not need to do that. Figure 4 shows both approaches’ code output differences. The function primitives work similarly to the older approach, requiring the desired object type and its instance as an input. But now, the desired object is not from a pre-defined set of classes but a small query defining the object, so the implementation to find the desired location changes. We use the text and image-aligned nature of CLIP embeddings to find the desired object, where the input description is passed to the model, and its corresponding embedding is used to find the object in O3D-SIM.

\ A cosine similarity is calculated between the embedding of the description and all the embeddings of our representation. These are ranked in a decreasing order, and the desired instance is selected. Once the instance is finalized, a goal corresponding to this instance is generated and passed to the navigation stack for autonomous navigation of the robot, hence achieving Language-Guided Navigation.

\

:::info Authors:

(1) Laksh Nanwani, International Institute of Information Technology, Hyderabad, India; this author contributed equally to this work;

(2) Kumaraditya Gupta, International Institute of Information Technology, Hyderabad, India;

(3) Aditya Mathur, International Institute of Information Technology, Hyderabad, India; this author contributed equally to this work;

(4) Swayam Agrawal, International Institute of Information Technology, Hyderabad, India;

(5) A.H. Abdul Hafez, Hasan Kalyoncu University, Sahinbey, Gaziantep, Turkey;

(6) K. Madhava Krishna, International Institute of Information Technology, Hyderabad, India.

:::


:::info This paper is available on arxiv under CC by-SA 4.0 Deed (Attribution-Sharealike 4.0 International) license.

:::

\

Market Opportunity
null Logo
null Price(null)
--
----
USD
null (null) Live Price Chart
Disclaimer: The articles reposted on this site are sourced from public platforms and are provided for informational purposes only. They do not necessarily reflect the views of MEXC. All rights remain with the original authors. If you believe any content infringes on third-party rights, please contact [email protected] for removal. MEXC makes no guarantees regarding the accuracy, completeness, or timeliness of the content and is not responsible for any actions taken based on the information provided. The content does not constitute financial, legal, or other professional advice, nor should it be considered a recommendation or endorsement by MEXC.
Tags:

You May Also Like

SEC decisions scrutinized as senator seeks records on crypto enforcement rollbacks

SEC decisions scrutinized as senator seeks records on crypto enforcement rollbacks

The post SEC decisions scrutinized as senator seeks records on crypto enforcement rollbacks appeared on BitcoinEthereumNews.com. U.S. securities regulators have
Share
BitcoinEthereumNews2026/03/31 08:08
BitGo lists HYPE token for trading

BitGo lists HYPE token for trading

The post BitGo lists HYPE token for trading appeared on BitcoinEthereumNews.com. Key Takeaways BitGo has added HYPE token to its supported trading assets. HYPE is the native token of the Hyperliquid protocol, a decentralized exchange and layer-1 blockchain. BitGo added HYPE token for trading today, expanding access to the digital asset from the Hyperliquid protocol. The custody and trading platform now supports HYPE, allowing institutional and retail clients to trade the token through BitGo’s services. Hyperliquid operates as a decentralized exchange and layer-1 blockchain focused on perpetual futures trading. Source: https://cryptobriefing.com/bitgo-lists-hype-token-hyperliquid/
Share
BitcoinEthereumNews2025/09/18 07:01
Crypto Supercycle in 2025? DeepSeek Ranks the Best Altcoins to Buy Right Now

Crypto Supercycle in 2025? DeepSeek Ranks the Best Altcoins to Buy Right Now

The post Crypto Supercycle in 2025? DeepSeek Ranks the Best Altcoins to Buy Right Now appeared on BitcoinEthereumNews.com. Crypto Supercycle in 2025? DeepSeek Ranks the Best Altcoins to Buy Right Now Sign Up for Our Newsletter! For updates and exclusive offers enter your email. As a crypto writer, Krishi splits his time between decoding the chaos of the markets and writing about it in a way that doesn’t put you to sleep. He’s been at it for nearly two years in the crypto trenches. Yes, he regrets missing the magnificent rallies that came before that (who doesn’t!), but he’s more than ready to put his money where his words are. Before diving headfirst into crypto, Krishi spent over five years writing for some of the biggest names in tech, including TechRadar, Tom’s Guide, and PC Gaming, covering everything from gadgets and cybersecurity to gaming and software. When he’s not scouring and writing about the latest happenings in crypto, Krishi trades the forex market while keeping crypto in his long-term HODL plans. He’s a Bitcoin believer, though he never lets that bias creep into his writing. This website uses cookies. By continuing to use this website you are giving consent to cookies being used. Visit our Privacy Center or Cookie Policy. I Agree Source: https://bitcoinist.com/crypto-supercycle-2025-best-altcoins-to-buy-now-deepseek/
Share
BitcoinEthereumNews2025/09/18 01:45