This article outlines the OW‑VISCap framework, which jointly detects, segments, and captions both seen and unseen objects within a video.This article outlines the OW‑VISCap framework, which jointly detects, segments, and captions both seen and unseen objects within a video.

Teaching AI to See and Speak: Inside the OW‑VISCap Approach

2025/11/04 17:46
3 min read

Abstract and 1. Introduction

  1. Related Work

    2.1 Open-world Video Instance Segmentation

    2.2 Dense Video Object Captioning and 2.3 Contrastive Loss for Object Queries

    2.4 Generalized Video Understanding and 2.5 Closed-World Video Instance Segmentation

  2. Approach

    3.1 Overview

    3.2 Open-World Object Queries

    3.3 Captioning Head

    3.4 Inter-Query Contrastive Loss and 3.5 Training

  3. Experiments and 4.1 Datasets and Evaluation Metrics

    4.2 Main Results

    4.3 Ablation Studies and 4.4 Qualitative Results

  4. Conclusion, Acknowledgements, and References

\ Supplementary Material

A. Additional Analysis

B. Implementation Details

C. Limitations

3 Approach

Given a video, our goal is to jointly detect, segment and caption object instances present in the video. Importantly, note that object instance categories may not be part of the training set (e.g., the parachutes shown in Fig. 3 (top row)), placing our goal in an open-world setting. To achieve this goal, a given video is first broken into short clips, each consisting of T frames. Each clip is processed using our approach OW-VISCap. We discuss merging of the results of each clip in Sec. 4.

\ We provide an overview of OW-VISCap to process each clip in Sec. 3.1. We then discuss our contributions: (a) introduction of open-world object queries in Sec. 3.2, (b) use of masked attention for object-centric captioning in Sec. 3.3, and (c) use of inter-query contrastive loss to ensure that the object qeries are different from each other in Sec. 3.4. In Sec. 3.5, we discuss the final training objective.

3.1 Overview

\ Both open- and closed-world object queries are processed by our specifically designed captioning head which yields an object-centric caption, a classification head which yields a category label, and a detection head which yields either a segmentation mask or a bounding-box.

\

\ We introduce an inter-query contrastive loss to ensure that the object queries are encouraged to differ from each other. We provide details in Sec. 3.4. For closed world objects, this loss helps in removing highly overlapping false positives. For open-world objects, it helps in the discovery of new objects.

\ Finally, we provide the full training objective in Sec. 3.5.

\

3.2 Open-World Object Queries

\

\

\ We first match the ground truth objects with the open-world predictions by minimizing a matching cost using the Hungarian algorithm [34]. The optimal matching is then used to calculate the final open-world loss.

\

\

3.3 Captioning Head

\

\

3.4 Inter-Query Contrastive Loss

\

\

3.5 Training

Our total training loss is

\ Table 1: Open-world tracking accuracy (OWTA) on the BURST validation and test sets for all, common (comm.) and uncommon (unc.) categories of objects. Onl. refers to online frame-by-frame processing. The best scores are highlighted in bold font, and the second-best scores are underlined.

\ Table 2: Dense video object captioning results on the VidSTG [57] dataset. Off. indicates offline methods and onl. refers to online methods.

\

:::info Authors:

(1) Anwesa Choudhuri, University of Illinois at Urbana-Champaign (anwesac2@illinois.edu);

(2) Girish Chowdhary, University of Illinois at Urbana-Champaign (girishc@illinois.edu);

(3) Alexander G. Schwing, University of Illinois at Urbana-Champaign (aschwing@illinois.edu).

:::


:::info This paper is available on arxiv under CC by 4.0 Deed (Attribution 4.0 International) license.

:::

\

Market Opportunity
null Logo
null Price(null)
--
----
USD
null (null) Live Price Chart
Disclaimer: The articles reposted on this site are sourced from public platforms and are provided for informational purposes only. They do not necessarily reflect the views of MEXC. All rights remain with the original authors. If you believe any content infringes on third-party rights, please contact crypto.news@mexc.com for removal. MEXC makes no guarantees regarding the accuracy, completeness, or timeliness of the content and is not responsible for any actions taken based on the information provided. The content does not constitute financial, legal, or other professional advice, nor should it be considered a recommendation or endorsement by MEXC.

You May Also Like

Siren Token Sheds 16.4% After 54% Retreat From All-Time High

Siren Token Sheds 16.4% After 54% Retreat From All-Time High

Siren token experienced a sharp 16.4% decline in the past 24 hours, trading at $0.247 as the market cap contracted by $34.4 million. Our analysis of on-chain metrics
Share
Blockchainmagazine2026/03/02 05:03
Privacy is ‘Constant Battle’ Between Blockchain Stakeholders and State

Privacy is ‘Constant Battle’ Between Blockchain Stakeholders and State

The post Privacy is ‘Constant Battle’ Between Blockchain Stakeholders and State appeared on BitcoinEthereumNews.com. Blockchain industry participants and regulators continue wrangling over privacy rights as the European Union’s sweeping Anti-Money Laundering (AML) rules look set to ban privacy-preserving tokens and anonymous crypto accounts starting in 2027. Credit institutions, financial institutions and crypto asset service providers (CASPs) will be prohibited from maintaining anonymous accounts or handling privacy-preserving cryptocurrencies under the EU’s new Anti-Money Laundering Regulation (AMLR) that will go into effect in 2027, Cointelegraph reported in May. Maintaining the right to access privacy-preserving coins like Monero (XMR) has been a “constant battle” between blockchain industry stakeholders and regulators, according to Anja Blaj, an independent legal consultant and policy expert at the European Crypto Initiative. “Once you think of how the states want to play out their policies, they want to establish control. They want to understand who the parties are that transact among themselves,” said Blaj, speaking during Cointelegraph’s daily live X spaces show on Sept. 3. “[The state] wants to understand that to be able to prevent whatever crime and scamming is happening, and we want to enforce the policies that we create as a society.” Her comments came as the EU ramped up its regulatory oversight of the crypto industry, building on the bloc’s Markets in Crypto-Assets Regulation (MiCA). Related: Swiss banks complete first blockchain-based legally binding payment Room for negotiation remains While the AML framework is final, regulatory experts still see potential for negotiation until it rolls out in 2027. Policymaking is a “continuous conversation,” meaning that “nothing is set in stone, even if the regulation is already out,” said Blaj. “There are still ways to either talk to the regulators, see how it’s going to play out, how it’s going to be enforced.” While there’s always room for negotiations with policymakers, the regulation concerning privacy-preserving cryptocurrencies and accounts is becoming “more…
Share
BitcoinEthereumNews2025/09/18 12:45
Santander’s Openbank Enables Bitcoin, Litecoin, POL, Ethereum, and Altcoin Trading for German Customers

Santander’s Openbank Enables Bitcoin, Litecoin, POL, Ethereum, and Altcoin Trading for German Customers

Santander’s digital bank has launched crypto trading in Germany, letting customers buy, sell, and hold these assets. At launch, Openbank customers in Germany can get their hands on Bitcoin, Ethereum, Cardano, Litecoin, and Polygon. Openbank, the digital arm of Banco Santander, has just rolled out a new crypto trading service for its retail customers in [...]]]>
Share
Crypto News Flash2025/09/18 04:00