1 |
Local-Global Context Aware Transformer for Language-Guided Video Segmentation ...
|
|
|
|
BASE
|
|
Show details
|
|
2 |
Compositional Temporal Grounding with Structured Variational Cross-Graph Correspondence Learning ...
|
|
|
|
BASE
|
|
Show details
|
|
3 |
Cumulative Effects of Physical, Chemical, and Biological Measures on Algae Growth Inhibition
|
|
|
|
In: Water; Volume 14; Issue 6; Pages: 877 (2022)
|
|
BASE
|
|
Show details
|
|
4 |
Rethinking Cross-modal Interaction from a Top-down Perspective for Referring Video Object Segmentation ...
|
|
|
|
BASE
|
|
Show details
|
|
5 |
Contrastive Video-Language Segmentation ...
|
|
|
|
Abstract:
We focus on the problem of segmenting a certain object referred by a natural language sentence in video content, at the core of formulating a pinpoint vision-language relation. While existing attempts mainly construct such relation in an implicit way, i.e., grid-level multi-modal feature fusion, it has been proven problematic to distinguish semantically similar objects under this paradigm. In this work, we propose to interwind the visual and linguistic modalities in an explicit way via the contrastive learning objective, which directly aligns the referred object and the language description and separates the unreferred content apart across frames. Moreover, to remedy for the degradation problem, we present two complementary hard instance mining strategies, i.e., Language-relevant Channel Filter and Relative Hard Instance Construction. They encourage the network to exclude visual-distinguishable feature and to focus on easy-confused objects during the contrastive training. Extensive experiments on two ...
|
|
Keyword:
Computation and Language cs.CL; Computer Vision and Pattern Recognition cs.CV; FOS Computer and information sciences
|
|
URL: https://dx.doi.org/10.48550/arxiv.2109.14131 https://arxiv.org/abs/2109.14131
|
|
BASE
|
|
Hide details
|
|
6 |
CLIP: A Dataset for Extracting Action Items for Physicians from Hospital Discharge Notes ...
|
|
|
|
BASE
|
|
Show details
|
|
7 |
Constructing a Psychometric Testbed for Fair Natural Language Processing ...
|
|
|
|
BASE
|
|
Show details
|
|
8 |
Juegos serios en web para la auto-protección y prevención del COVID-19: Desarrollo y pruebas de usabilidad
|
|
|
|
In: Comunicar: Revista científica iberoamericana de comunicación y educación, ISSN 1134-3478, Nº 69, 2021 (Ejemplar dedicado a: Participación ciudadana en la esfera digital), pags. 97-111 (2021)
|
|
BASE
|
|
Show details
|
|
10 |
ActBERT: Learning Global-Local Video-Text Representations ...
|
|
|
|
BASE
|
|
Show details
|
|
11 |
Speech-to-Singing Conversion in an Encoder-Decoder Framework ...
|
|
|
|
BASE
|
|
Show details
|
|
12 |
Symbiotic Attention with Privileged Information for Egocentric Action Recognition ...
|
|
|
|
BASE
|
|
Show details
|
|
13 |
Grounded and Controllable Image Completion by Incorporating Lexical Semantics ...
|
|
|
|
BASE
|
|
Show details
|
|
15 |
Measurement of $W^{\pm}$-boson and $Z$-boson production cross-sections in $pp$ collisions at $\sqrt{s}=2.76$ TeV with the ATLAS detector
|
|
|
|
BASE
|
|
Show details
|
|
16 |
Baidu-UTS Submission to the EPIC-Kitchens Action Recognition Challenge 2019 ...
|
|
|
|
BASE
|
|
Show details
|
|
17 |
Китайско-русский параллельный дискурсивный корпус: выравнивание на уровне клаузы и статистический анализ ; Chinese-Russian Parallel Discourse Corpus: Alignment of Clauses and Statistical Analysis
|
|
|
|
BASE
|
|
Show details
|
|
18 |
Extensive translation of circular RNAs driven by N6-methyladenosine
|
|
|
|
BASE
|
|
Show details
|
|
20 |
Китайско-русский параллельный корпус с дискурсивно-структурной разметкой: теоретические аспекты ; Theoretical aspects of building a Chinese-Russian Parallel Corpus with discourse-structure annotation
|
|
|
|
BASE
|
|
Show details
|
|
|
|