2 |
On the Use of Linguistic Features for the Evaluation of Generative Dialogue Systems ...
|
|
|
|
Abstract:
Automatically evaluating text-based, non-task-oriented dialogue systems (i.e., `chatbots') remains an open problem. Previous approaches have suffered challenges ranging from poor correlation with human judgment to poor generalization and have often required a gold standard reference for comparison or human-annotated data. Extending existing evaluation methods, we propose that a metric based on linguistic features may be able to maintain good correlation with human judgment and be interpretable, without requiring a gold-standard reference or human-annotated data. To support this proposition, we measure and analyze various linguistic features on dialogues produced by multiple dialogue models. We find that the features' behaviour is consistent with the known properties of the models tested, and is similar across domains. We also demonstrate that this approach exhibits promising properties such as zero-shot generalization to new domains on the related task of evaluating response relevance. ...
|
|
Keyword:
Computation and Language cs.CL; FOS Computer and information sciences; Machine Learning cs.LG
|
|
URL: https://arxiv.org/abs/2104.06335 https://dx.doi.org/10.48550/arxiv.2104.06335
|
|
BASE
|
|
Hide details
|
|
3 |
TorontoCL at CMCL 2021 Shared Task: RoBERTa with Multi-Stage Fine-Tuning for Eye-Tracking Prediction ...
|
|
|
|
BASE
|
|
Show details
|
|
4 |
Quantifying the Task-Specific Information in Text-Based Classifications ...
|
|
|
|
BASE
|
|
Show details
|
|
5 |
An {E}valuation of {D}isentangled {R}epresentation {L}earning for {T}exts ...
|
|
|
|
BASE
|
|
Show details
|
|
6 |
How is BERT surprised? Layerwise detection of linguistic anomalies ...
|
|
|
|
BASE
|
|
Show details
|
|
7 |
Comparing Pre-trained and Feature-Based Models for Prediction of Alzheimer's Disease Based on Speech
|
|
|
|
In: Front Aging Neurosci (2021)
|
|
BASE
|
|
Show details
|
|
8 |
Identification of primary and collateral tracks in stuttered speech
|
|
|
|
In: LREC 2020 - 12th Conference on Language Resources and Evaluation ; https://hal.archives-ouvertes.fr/hal-02959454 ; LREC 2020 - 12th Conference on Language Resources and Evaluation, May 2020, Marseille, France (2020)
|
|
BASE
|
|
Show details
|
|
9 |
Semantic coordinates analysis reveals language changes in the AI field ...
|
|
|
|
BASE
|
|
Show details
|
|
11 |
To BERT or Not To BERT: Comparing Speech and Language-based Approaches for Alzheimer's Disease Detection ...
|
|
|
|
BASE
|
|
Show details
|
|
12 |
An information theoretic view on selecting linguistic probes ...
|
|
|
|
BASE
|
|
Show details
|
|
13 |
Examining the rhetorical capacities of neural language models ...
|
|
|
|
BASE
|
|
Show details
|
|
14 |
A textual analysis of US corporate social responsibility reports
|
|
|
|
BASE
|
|
Show details
|
|
15 |
Lexical Features Are More Vulnerable, Syntactic Features Have More Predictive Power ...
|
|
|
|
BASE
|
|
Show details
|
|
16 |
Representation Learning for Discovering Phonemic Tone Contours ...
|
|
|
|
BASE
|
|
Show details
|
|
18 |
The Effect of Heterogeneous Data for Alzheimer's Disease Detection from Speech ...
|
|
|
|
BASE
|
|
Show details
|
|
19 |
Detecting cognitive impairments by agreeing on interpretations of linguistic features ...
|
|
|
|
BASE
|
|
Show details
|
|
20 |
Deconfounding age effects with fair representation learning when assessing dementia ...
|
|
|
|
BASE
|
|
Show details
|
|
|
|