Dataset Information

Natural Language Processing markers in first episode psychosis and people at clinical high-risk.

ABSTRACT: Recent work has suggested that disorganised speech might be a powerful predictor of later psychotic illness in clinical high risk subjects. To that end, several automated measures to quantify disorganisation of transcribed speech have been proposed. However, it remains unclear which measures are most strongly associated with psychosis, how different measures are related to each other and what the best strategies are to collect speech data from participants. Here, we assessed whether twelve automated Natural Language Processing markers could differentiate transcribed speech excerpts from subjects at clinical high risk for psychosis, first episode psychosis patients and healthy control subjects (total N = 54). In-line with previous work, several measures showed significant differences between groups, including semantic coherence, speech graph connectivity and a measure of whether speech was on-topic, the latter of which outperformed the related measure of tangentiality. Most NLP measures examined were only weakly related to each other, suggesting they provide complementary information. Finally, we compared the ability of transcribed speech generated using different tasks to differentiate the groups. Speech generated from picture descriptions of the Thematic Apperception Test and a story re-telling task outperformed free speech, suggesting that choice of speech generation method may be an important consideration. Overall, quantitative speech markers represent a promising direction for future clinical applications.

SUBMITTER: Morgan SE

PROVIDER: S-EPMC8669009 | biostudies-literature | 2021 Dec

REPOSITORIES: biostudies-literature

ACCESS DATA

Publications

Natural Language Processing markers in first episode psychosis and people at clinical high-risk.

Morgan Sarah E SE Diederen Kelly K Vértes Petra E PE Ip Samantha H Y SHY Wang Bo B Thompson Bethany B Demjaha Arsime A De Micheli Andrea A Oliver Dominic D Liakata Maria M Fusar-Poli Paolo P Spencer Tom J TJ McGuire Philip P

Translational psychiatry 20211213 1

Recent work has suggested that disorganised speech might be a powerful predictor of later psychotic illness in clinical high risk subjects. To that end, several automated measures to quantify disorganisation of transcribed speech have been proposed. However, it remains unclear which measures are most strongly associated with psychosis, how different measures are related to each other and what the best strategies are to collect speech data from participants. Here, we assessed whether twelve autom ...[more]

PMID: 34903724

Similar Datasets

Project description:Language production has often been described as impaired in psychiatric diseases such as in psychosis. Nevertheless, little is known about the characteristics of linguistic difficulties and their relation with other cognitive domains in patients with a first episode of psychosis (FEP), either affective or non-affective. To deepen our comprehension of linguistic profile in FEP, 133 patients with FEP (95 non-affective, FEP-NA; 38 affective, FEP-A) and 133 healthy controls (HC) were assessed with a narrative discourse task. Speech samples were systematically analyzed with a well-established multilevel procedure investigating both micro- (lexicon, morphology, syntax) and macro-linguistic (discourse coherence, pragmatics) levels of linguistic processing. Executive functioning and IQ were also evaluated. Both linguistic and neuropsychological measures were secondarily implemented with a machine learning approach in order to explore their predictive accuracy in classifying participants as FEP or HC. Compared to HC, FEP patients showed language production difficulty at both micro- and macro-linguistic levels. As for the former, FEP produced shorter and simpler sentences and fewer words per minute, along with a reduced number of lexical fillers, compared to HC. At the macro-linguistic level, FEP performance was impaired in local coherence, which was paired with a higher percentage of utterances with semantic errors. Linguistic measures were not correlated with any neuropsychological variables. No significant differences emerged between FEP-NA and FEP-A (p≥0.02, after Bonferroni correction). Machine learning analysis showed an accuracy of group prediction of 76.36% using language features only, with semantic variables being the most impactful. Such a percentage was enhanced when paired with clinical and neuropsychological variables. Results confirm the presence of language production deficits already at the first episode of the illness, being such impairment not related to other cognitive domains. The high accuracy obtained by the linguistic set of features in classifying groups support the use of machine learning methods in neuroscience investigations.

Dataset Information

Natural Language Processing markers in first episode psychosis and people at clinical high-risk.

Publications

Natural Language Processing markers in first episode psychosis and people at clinical high-risk.

Similar Datasets

OmicsDI is part of the ELIXIR infrastructure

Tweets