Unknown

Dataset Information

0

CELER: A 365-Participant Corpus of Eye Movements in L1 and L2 English Reading.


ABSTRACT: We present CELER (Corpus of Eye Movements in L1 and L2 English Reading), a broad coverage eye-tracking corpus for English. CELER comprises over 320,000 words, and eye-tracking data from 365 participants. Sixty-nine participants are L1 (first language) speakers, and 296 are L2 (second language) speakers from a wide range of English proficiency levels and five different native language backgrounds. As such, CELER has an order of magnitude more L2 participants than any currently available eye movements dataset with L2 readers. Each participant in CELER reads 156 newswire sentences from the Wall Street Journal (WSJ), in a new experimental design where half of the sentences are shared across participants and half are unique to each participant. We provide analyses that compare L1 and L2 participants with respect to standard reading time measures, as well as the effects of frequency, surprisal, and word length on reading times. These analyses validate the corpus and demonstrate some of its strengths. We envision CELER to enable new types of research on language processing and acquisition, and to facilitate interactions between psycholinguistics and natural language processing (NLP).

SUBMITTER: Berzak Y 

PROVIDER: S-EPMC9692049 | biostudies-literature | 2022

REPOSITORIES: biostudies-literature

altmetric image

Publications

CELER: A 365-Participant Corpus of Eye Movements in L1 and L2 English Reading.

Berzak Yevgeni Y   Nakamura Chie C   Smith Amelia A   Weng Emily E   Katz Boris B   Flynn Suzanne S   Levy Roger R  

Open mind : discoveries in cognitive science 20220701


We present CELER (<b>C</b>orpus of <b>E</b>ye Movements in <b>L</b>1 and L2 <b>E</b>nglish <b>R</b>eading), a broad coverage eye-tracking corpus for English. CELER comprises over 320,000 words, and eye-tracking data from 365 participants. Sixty-nine participants are L1 (first language) speakers, and 296 are L2 (second language) speakers from a wide range of English proficiency levels and five different native language backgrounds. As such, CELER has an order of magnitude more L2 participants tha  ...[more]

Similar Datasets

| S-EPMC9005965 | biostudies-literature
| S-EPMC11607845 | biostudies-literature
| S-EPMC8513778 | biostudies-literature
| S-EPMC6363172 | biostudies-literature
| S-EPMC7157570 | biostudies-literature
| S-EPMC7752842 | biostudies-literature
| S-EPMC6722069 | biostudies-literature
| S-EPMC4630541 | biostudies-literature
| S-EPMC10292930 | biostudies-literature