Full list of my publications, most recent first. See also Google Scholar and ORCID for citation counts and up-to-date metrics.
A central question in cognition is how representations are integrated across different modalities, such as language and vision. One prominent hypothesis posits the existence of an abstract, prelinguistic “language of vision” as a representational system that organizes meaning compositionally, enabling cross‐modal integration. This hypothesis predicts that the language of vision operates universally, independent of linguistic surface features such as word order. We conducted eye‐tracking experiments where participants described visual scenes in English, Portuguese, and Japanese. By analyzing spoken descriptions alongside eye‐movement sequences divided into planning and articulation phases, we demonstrate that semantic similarity between sentences strongly predicts the similarity of associated scan patterns in all three languages, even across scenes and between sentences in different languages. In contrast, the effect of syntactic constraints was secondary and transient: it was restricted to within‐language and within‐scene comparisons, and temporally confined to the early planning phase of the utterance. Our findings support an interactive account of cross‐modal coordination in which a universal language of vision provides stable semantic scaffolding, while syntax serves as a local constraint, primarily active during message linearization.
Visual working memory (VWM) operates within structured environments, yet it remains debated whether representations are guided primarily by the global scene configuration (gist) or the intrinsic features of objects. This question extends to cognitive ageing, where declines in VWM co-occur with a relative preservation of global information over fine-grained visual details. To disentangle these influences, younger and older adults detected changes to an object’s identity, location, or both, while its semantic consistency with the scene was manipulated. Unlike our previous work, where location changes disrupted the spatial layout, here we preserved the layout by swapping the critical object, thereby isolating memory for object-location binding from sensitivity to global disruptions. Across age groups, conjunctive changes were detected more accurately than single-feature changes. Crucially, detecting location changes was significantly more difficult when the layout was preserved (swaps) than when it was disrupted (displacements). This demonstrates that while layout disruptions provide salient cues, recalling an object’s location from a stable configuration requires retrieving its intrinsic features. This is further supported by a detection advantage for semantically inconsistent objects (e.g., a torch in a bathroom), which was observed specifically in younger adults, suggesting an age-related decline in the strategic use of contextual violations. Eye-tracking during retrieval revealed longer fixations for inconsistent objects, indicating increased effort in integrating them. In contrast, consistent objects were fixated faster, a novel retrieval-phase finding likely driven by memory from their initial encoding as inconsistent. While core attentional mechanisms were preserved with age, our results argue that when global structure remains unaltered, VWM is guided by hierarchical representations of object features, in which semantic meaning plays a central, organising role.
• We manipulated task goals to examine how partners coordinate gaze and speech. • Gaze alignment was stronger in a structured task with fewer degrees of freedom. • Late-trial dips in gaze alignment hindered performance in a more flexible task. • Speech and gaze exhibit compensatory alignment across different task demands. • Findings support a dynamic, multimodal framework extending beyond Interactive Alignment. Collaborative task performance is assumed to benefit from interpersonal coordination between interacting individuals. Prominent views of language use and social behavior, including the Interactive Alignment Model (IAM; Pickering & Garrod, 2004 ), support this view by building on tasks that require monitoring a partner’s perspective (e.g., in route planning), proposing that behavioral alignment enables conceptual convergence. However, the role of alignment in tasks requiring complementarity (e.g., a “divide and conquer” strategy during joint visual search) remains underexplored. We address this gap by manipulating task goals (route planning vs. visual search) as forty dyads completed ten trials involving subway maps while their eye movements and speech were co-registered. We used Cross Recurrence Quantification Analysis (CRQA) to examine the temporal relationships between partners’ eye fixations and word sequences, generating measures that reveal similarity and dynamic coupling. Dyads exhibited more gaze alignment in route planning than visual search across a range of CRQA metrics. Gaze alignment also varied across the trial and related differently to accuracy: in visual search, greater alignment late in the trial predicted better performance. In speech, route planning prompted longer and more entropic word sequences, but lower overall recurrence than visual search. This finding suggests that the two modalities organize in a compensatory fashion to support distinct task demands. These results support a theoretical framework more general than IAM, in which interactive alignment emerges as a consequence of dynamic adaptation to task goals. Overall, task goals constrain how people coordinate behavior and offer insights into how collaborating partners distribute their multimodal contributions.
Binding, a critical cognitive process likely mediated by attention, is essential for creating coherent object representations within a scene. This process is vulnerable in individuals with dementia, who exhibit deficits in visual working memory (VWM) binding, primarily tested using abstract arrays of standalone objects. To explore how binding operates in more realistic settings across the lifespan, we examined the impact of object saliency and semantic consistency on VWM binding and the role of overt attention. Using an eye-tracking change detection task, we compared younger adults, healthy older adults, and individuals with Mild Cognitive Impairment (MCI). Participants were presented with naturalistic scenes and asked to detect changes in the identity and/or location of objects that were either semantically consistent or inconsistent with their scene context. Across all age groups, semantically inconsistent objects were prioritised during encoding, leading to better change detection than consistent objects. Highly salient objects decreased the inconsistency advantage while being detrimental to detection accuracy when inspected at longer latencies to the first fixation. Longer fixation durations on the critical object were beneficial for recognition. In contrast, delayed initial inspection or frequent subsequent fixations on other objects were detrimental to detection, regardless of age or cognitive impairment. These findings challenge the notion of generalised semantic memory impairment in the prodromal stages of dementia and highlight the importance of efficient attentional control in supporting VWM binding, even in the face of cognitive decline. Overall, preserved low-level and high-level mechanisms of object-scene integration can compensate for age-related cognitive decline, enabling successful binding in naturalistic contexts.
• The speed of visual object recognition is strongly affected by the scene background. • Prior research in this area has tended to use non-ecological scenes and tasks. • Seventy-one adults completed visual search of naturalistic scenes during fMRI. • Participants responded slower when the object was congruent with the background. • Greater dorsal frontoparietal activation observed in such object-congruent trials. Visual attention allows us to navigate complex environments by selecting behaviorally relevant stimuli while suppressing distractors, through a dynamic balance between top-down and bottom-up mechanisms. Extensive attention research has examined the object-context relationship. Some studies have shown that incongruent object-context associations are processed faster, likely due to semantic mismatch-related attentional capture, while others have suggested that schema-driven facilitation may enhance object recognition when the object and context are congruent. Beyond the conflicting findings, translation of this work to real world contexts has been difficult due to the use of non-ecological scenes and stimuli when investigating the object-context congruency relationship. To address this, we employed a goal-directed visual search task and naturalistic indoor scenes during functional MRI (fMRI). Seventy-one healthy adults searched for a target object, either congruent or incongruent within the scene context, following a word cue. We collected accuracy and response time behavioral data, and all fMRI data were processed following standard pipelines, with statistical maps thresholded at p < .05 following multiple comparisons correction. Our results indicated faster response times for incongruent relative to congruent trials, likely reflecting the so-called pop-out effect of schema violations in the incongruent condition. Our neural results indicated that congruent elicited greater activation than incongruent trials in the dorsal frontoparietal attention network and the precuneus, likely reflecting sustained top-down attentional control to locate the targets that blend more seamlessly into the context. These findings highlight the flexible interplay between top-down and bottom-up mechanisms in real-world visual search, emphasizing the dominance of schema-guided top-down processes in congruent contexts and rapid attention capture in incongruent contexts.
Recent conceptualisations of bilingualism are moving away from strict categorisations, towards continuous approaches. This study supports this trend by combining empirical psycholinguistics data with machine learning classification modelling. Support vector classifiers were trained on two datasets of coded productions by Italian speakers to predict the class they belonged to (“monolingual”, “attriters” and “heritage”). All classes can be predicted above chance (>33%), even if the classifier's performance substantially varies, with monolinguals identified much better ( f -score >70%) than attriters ( f -score <50%), which are instead the most confusable class. Further analyses of the classification errors expressed in the confusion matrices qualify that attriters are identified as heritage speakers nearly as often as they are correctly classified. Cluster clitics are the most identifying features for the classification performance. Overall, this study supports a conceptualisation of bilingualism as a continuum of linguistic behaviours rather than sets of a priori established classes.
Although long-term visual memory (LTVM) has a remarkable capacity, the fidelity of its episodic representations can be influenced by at least two intertwined interference mechanisms during the encoding of objects belonging to the same category: the capacity to hold similar episodic traces (e.g., different birds ) and the conceptual similarity of the encoded traces (e.g., a sparrow shares more features with a robin than with a penguin ). The precision of episodic traces can be tested by having participants discriminate lures (unseen objects) from targets (seen objects) representing different exemplars of the same concept (e.g., two visually similar penguins ), which generates interference at retrieval that can be solved if efficient pattern separation happened during encoding. The present study examines the impact of within-category encoding interference on the fidelity of mnemonic object representations, by manipulating an index of cumulative conceptual interference that represents the concurrent impact of capacity and similarity. The precision of mnemonic discrimination was further assessed by measuring the impact of visual similarity between targets and lures in a recognition task. Our results show a significant decrement in the correct identification of targets for increasing interference. Correct rejections of lures were also negatively impacted by cumulative interference as well as by the visual similarity with the target. Most interestingly though, mnemonic discrimination for targets presented with a visually similar lure was more difficult when objects were encoded under lower, not higher, interference. These findings counter a simply additive impact of interference on the fidelity of object representations providing a finer-grained, multi-factorial, understanding of interference in LTVM.
Objective: Retaining the identity or location of decontextualized objects in visual short-term working memory (VWM) is impaired by healthy and pathological ageing, but research remains inconclusive on whether these two features are equally impacted by it. Moreover, it is unclear whether similar impairments would manifest in naturalistic visual contexts. Method: (changed in location and identity). Results: features took longer to be fixated for the first time but required a shorter first pass compared to changes in identity alone which displayed the opposite pattern. Conclusions: Locations of objects are better remembered than their identities; memory for changes is best when involving both features. These mechanisms are spared by pathological ageing as indicated by the similarity between groups besides trivial differences in overall performance. These findings demonstrate that VWM mechanisms in the context of naturalistic scene information are preserved in people with MCI. (PsycInfo Database Record (c) 2023 APA, all rights reserved).
The ability to maintain visual working memory (VWM) associations about the identity and location of objects has at times been found to decrease with age. To date, however, this age-related difficulty was mostly observed in artificial visual contexts (e.g., object arrays), and so it is unclear whether it may manifest in naturalistic contexts, and in which ways. In this eye-tracking study, 26 younger and 24 healthy older adults were asked to detect changes in a critical object situated in a photographic scene (192 in total), about its identity (the object becomes a different object but maintains the same position), location (the object only changes position) or both (the object changes in location and identity). Aging was associated with a lower change detection performance. A change in identity was harder to detect than a location change, and performance was best when both features changed, especially in younger adults. Eye movements displayed minor differences between age groups (e.g., shorter saccades in older adults) but were similarly modulated by the type of change. Latencies to the first fixation were longer and the amplitude of incoming saccades was larger when the critical object changed in location. Once fixated, the target object was inspected for longer when it only changed in identity compared to location. Visually salient objects were fixated earlier, but saliency did not affect any other eye movement measures considered, nor did it interact with the type of change. Our findings suggest that even though aging results in lower performance, it does not selectively disrupt temporary bindings of object identity, location, or their association in VWM, and highlight the importance of using naturalistic contexts to discriminate the cognitive processes that undergo detriment from those that are instead spared by aging.
Recurrence quantification analysis is a widely used method for characterizing patterns in time series. This article presents a comprehensive survey for conducting a wide range of recurrencebased analyses to quantify the dynamical structure of single and multivariate time series and capture coupling properties underlying leader-follower relationships. The basics of recurrence quantification analysis (RQA) and all its variants are formally introduced step-by-step from the simplest autorecurrence to the most advanced multivariate case. Importantly, we show how such RQA methods can be deployed under a single computational framework in R using a substantially renewed version of our crqa 2.0 package. This package includes implementations of several recent advances in recurrencebased analysis, among them applications to multivariate data and improved entropy calculations for categorical data. We show concrete applications of our package to example data, together with a detailed description of its functions and some guidelines on their usage.
Objective: Long-term visual memory representations, measured by recognition performance, degrade as a function of semantic interference, and their strength is related to eye-movement responses. Even though clinical research has examined interference mechanisms in pathological cognitive aging and explored the diagnostic potential of eye-movements in this context, little is known about their interaction in long-term visual memory. Method: An eye-tracking study compared a Mild Cognitive Impaired group with healthy adults. Participants watched a stream of 129 naturalistic images from different semantic categories, presented at different frequencies (1, 6, 12, 24) to induce semantic interference (SI), then asked in a 2-Alternative Forced Choice paradigm to verbally recognize the scene they remembered (old/novel). Results: Recognition accuracy of both groups was negatively impacted by SI, especially in healthy adults. A wider distribution of overt attention across the scene predicted better recognition, especially by the Mild Cognitive Impaired (MCI) participants, although these fixation patterns were influenced by SI. MCI compensated the detrimental effect of SI by focusing overt attention during encoding and so accruing distinctive details of the scene. During recognition, MCI participants widened overt attention to boost retrieval. Independently of the group: (a) the re-instatement of fixations indicated a more successful recall and increased as a function of SI; and (b) attending visually salient regions negatively impacted on recognition accuracy, although the reliance on such regions grew as SI increased. Conclusions: Effects of SI on long-term memory were reduced in MCI participants. They used different oculomotor strategies compared to healthy adults to compensate for its detrimental effects. (PsycInfo Database Record (c) 2021 APA, all rights reserved).
Alzheimer’s disease (AD) patients underperform on a range of tasks requiring semantic processing, but it is unclear whether this impairment is due to a generalised loss of semantic knowledge or to issues in accessing and selecting such information from memory. The objective of this eye-tracking visual search study was to determine whether semantic expectancy mechanisms known to support object recognition in healthy adults are preserved in AD patients. Furthermore, as AD patients are often reported to be impaired in accessing information in extra-foveal vision, we investigated whether that was also the case in our study. Twenty AD patients and 20 age-matched controls searched for a target object among an array of distractors presented extra-foveally. The distractors were either semantically related or unrelated to the target (e.g., a car in an array with other vehicles or kitchen items). Results showed that semantically related objects were detected with more difficulty than semantically unrelated objects by both groups, but more markedly by the AD group. Participants looked earlier and for longer at the critical objects when these were semantically unrelated to the distractors. Our findings show that AD patients can process the semantics of objects and access it in extra-foveal vision. This suggests that their impairments in semantic processing may reflect difficulties in accessing semantic information rather than a generalised loss of semantic memory.
In vision science, a particularly controversial topic is whether and how quickly the semantic information about objects is available outside foveal vision. Here, we aimed at contributing to this debate by coregistering eye movements and EEG while participants viewed photographs of indoor scenes that contained a semantically consistent or inconsistent target object. Linear deconvolution modeling was used to analyze the ERPs evoked by scene onset as well as the fixation-related potentials (FRPs) elicited by the fixation on the target object (t) and by the preceding fixation (t − 1). Object–scene consistency did not influence the probability of immediate target fixation or the ERP evoked by scene onset, which suggests that object–scene semantics was not accessed immediately. However, during the subsequent scene exploration, inconsistent objects were prioritized over consistent objects in extrafoveal vision (i.e., looked at earlier) and were more effortful to process in foveal vision (i.e., looked at longer). In FRPs, we demonstrate a fixation-related N300/N400 effect, whereby inconsistent objects elicit a larger frontocentral negativity than consistent objects. In line with the behavioral findings, this effect was already seen in FRPs aligned to the pretarget fixation t − 1 and persisted throughout fixation t, indicating that the extraction of object semantics can already begin in extrafoveal vision. Taken together, the results emphasize the usefulness of combined EEG/eye movement recordings for understanding the mechanisms of object–scene integration during natural viewing.
Eye-tracking studies using arrays of objects have demonstrated that some high-level processing of object semantics can occur in extra-foveal vision, but its role on the allocation of early overt attention is still unclear. This eye-tracking visual search study contributes novel findings by examining the role of object-to-object semantic relatedness and visual saliency on search responses and eye-movement behaviour across arrays of increasing size (3, 5, 7). Our data show that a critical object was looked at earlier and for longer when it was semantically unrelated than related to the other objects in the display, both when it was the search target (target-present trials) and when it was a target’s semantically related competitor (target-absent trials). Semantic relatedness effects manifested already during the very first fixation after array onset, were consistently found for increasing set sizes, and were independent of low-level visual saliency, which did not play any role. We conclude that object semantics can be extracted early in extra-foveal vision and capture overt attention from the very first fixation. These findings pose a challenge to models of visual attention which assume that overt attention is guided by the visual appearance of stimuli, rather than by their semantics.
During the visual search, cognitive control mechanisms activate to inhibit distracting information and efficiently orient attention towards contextually relevant regions likely to contain the search target. Cognitive ageing is known to hinder cognitive control mechanisms, however little is known about their interplay with contextual expectations, and their role in visual search. In two eye-tracking experiments, we compared the performance of a younger and an older group of participants searching for a target object varying in semantic consistency with the search scene (e.g., a basket of bread vs. a clothes iron in a restaurant scene) after being primed with contextual information either congruent or incongruent with it (e.g., a restaurant vs. a bathroom). Primes were administered either as scenes (Experiment 1) or words (Experiment 2, which included scrambled words as neutral primes). Participants also completed two inhibition tasks (Stroop and Flanker) to assess their cognitive control. Older adults had greater difficulty than younger adults when searching for inconsistent objects, especially when primed with congruent information (Experiment 1), or a scrambled word (neutral condition, Experiment 2). When the target object violates the semantics of the search context, congruent expectations or perceptual distractors, have to be suppressed through cognitive control, as they are irrelevant to the search. In fact, higher cognitive control, especially in older participants, was associated with better target detection in these more challenging conditions, although it did not influence eye-movement responses. These results shed new light on the links between cognitive control, contextual expectations and visual attention in healthy ageing.
Humans have the fascinating capacity of processing nonverbal visual cues to understand and anticipate the actions of other humans. This “intention reading” ability is underpinned by shared motor repertoires and action models, which we use to interpret the intentions of others as if they were our own. We investigate how different cues contribute to the legibility of human actions during interpersonal interactions. Our first contribution is a publicly available dataset with recordings of human body motion and eye gaze, acquired in an experimental scenario with an actor interacting with three subjects. From these data, we conducted a human study to analyze the importance of different nonverbal cues for action perception. As our second contribution, we used motion/gaze recordings to build a computational model describing the interaction between two persons. As a third contribution, we embedded this model in the controller of an iCub humanoid robot and conducted a second human study, in the same scenario with the robot as an actor, to validate the model's “intention reading” capability. Our results show that it is possible to model (nonverbal) signals exchanged by humans during interaction, and how to incorporate such a mechanism in robotic systems with the twin goal of being able to “read” human action intentionsand acting in a way that is legible by humans.
Background and aims Implicit learning mechanisms associated with detecting structural regularities have been proposed to underlie both the long-term acquisition of linguistic structure and a short-term tendency to repeat linguistic structure across sentences (structural priming) in typically developing children. Recent research has suggested that a deficit in such mechanisms may explain the inconsistent trajectory of language learning displayed by children with Developmental Learning Disorder. We used a structural priming paradigm to investigate whether a group of children with Developmental Learning Disorder showed impaired implicit learning of syntax (syntactic priming) following individual syntactic experiences, and the time course of any such effects. Methods Five- to six-year-old Italian-speaking children with Developmental Learning Disorder and typically developing age-matched and language-matched controls played a picture-description-matching game with an experimenter. The experimenter’s descriptions were systematically manipulated so that children were exposed to both active and passive structures, in a randomized order. We investigated whether children’s descriptions used the same abstract syntax (active or passive) as the experimenter had used on an immediately preceding turn (no-delay) or three turns earlier (delay). We further examined whether children’s syntactic production changed with increasing experience of passives within the experiment. Results Children with Developmental Learning Disorder’s syntactic production was influenced by the syntax of the experimenter’s descriptions in the same way as typically developing language-matched children, but showed a different pattern from typically developing age-matched children. Children with Developmental Learning Disorder were more likely to produce passive syntax immediately after hearing a passive sentence than an active sentence, but this tendency was smaller than in typically developing age-matched children. After two intervening sentences, children with Developmental Learning Disorder no longer showed a significant syntactic priming effect, whereas typically developing age-matched children did. None of the groups showed a significant effect of cumulative syntactic experience. Conclusions Children with Developmental Learning Disorder show a pattern of syntactic priming effects that is consistent with an impairment in implicit learning mechanisms that are associated with the detection and extraction of abstract structural regularities in linguistic input. Results suggest that this impairment involves reduced initial learning from each syntactic experience, rather than atypically rapid decay following intact initial learning. Implications Children with Developmental Learning Disorder may learn less from each linguistic experience than typically developing children, and so require more input to achieve the same learning outcome with respect to syntax. Structural priming is an effective technique for manipulating both input quality and quantity to determine precisely how Developmental Learning Disorder is related to language input, and to investigate how input tailored to take into account the cognitive profile of this population can be optimised in designing interventions.
When people communicate, they coordinate a wide range of linguistic and non‐linguistic behaviors. This process of coordination is called alignment, and it is assumed to be fundamental to successful communication. In this paper, we question this assumption and investigate whether disalignment is a more successful strategy in some cases. More specifically, we hypothesize that alignment correlates with task success only when communication is interactive. We present results from a spot‐the‐difference task in which dyads of interlocutors have to decide whether they are viewing the same scene or not. Interactivity was manipulated in three conditions by increasing the amount of information shared between interlocutors (no exchange of feedback, minimal feedback, full dialogue). We use recurrence quantification analysis to measure the alignment between the scan‐patterns of the interlocutors. We found that interlocutors who could not exchange feedback aligned their gaze more, and that increased gaze alignment correlated with decreased task success in this case. When feedback was possible, in contrast, interlocutors utilized it to better organize their joint search strategy by diversifying visual attention. This is evidenced by reduced overall alignment in the minimal feedback and full dialogue conditions. However, only the dyads engaged in a full dialogue increased their gaze alignment over time to achieve successful performances. These results suggest that alignment per se does not imply communicative success, as most models of dialogue assume. Rather, the effect of alignment depends on the type of alignment, on the goals of the task, and on the presence of feedback.
Long-term memory capacity for visual scenes is impressive. However, as shown in studies with young adults, the fidelity of the representation is degraded by semantic interference (SI): the higher the number of images from the same semantic category, the worse their retrieval. We investigated whether and how SI impacts long-term visual memory of healthy older participants and people in pre-dementia stage. 16 participants with a diagnosis of amnestic Mild Cognitive Impairment, a precursor of Alzheimer's Disease, (aMCI - 2 Female, 73.6±9.3 year) and 17 healthy age-matched control (12 Female, 68.9 ± 8.2 year) watched a stream of 129 naturalistic images drawn from 12 different scenarios (e.g., kitchen), each presented for 3 seconds, with an inter-stimulus interval of 500ms. SI was implemented by presenting a different number of scenes from each category (i.e., 24, 12, 6, 1). The participants were asked to verbally recognize the presented scenes in a 2 Alternative Force Choice (old/new) with both scenes belonging to the same semantic category. All participants had a performance significantly above chance under a binomial test, and no less than 60% per cent. Figure 1 reports the d-prime scores for both groups. Older adults can recognize a high number of scenes and their performance is influenced by SI. The two groups differed in overall performance, as expected. However, the decrement in performance as a function of the degree of SI is not different.
Expectancy mechanisms are routinely used by the cognitive system in stimulus processing and in anticipation of appropriate responses. Electrophysiology research has documented negative shifts of brain activity when expectancies are violated within a local stimulus context (e.g., reading an implausible word in a sentence) or more globally between consecutive stimuli (e.g., a narrative of images with an incongruent end). In this EEG study, we examine the interaction between expectancies operating at the level of stimulus plausibility and at more global level of contextual congruency to provide evidence for, or against, a disassociation of the underlying processing mechanisms. We asked participants to verify the congruency of pairs of cross-modal stimuli (a sentence and a scene), which varied in plausibility. ANOVAs on ERP amplitudes in selected windows of interest show that congruency violation has longer-lasting (from 100 to 500ms) and more widespread effects than plausibility violation (from 200 to 400ms). We also observed critical interactions between these factors, whereby incongruent and implausible pairs elicited stronger negative shifts than their congruent counterpart, both early on (100-200ms) and between 400-500ms. Our results suggest that the integration mechanisms are sensitive to both global and local effects of expectancy in a modality independent manner. Overall, we provide novel insights into the interdependence of expectancy during meaning integration of cross-modal stimuli in a verification task.
Human to human sensorimotor interaction can only be fully understood by modeling the patterns of bodily synchronization and reconstructing the underlying mechanisms of optimal cooperation. We designed a tower-building task to address such a goal. We recorded upper body kinematics of dyads and focused on the velocity profiles of the head and wrist. We applied recurrence quantification analysis to examine the dynamics of synchronization within, and across the experimental trials, to compare the roles of leader and follower. Our results show that the leader was more auto-recurrent than the follower to make his/her behavior more predictable. When looking at the cross-recurrence of the dyad, we find different patterns of synchronization for head and wrist motion. On the wrist, dyads synchronized at short lags, and such a pattern was weakly modulated within trials, and invariant across them. Head motion, instead, synchronized at longer lags and increased both within and between trials: a phenomenon mostly driven by the leader. Our findings point at a multilevel nature of human to human sensorimotor synchronization, and may provide an experimentally solid benchmark to identify the basic primitives of motion, which maximize behavioral coupling between humans and artificial agents.
The human sentence processor is able to make rapid predictions about upcoming linguistic input. For example, upon hearing the verb eat , anticipatory eye‐movements are launched toward edible objects in a visual scene (Altmann & Kamide, 1999). However, the cognitive mechanisms that underlie anticipation remain to be elucidated in ecologically valid contexts. Previous research has, in fact, mainly used clip‐art scenes and object arrays, raising the possibility that anticipatory eye‐movements are limited to displays containing a small number of objects in a visually impoverished context. In Experiment 1, we confirm that anticipation effects occur in real‐world scenes and investigate the mechanisms that underlie such anticipation. In particular, we demonstrate that real‐world scenes provide contextual information that anticipation can draw on: When the target object is not present in the scene, participants infer and fixate regions that are contextually appropriate (e.g., a table upon hearing eat ). Experiment 2 investigates whether such contextual inference requires the co‐presence of the scene, or whether memory representations can be utilized instead. The same real‐world scenes as in Experiment 1 are presented to participants, but the scene disappears before the sentence is heard. We find that anticipation occurs even when the screen is blank, including when contextual inference is required. We conclude that anticipatory language processing is able to draw upon global scene representations (such as scene type) to make contextual inferences. These findings are compatible with theories assuming contextual guidance, but posit a challenge for theories assuming object‐based visual indices.
We investigated the production of subject relative clauses (SRc) in Italian pre-school children with Specific Language Impairment (SLI) and age-matched typically-developing children (TD) controls. In a structural priming paradigm, children described pictures after hearing the experimenter produce a bare noun or an SRc description, as part of a picture matching task. In a sentence repetition task, children repeated SRc. In the priming paradigm, children with SLI produced SRc after hearing the experimenter use SRc with the same or different lexical content; the magnitude of this priming effect was the same as in TDC. However, children with SLI showed a smaller cumulative priming effect than TDC. Children with SLI showed superior SRc performance in picture-matching than in sentence repetition. We propose that children with SLI have an abstract representation of SRc that can be facilitated by prior exposure, but exhibit impaired implicit learning mechanisms.
Psycholinguistic research using the visual world paradigm has shown that the processing of sentences is constrained by the visual context in which they occur. Recently, there has been growing interest in the interactions observed when both language and vision provide relevant information during sentence processing. In three visual world experiments on syntactic ambiguity resolution, we investigate how visual and linguistic information influence the interpretation of ambiguous sentences. We hypothesize that (1) visual and linguistic information both constrain which interpretation is pursued by the sentence processor, and (2) the two types of information act upon the interpretation of the sentence at different points during processing. In Experiment 1, we show that visual saliency is utilized to anticipate the upcoming arguments of a verb. In Experiment 2, we operationalize linguistic saliency using intonational breaks and demonstrate that these give prominence to linguistic referents. These results confirm prediction (1). In Experiment 3, we manipulate visual and linguistic saliency together and find that both types of information are used, but at different points in the sentence, to incrementally update its current interpretation. This finding is consistent with prediction (2). Overall, our results suggest an adaptive processing architecture in which different types of information are used when they become available, optimizing different aspects of situated language processing.
This paper describes the R package crqa to perform cross-recurrence quantification analysis of two time series of either a categorical or continuous nature. Streams of behavioral information, from eye movements to linguistic elements, unfold over time. When two people interact, such as in conversation, they often adapt to each other, leading these behavioral levels to exhibit recurrent states. In dialog, for example, interlocutors adapt to each other by exchanging interactive cues: smiles, nods, gestures, choice of words, and so on. In order for us to capture closely the goings-on of dynamic interaction, and uncover the extent of coupling between two individuals, we need to quantify how much recurrence is taking place at these levels. Methods available in crqa would allow researchers in cognitive science to pose such questions as how much are two people recurrent at some level of analysis, what is the characteristic lag time for one person to maximally match another, or whether one person is leading another. First, we set the theoretical ground to understand the difference between "correlation" and "co-visitation" when comparing two time series, using an aggregative or cross-recurrence approach. Then, we describe more formally the principles of cross-recurrence, and show with the current package how to carry out analyses applying them. We end the paper by comparing computational efficiency, and results' consistency, of crqa R package, with the benchmark MATLAB toolbox crptoolbox (Marwan, 2013). We show perfect comparability between the two libraries on both levels.
The role of the task has received special attention in visual-cognition research because it can provide causal explanations of goal-directed eye-movement responses. The dependency between visual attention and task suggests that eye movements can be used to classify the task being performed. A recent study by Greene, Liu, and Wolfe (2012), however, fails to achieve accurate classification of visual tasks based on eye-movement features. In the present study, we hypothesize that tasks can be successfully classified when they differ with respect to the involvement of other cognitive domains, such as language processing. We extract the eye-movement features used by Greene et al. as well as additional features from the data of three different tasks: visual search, object naming, and scene description. First, we demonstrated that eye-movement responses make it possible to characterize the goals of these tasks. Then, we trained three different types of classifiers and predicted the task participants performed with an accuracy well above chance (a maximum of 88% for visual search). An analysis of the relative importance of features for classification accuracy reveals that just one feature, i.e., initiation time, is sufficient for above-chance performance (a maximum of 79% accuracy in object naming). Crucially, this feature is independent of task duration, which differs systematically across the three tasks we investigated. Overall, the best task classification performance was obtained with a set of seven features that included both spatial information (e.g., entropy of attention allocation) and temporal components (e.g., total fixation on objects) of the eye-movement record. This result confirms the task-dependent allocation of visual attention and extends previous work by showing that task classification is possible when tasks differ in the cognitive processes involved (purely visual tasks such as search vs. communicative tasks such as scene description).
An ongoing issue in visual cognition concerns the roles played by low- and high-level information in guiding visual attention, with current research remaining inconclusive about the interaction between the two. In this study, we bring fresh evidence into this long-standing debate by investigating visual saliency and contextual congruency during object naming (Experiment 1), a task in which visual processing interacts with language processing. We then compare the results of this experiment to data of a memorization task using the same stimuli (Experiment 2). In Experiment 1, we find that both saliency and congruency influence visual and naming responses and interact with linguistic factors. In particular, incongruent objects are fixated later and less often than congruent ones. However, saliency is a significant predictor of object naming, with salient objects being named earlier in a trial. Furthermore, the saliency and congruency of a named object interact with the lexical frequency of the associated word and mediate the time-course of fixations at naming. In Experiment 2, we find a similar overall pattern in the eye-movement responses, but only the congruency of the target is a significant predictor, with incongruent targets fixated less often than congruent targets. Crucially, this finding contrasts with claims in the literature that incongruent objects are more informative than congruent objects by deviating from scene context and hence need a longer processing. Overall, this study suggests that different sources of information are interactively used to guide visual attention on the targets to be named and raises new questions for existing theories of visual attention.
Object detection and identification are fundamental to human vision, and there is mounting evidence that objects guide the allocation of visual attention. However, the role of objects in tasks involving multiple modalities is less clear. To address this question, we investigate object naming, a task in which participants have to verbally identify objects they see in photorealistic scenes. We report an eye-tracking study that investigates which features (attentional, visual, and linguistic) influence object naming. We find that the amount of visual attention directed toward an object, its position and saliency, along with linguistic factors such as word frequency, animacy, and semantic proximity, significantly influence whether the object will be named or not. We then ask how features from different modalities are combined during naming, and find significant interactions between saliency and position, saliency and linguistic features, and attention and position. We conclude that when the cognitive system performs tasks such as object naming, it uses input from one modality to constraint or enhance the processing of other modalities, rather than processing each input modality independently.
Most everyday tasks involve multiple modalities, which raises the question of how the processing of these modalities is coordinated by the cognitive system. In this paper, we focus on the coordination of visual attention and linguistic processing during speaking. Previous research has shown that objects in a visual scene are fixated before they are mentioned, leading us to hypothesize that the scan pattern of a participant can be used to predict what he or she will say. We test this hypothesis using a data set of cued scene descriptions of photo‐realistic scenes. We demonstrate that similar scan patterns are correlated with similar sentences, within and between visual scenes; and that this correlation holds for three phases of the language production process (target identification, sentence planning, and speaking). We also present a simple algorithm that uses scan patterns to accurately predict associated sentences by utilizing similarity‐based retrieval.