Home

I thought to share this short essay “The space between silence: Human interactions beyond spoken words” that I wrote as an invited paper for the conference on Human Interaction in the Age of the Internet.

The conference was organized by the Jane and Aatos Erkko Visiting Professor Kevin Durrheim with the Helsinki Collegium for Advanced Studies this week (May 11–12, 2026) at University of Helsinki. More about the conference: https://blogs.helsinki.fi/erkko-conference2026/

It was an awesome interdisciplinary gathering of scholars across social psychology, communication, linguistics, anthropology, and computational research to explore how digital infrastructures are reshaping social life.

I’d love any thoughts or feedback as I develop this into a full paper or proposal for a new (collaborative) research project in the near future.

Brief excerpts: 

How do we make sense of the nature of human interactions that go beyond spoken words expressed in conversations and written articulation? These include Paralinguistic cues, the non-verbal vocal elements of speech, Kinesics, facial expressions and gestures, and Proxemics, how humans use physical space, personal distance and body position for implicit or explicit communicative purposes.

Such non-spoken gestures and silent interactions are rich forms of human expression; they are rarely if at all captured by machine learning systems or natural language models. This richness of communication embodies a great deal more than approval, disapproval or transmission of simplistic messages during extended interactions among individuals or groups, but encapsulate far more socio-cultural and contextual meaning that shapes our situational understanding and the affect of intended or unintended expressions.

All these embodied, gestural, place-based, and situationally-defined facets, engaging paralinguistic, kinesics and proxemic phenomenon, greatly shape human interactions and communication both implicitly and explicitly. However, these are almost entirely ignored or unaccounted for in most of our digital interactions, particularly with and by AI-based conversational systems (chatbots) and embodied AI systems including advanced robotics.

Today’s AI systems remain far more simplistic despite the rapid evolution of large language models (LLMs) which are primarily trained on textual archives, conversational interactions or multi-modal data while having limited embedded perception or situational context, to learn and engage with the richness of human non-vocal paralingusitic cues, kinesics and proxemics that even infants perceive, understand and manipulate tactfully.

My provocation is not about making better machine learning systems, but to consider what makes us truly human and the deeper challenges of computational understanding, while humans continue to evolve their expressive communication gestures shaped by sociocultural, embodied and cognitive learning to better perceive, enrich and obfuscate the relational codes and meanings in their interactions over time.