top of page

Training Robots to Read Emotions

By Aarav Shah,

Technology Columnist; The Lawrenceville School, NJ


Both robots and artificial intelligence have come a long way, developing from returning speech input manually to displaying dexterity and other physical exercises. Yet, with the question of robots' feasibility in the workplace and the potential replacement of human jobs, one thing is certain. If robots are heightened to initiate this major turning point in the world's history, then they need to be able to analyze and empathize with humans' facial expressions and internal emotions. Take, for reference, a robot acting as a waiter or waitress at a restaurant. The robot's job would be to pick up food from the kitchen, drop it off at the correct table, and continue that cycle. Yet, what if the robot made an error and dropped off the food at the wrong table, and the customer displayed a rather annoyed, angry, or even confused facial expression? As of today's current model, the robot would not be able to empathize with the customer, understand the unforeseen frustration, and bring the dish to the correct table. 



With an objective to improve robots' ability to understand humans' emotions, researchers led by Seung Chan Hong recently designed a study that trained collaborative robots to diagnose a human's feelings with the help of contextual factors. Forty volunteers were involved in the study, and the outcome was that robots could only go to a certain extent in managing to correctly understand the feelings of humans based on their interpretation of the available context. This inadequate capability demonstrated by the robots encouraged Seung Chan Hong to pursue this topic, further training robots to effectively interpret human emotions through the usage of a visual language model (VLM).


A visual language model is a large language database that has the ability to take in the input of visual images, hence its name. To test whether the VLM was successful in training the robots, the research team had participants of the study watch videos of robots responding to human behavior and emotions, as well as subsequently judge the robot's diagnosis of the humans' feelings. In the end, the humans were able to take more context into consideration and provide countless alternatives for why the human might have been acting in that certain way. Important to note, though, this second experiment had been a large improvement from conventional AI systems, for this VLM-trained robot achieved a score of 0.86 whereas the traditional AI-trained robot received a score of 0.77. This was based off of a scale from 0 to 1, the higher the number demonstrating that the robot's diagnosis was more closely related to the human's interpretation. In yet another follow-up experiment, Seung Chan Hong had forty volunteers interact with a VLM-trained robot. The goal of the robot was to purposefully make an error and respond in a way such that the concerns of the human were addressed at that moment. Out of the forty participants part of this specific portion of the multi-phase study, thirty-one participants agreed that this empathetic response was far better than a standardized message. Even then, in a later study, many participants expressed their unhappiness of the robot's inability to perform the task to begin with. 


There are many improvements to be made in robots before they enter the mainstream workplace and social environment, yet this experiment was one step towards attempting to improve their outward display of empathy.


Works Cited 


Hampson, Michelle. (2026, June 13). Visual Language Models Train Robots to Read Human Emotions. Retrieved from 


Nvidia. (n.d). What Are Vision Language Models? Retrieved from 

bottom of page