CompCog: Deep causal inference grounds the perception of cognitive objects in speech
INSTITUTION
University of Southern California, CA
PRINCIPAL INVESTIGATOR
Dani Byrd
FUNDING
$600K
YEAR
2023
MOONBASE SCORE
Still being scored
LOADING MOONBASE SCORE
Abstract
Artificial Intelligence systems are becoming more and more important in society, and their performance has improved enormously in recent years. Yet, we still do not understand how these systems actually work, and how they emulate human performance, to the extent that they do. In this work, novel methods are developed for probing the inner working of these systems by comparing their internal computations with corresponding computations humans perform. The specific skill we probe is speech recognition—a highly complex process, as speech is a richly variable, information dense, and quickly transmitted medium of communication. One of the ways that the human brain deals with this complexity during speech recognition is by engaging not only brain areas responsible for listening but also areas crucial to the production of speech. This suggests that human cognition is aware of the systems in the world that cause speech—the movements of the lips, tongue, and other vocal organs. Do current artificial intelligence systems also develop such a deep causal understanding of speech? In this work we answer this question by delving into the mathematical models of both human and machine knowledge in these systems. Our work has two major goals. The first is technological: understanding how artificial intelligence systems actually work on the inside, which is ultimately a necessary step in directing their abilities to societal benefit. The second is scientific: before widespread use in society, artificial intelligence systems were developed by cognitive scientists to understand human cognition, and by probing the inner workings of state-of-the-art machine learning as a cognitive model we may be able to better understand how humans perceive speech. In this work, therefore, science and technology further each other, as they have done successfully in the past. This research program specifically probes the relationship between the production and perception of speech in humans and computers. To do so, speakers of three languages (English, Russian, Korean) are imaged using a real time Magnetic Resonance Imaging (MRI), which shows in vivid detail how the speech articulators move. Speaker’s speech audio signals are recorded simultaneously. The data are analyzed using mathematical models of speech production, modern speech recognition systems, and mathematical models of how human neural rhythms analyze speech. Experimental manipulations unveil how the representations in each of the systems corresponds to those in the others. This strategy inform us about the science of human cognition hand-in-hand with illuminating the black-box technology of machine emulation of the human capability. In the future, in addition to advancing science and technology, we anticipate the application of this knowledge to the creation of novel small-sized speech recognition systems that can assist in the documentation of endangered languages. This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
Are you the primary organization running this research?
The two tools below are built for the principal investigator & host institution behind this project.