San Francisco State University

Computer Science
DOD funds speech recognition project

Reproduced with permission from NeXT Computer, Inc.
A Reference Guide to NeXT in Higher Education, Fall 1992
ยช 1992 NeXT Computer, Inc.

Thomas Holton, associate professor of engineering, and two undergraduate computer science students, Joel Miller and Lee Worden, are using models of auditory signal processing to develop improved algorithms for the computer recognition of speech. The entire project, which is funded by the Department of Defense, is taking place on NeXT machines, using the graphics and DSP functions extensively.

"The purpose of our work," says Holton, "is to understand the fundamental strategy used by the auditory system to process speech signals and to apply this understanding to the design of improved algorithms for robust computer recognition of speech features."

According to Holton, "My background is in hearing, experimental neuropsychology and electrical engineering, and this project combines my interests in these areas. The project also reflects my philosophical feeling that speech is a fundamentally human activity, and that current engineering approaches to speech recognition-e.g. spectrographic approaches-do not process speech the way humans do. In fact, our approach shows significant promise compared to the conventional spectrographically based algorithms."

In collaboration with colleagues Steve Love and Steve Gill of Votan Corporation, Holton has developed a comprehensive model of peripheral and early central auditory signal processing. The components include a detailed three-dimensional hydro-mechanical model of the cochlea, which describes the fluid-wave processes underlying basilar-membrane motion, a biophysical description of mechano-electric transduction in cochlear hair cells, a description of the time-dependent synaptic chemistry of hair cells and auditory-nerve fibers, and a `micro-neural-net' description of neural signal processing in the cochlear nucleus.

"A comparison of the predictions of the model with experimental physiological data obtained in response to simple (tonal) and complex (speech) stimuli suggests that the model accurately describes essential features of auditory signal processing," says Holton.

The researcher says he purchased the NeXT machines for the project because they had "the right mix of features at a reasonable price."

"The auditory-model algorithms that form the core of our approach are highly DSP-intensive, and the DSP56001 co-processor in the NeXT was initially a strong selling point," he explains. "The applications we've written also require sound recording, editing and playback, and high-resolution displays. The PostScript display capability of the NeXT machine, as well as NeXT's development environment and set of supplied objects (e.g. SoundKit) have been instrumental in helping us get our applications up and running fast, and helping us maintain them."

For more information, please contact:

Thomas Holton
Associate Professor of Engineering
San Francisco State University
1600 Holloway Avenue
San Francisco, CA 94132
(415) 338-1529
th@ernie.sfsu.edu