SonicSenses

Sound, vision & multisensory perception

Visualizing Sound: How the Brain Integrates What We Hear and See

How hearing and vision influence each other, when the brain treats two signals as one event, and why watching sound being visualised is an experience rather than a treatment.

12 min read

The short answer

Hearing and vision are processed by separate systems that influence each other constantly. When an auditory and a visual signal arrive close enough in time and space, and the context makes it plausible that they came from the same event, perception tends to combine them, and the more reliable signal usually dominates. This is well demonstrated in humans - in audiovisual speech, in illusions where what you see changes what you hear, and in synchrony judgements. Integration is not automatic or universal, though. It depends on timing, spatial layout, attention, prior expectation and the reliability of each signal, and the size of these effects varies considerably between people. A momentary audiovisual interaction is also not evidence of durable neural reorganisation.

Why this matters for sound and music

SonicSenses turns analysed features of audio into moving visuals, so the honest question is what pairing sound with a picture actually does perceptually. The answer is interesting and specific, and it stops well short of the cognitive and therapeutic claims usually attached to audiovisual products.

What multisensory integration means, precisely

Hearing and vision start out as separate physical measurements: pressure changes at the eardrum, light at the retina. They are processed by dedicated pathways, and they also converge - in subcortical structures such as the superior colliculus, in association cortex, and through direct influences on early sensory areas. The single-neuron work that grounds this field showed that responses to a combined audiovisual event are not simply the sum of the two unimodal responses, and that whether they enhance or suppress depends on the timing and spatial relationship between them.

It is tempting to summarise this as 'the brain combines sound and vision into one unified signal'. That is too strong. Integration happens at multiple levels and in a graded, situation-dependent way. Sometimes signals are bound into one perceived event, sometimes one modality biases the other slightly, and sometimes they are kept separate because the brain treats them as unrelated.

A useful framing from human psychophysics is reliability weighting. When two senses report on the same property, observers behave roughly as though they combine the estimates in proportion to how precise each one is. Vision usually wins for spatial judgements; hearing usually wins for fine timing. That single idea explains a large amount of otherwise confusing evidence, though the precise fit to any given task is an active research question.

  • Temporal alignment

    How close in time the two signals arrive - the strongest single cue for binding.

  • Spatial alignment

    Whether they appear to come from the same place.

  • Reliability

    How precise each signal is; the more precise one carries more weight.

  • Attention and context

    What the observer is attending to, and whether one event plausibly produced both signals.

Where vision changes what you hear

The best-known demonstration is audiovisual speech. When a seen mouth movement is dubbed with a different spoken syllable, many listeners report hearing a third syllable that matches neither input. The effect is genuine and has been replicated for decades. It is also far more variable than popular retellings admit: susceptibility differs strongly across individuals, stimuli, syllables and languages, so it should be described as a robust phenomenon in some conditions rather than as something that happens to everyone.

A second class of demonstrations runs the other way in kind but the same way in principle. In the sound-induced flash illusion, a single brief flash accompanied by two rapid beeps is often perceived as two flashes, which is a case of hearing biasing vision on a timing judgement - exactly the direction reliability weighting predicts.

What none of this shows is that any visual stimulus alters auditory perception. These effects appear under specific conditions: close timing, plausible common cause, and a task where the influencing modality is the more reliable one. Arbitrary visuals paired with arbitrary audio mostly do not produce them.

  • Audiovisual speech influence: well demonstrated, highly variable between people and languages.
  • Sound-induced flash illusion: well replicated, timing-specific.
  • Spatial capture of sound by vision: well documented for co-located, plausibly linked events.
  • Any visual accompaniment changing what you hear: not supported.

Timing: binding windows are ranges, not thresholds

Audiovisual events in the world rarely arrive at the senses simultaneously. Sound travels slowly compared with light, and neural transduction times differ, so the perceptual system tolerates a spread of small asynchronies and still treats the signals as one event. Researchers call this tolerance the temporal binding window.

It is important not to quote one universal millisecond figure. Reported windows differ by task - simultaneity judgement, temporal-order judgement, or susceptibility to an illusion - by stimulus complexity, with speech tolerating larger asynchronies than simple flashes and beeps, and by individual. Reviews of the construct emphasise exactly this methodological dependence, and also note that the window has been studied as an individual-difference measure in several clinical populations without that establishing a treatment target.

One genuinely training-based finding belongs here. In a laboratory study, adults given repeated feedback on simultaneity judgements showed a narrowed binding window that persisted for at least a week. That is perceptual learning from an explicit, feedback-driven task - not something demonstrated for passively watching audio-reactive visuals.

  • Simultaneity judgement

    Asking whether two signals occurred at the same time.

  • Temporal-order judgement

    Asking which came first - a different task with different numbers.

  • Speech versus simple stimuli

    Speech tolerates much larger asynchronies than flashes and beeps.

  • Individual variation

    Windows differ substantially between people, which is why single thresholds mislead.

Cross-modal correspondences: consistent, but not laws

Separately from full integration, people show consistent associations between features in different senses. Higher pitches are matched with higher spatial positions, with smaller objects, with brighter colours and with sharper shapes; louder sounds with larger size; faster tempos with faster visual motion. These correspondences are measurable in reaction-time and matching tasks and appear early in development for some pairings.

The tutorial literature is careful about their origins. Some correspondences plausibly reflect statistical regularities in the environment, some reflect shared magnitude coding, and some reflect language and culture. Not all of them generalise across cultures, and the ones that do are still tendencies rather than fixed biological mappings. Presenting a pitch-to-colour mapping as a universal law would misrepresent the evidence.

This matters directly for any sound visualisation. Design choices that follow common correspondences - bright, high-placed, small-scale visuals for high-frequency energy, for example - tend to feel congruent to many viewers. That is a design and aesthetics finding, not a claim about a correct or true visual form for a sound.

What SonicSenses actually does with audio

Describing the product accurately matters more here than anywhere else on the site. The SonicSenses visualizer runs the incoming audio through a Web Audio analyser and derives a small set of measurable features each frame: overall level from the waveform, energy in low, mid and high frequency bands, a spectral centroid used as a brightness measure, onset and transient detection with an estimated beat rate, and a stereo balance measure taken from separate left and right analysers where the source allows it. Those values drive the visual parameters of the selected world.

So there are three distinct things, and conflating them is where scientific credibility is usually lost. There is the sound itself, a physical pressure wave. There is the data extracted from that sound, a handful of numeric features computed per frame. And there is the visual representation, a designed rendering driven by those numbers.

SonicSenses translates analysed features of sound into a changing visual representation, giving users an experiential way to explore relationships between auditory structure and visual form. A SonicSenses image is not a brain scan, not a representation of neural activity, and not a literal picture of sound waves travelling through the brain or the body.

  • Extracted features: level, low/mid/high band energy, spectral brightness, onsets and beat rate, stereo balance.
  • Not extracted or displayed: neural signals, physiological state, emotional content, or any clinical measure.
  • Purpose: experience and education about audio structure, not assessment or treatment.

Interaction now versus durable change later

The most common overstatement in this area is treating a perceptual effect as evidence of rewiring. It is worth separating four things. An immediate multisensory interaction happens while the stimuli are present and stops when they stop. Adaptation is a short-lived recalibration after sustained exposure, such as shifting perceived simultaneity after repeated exposure to a fixed audiovisual lag. Perceptual learning is a lasting improvement in discrimination, and typically requires attention, repetition and feedback. Durable neuroplasticity is structural or long-term functional reorganisation, which needs longitudinal evidence to claim.

Audiovisual research is dominated by the first two categories. Studies that reach the third are explicitly training studies with defined tasks and feedback. None of the standard demonstrations - the speech illusion, the flash illusion, cross-modal matching - are evidence of the fourth.

Applied to SonicSenses, that puts the honest claim in a narrow, defensible place. Watching sound rendered visually is engaging, and it can make audio structure more noticeable. It is not a demonstrated route to improved sensory integration, cognition or creativity, and we do not claim it is.

What we know

  • Auditory and visual signals influence each other, and combined responses are not simply the sum of unimodal ones.
  • Whether signals are bound depends on timing, spatial layout, plausibility of a common cause, attention and reliability.
  • Vision tends to dominate spatial judgements; hearing tends to dominate fine timing judgements.
  • Temporal binding windows vary by task, stimulus type and individual.
  • Cross-modal correspondences between pitch, brightness, size, shape and motion are consistently measurable in the general population.

What remains uncertain

  • How well reliability-weighting models describe integration outside simple laboratory tasks.
  • How much of any given cross-modal correspondence is environmental statistics, shared magnitude coding, or language and culture.
  • Why McGurk susceptibility varies so widely across individuals and languages.
  • Whether laboratory narrowing of the binding window transfers to any everyday task.

What this does not prove

  • That combining sound and visuals rewires the brain.
  • That audio-reactive visuals improve cognition, attention or sensory integration.
  • That audiovisual stimulation treats neurological or sensory disorders.
  • That the brain simply merges hearing and vision into one unified signal.
  • That any visual accompaniment changes what a listener hears.

Practical meaning

  • Expect visuals to change how you notice audio structure, not to change your hearing.
  • Congruent mappings - brightness with high-frequency energy, motion with tempo - tend to feel right because of common correspondences, not because they are correct.
  • Tight synchrony matters for the sense that sound and image belong together; large lags break it.
  • Treat any product claiming that watching visuals rewires or heals as making a claim this literature does not support.

Frequently asked questions

Does the brain combine sound and vision into one signal?
Not as a simple rule. Auditory and visual signals interact constantly, and under the right conditions - close timing, plausible common cause, compatible spatial layout - they are perceived as one event. In other conditions they stay separate, or one merely biases the other. Integration is graded and context-dependent rather than automatic.
Is the McGurk effect real?
Yes, and it has been replicated since 1976. But it is not universal. Susceptibility varies substantially by individual, syllable, talker and language, so it is best described as a robust demonstration that vision can influence speech perception rather than as something everyone experiences.
How closely do sound and visuals need to be synchronised?
There is no single number. Tolerance for audiovisual asynchrony depends on the task, the type of stimulus - speech tolerates far more than beeps and flashes - and the individual. Reviews of the temporal binding window emphasise that reported values are method-dependent, so quoting one universal millisecond threshold is misleading.
Does watching a music visualizer change your brain?
There is no evidence for that. Audiovisual research mostly documents immediate perceptual interactions and short-lived adaptation. Studies showing lasting perceptual change involve explicit training with feedback. SonicSenses is an experiential and educational tool, not a treatment or a training protocol.
Is a SonicSenses visual a picture of the sound, or of the brain?
Neither. It is a designed rendering of a small set of features measured from the audio - level, frequency-band energy, spectral brightness, onsets and beat rate, and stereo balance. It shows nothing about neural activity.

References & further reading

  1. Stein, B. E., & Stanford, T. R. (2008). Multisensory integration: current issues from the perspective of the single neuron. Nature Reviews Neuroscience DOI: 10.1038/nrn2331
  2. Ernst, M. O., & Banks, M. S. (2002). Humans integrate visual and haptic information in a statistically optimal fashion. Nature DOI: 10.1038/415429a
  3. McGurk, H., & MacDonald, J. (1976). Hearing lips and seeing voices. Nature DOI: 10.1038/264746a0
  4. Shams, L., Kamitani, Y., & Shimojo, S. (2000). What you see is what you hear. Nature DOI: 10.1038/35048669
  5. Vroomen, J., & Keetels, M. (2010). Perception of intersensory synchrony: a tutorial review. Attention, Perception, & Psychophysics DOI: 10.3758/APP.72.4.871
  6. Wallace, M. T., & Stevenson, R. A. (2014). The construct of the multisensory temporal binding window and its dysregulation in developmental disabilities. Neuropsychologia DOI: 10.1016/j.neuropsychologia.2014.08.005
  7. Powers, A. R., Hillock, A. R., & Wallace, M. T. (2009). Perceptual training narrows the temporal window of multisensory binding. The Journal of Neuroscience DOI: 10.1523/jneurosci.3501-09.2009
  8. Talsma, D., Senkowski, D., Soto-Faraco, S., & Woldorff, M. G. (2010). The multifaceted interplay between attention and multisensory integration. Trends in Cognitive Sciences DOI: 10.1016/j.tics.2010.06.008
  9. Spence, C. (2011). Crossmodal correspondences: a tutorial review. Attention, Perception, & Psychophysics DOI: 10.3758/s13414-010-0073-7

This article is an educational summary of publicly available research and is not medical advice. It does not diagnose, treat, or cure any medical or psychiatric condition. Where evidence is emerging or mixed, we say so. Consult a qualified professional for personal guidance.