Capo
So funktioniert es

Why slowed audio sounds underwater

Elyra-Labs-Team4 Min. LesezeitAktualisiert am 21. Aug. 2026
Die kurze Antwort

Slowing a recording the old way drags the pitch down with it, because you are playing the same waveform more slowly. Keeping the pitch fixed means rebuilding the sound from short overlapping grains, and when those grains do not line up you hear smearing, a metallic edge, or a voice that sounds submerged.

  1. Play the demo below in Varispeed. Pitch falls with speed.
  2. Switch to Independent. Speed moves, pitch does not.
  3. The difference is what costs the artefacts.

“Slow it down” describes two different operations. They sound different, and most questions about pitch shifting come from expecting one and getting the other.

Hear itVarispeed or independent — hear the difference

Two ways of slowing the same tone. Varispeed is tape: the pitch falls with the speed. Independent is what Capo does by default: the speed moves and the pitch stays. Move the speed slider in each mode and watch the resulting pitch.

Speed1.00× · 0 st

Synthesised live in your browser with the Web Audio API — the same technology Capo uses. Nothing is downloaded.

Varispeed: the honest one

Play a record at half speed and every frequency in it halves. Pitch is how often a waveform repeats per second, so halving the playback rate halves the repetition rate: an octave lower.

Tape does this. So does a turntable, and so does changing the sample rate. It sounds natural because it is what physically happens — and it is no use for practising along, because the recording is now in a different key from your instrument.

Time-stretching: the useful one

Holding the pitch steady means rebuilding the audio. The engine cuts it into grains a few tens of milliseconds long, then re-schedules them: spaced further apart to slow down, overlapped to speed up. Each grain still plays at its original rate, so the pitch stays put. Only the timing changes.

The grain boundaries are where it goes wrong. At every join, two pieces of waveform meet that were never adjacent; if they meet out of phase the result is a click, a smear or a faint ring. That happens thousands of times a second, and the accumulation is what people hear as watery, phasey or underwater.

Why it gets worse at extremes

At 0.9× the grains barely have to move, and the artefacts are inaudible. At 0.25×, each grain has to be stretched or repeated four times over, so there are four times as many joins and each one is under more strain. That is why the floor is 0.25× and not lower: past that point the errors stop being artefacts and start being the dominant sound.

SpeedWhat you hearUseful for
0.9–1.0×Nothing. It is transparent.Comfortable listening
0.6–0.8×Clean on most materialLearning a part
0.4–0.6×Slight smear on cymbals and sustained vocalsPicking apart a fast run
0.25–0.4×Audible watery qualityIdentifying one specific event

Which is why there is a quality setting

Better grain placement costs processing. Capo ships three time-stretch engines and the Audio quality setting picks between them: the lightest is cheap and smears sooner, the heaviest holds together further down. Below about 0.5× the difference is obvious; above 0.8× you will struggle to hear it at all.

Settings · audio & playback
Audio & playback in Settings: Audio quality, Skip interval and Count-in duration
Audio & playback in Settings. Low and Med are free; High carries a Pro badge.
1Audio & playback
1aAudio quality1bSkip interval1cCount-in duration

The reasoning behind having three rather than one is in why there are three engines and not one.

Voices are the hard case

Voices suffer more than instruments. A voice carries resonances set by the size of the throat and mouth — formants — and listeners are sensitive to them because they encode how big the speaker is. Smear those and the change registers immediately, even when the listener cannot name it.

That sensitivity is also why transposing a voice needs a separate correction — covered in singing a song that sits four semitones too high.

Hear the difference for yourself.

Three engines, switchable · free, no account
Zu Chrome hinzufügen – gratis

Two things this does not settle

It does not explain why some recordings survive slowing far better than others. Dense, loud, heavily limited masters have less room between transients and tend to smear earlier than sparse acoustic recordings, but that is a rule of thumb rather than a rule.

It also does not cover pitch shifting, which uses the same grain machinery in the opposite direction and has its own failure modes.

Fragen

Why does slowing a video normally lower the pitch?
Because pitch is how often a waveform repeats per second. Play it more slowly and it repeats less often, which is a lower note. That is varispeed.
How does Capo slow audio without changing the key?
It cuts the audio into short grains and re-schedules them, so each grain still plays at its original rate. Only the timing changes.
Why does heavily slowed audio sound watery?
Every grain boundary joins two pieces of waveform that were never adjacent. At 0.25× there are four times as many joins as at full speed, and the small errors accumulate.
What is the lowest usable speed?
0.25× is the floor. Between 0.4× and 0.6× you may hear slight smearing on cymbals and sustained vocals; above 0.8× it is effectively transparent.

Weiterlesen

Setz einen Capo auf alles, was du hörst.

Zu Chrome hinzufügen – gratisAuch für Firefox & Edge · v2.7.6 · kein Konto