Experiments 22
Any Semblance of My Speech
In 1969 Alvin Lucier recorded himself speaking, played the tape back into the room, recorded that, and kept going until his words had dissolved into the room’s own ringing. Here the room is a machine. It hears your sentence, learns your voice from it and says it back. Then it listens only to itself, again and again.
Rust · Candle · WebSocket · WebRTC · Opus · WhisperView source
Example · The readerVoice 100%Words 14 of 14
I am the master of my fate. I am the captain of my soul.
RecordUp to 8 seconds
Record one sentence, or turn Sound on to hear this example.
Notes
The piece this borrows
In 1969 the composer Alvin Lucier read a short text into a tape recorder, played the tape back through a loudspeaker into the same room, and recorded that. Then he did it again, and again. Every room rings at certain pitches, and each pass made those pitches a little louder and everything else a little quieter. After enough passes the words were gone and the room was playing itself, in the rhythm of his sentences.
The text he read describes the process. The title here is from it: he says he will go on “until the resonant frequencies of the room reinforce themselves so that any semblance of my speech, with perhaps the exception of rhythm, is destroyed.” The piece is I Am Sitting in a Room.
What a generation is here
There is no room and no loudspeaker. There are two models, and each generation runs three steps on the recording before it and on nothing else.
A speech recognizer writes down the words. A voice model studies the same recording and builds a voice from it. Then it says those words in that voice, and that becomes the recording. Your own take is used exactly once. From generation two on, the machine is listening only to itself.
What goes first
LibriVox once had fourteen volunteers read the same poem, William Ernest Henley’s “Invictus.” I took the last two lines from each reading, “I am the master of my fate, I am the captain of my soul,” and gave every voice 36 generations.
The voice goes first, and it goes smoothly. Scored against the original reader, likeness averaged 0.92 after one generation, 0.52 after five and 0.31 after ten. All fourteen were below half by generation nine, most of them by six.
The words last longer, because the recognizer copes well with a clean voice even when it is no longer yours. The first changed word came at generation eight on the median, as early as three and as late as twelve. Each slip is then kept by every later generation. In thirteen of the fourteen runs “fate” became “feet.”
None of them still said the line at generation 36. One settled on “I am the master of my speech, I am the captain of my soul” and held it for more than twenty generations. Others ended at “I’m the king.”, “Thank you.”, “Okay.”, “I know.” and “Shh.” Short stock phrases like those are what this recognizer falls back on when a recording is too far gone to make out, and once a run reaches one it rarely leaves.
The recorded run on this page is one of the fourteen, read by Ian King. It passes through “I don’t know the master of my feet. I don’t know the captain of my soul.” and ends at “I wonder what it is.”
Do the voices all turn into one voice? Partly. Two different readers start at a likeness of 0.26 to each other. By generation seven their copies score 0.54 against each other, which is closer than the 0.40 each scores against the person it started as. After about generation eighteen they drift apart again.
Reading the panel
Voice is how close each generation’s voiceprint is to yours, as a percentage, measured by the voice model’s own speaker encoder. Words is how many of the words it first heard from you are still there, in order. Words in red are not yours.
Sound is the average spectrum of each generation, low pitches at the bottom. Watch for thin bright bands that appear after a few generations and stay. Those are pitches the voice model’s output favors, and since each generation learns from the last, they get reinforced. They are the nearest thing here to Lucier’s room.
The line
Direct puts nothing between one generation and the next. The other two send each recording through a bad phone call first, using Opus, the codec most voice and video calls run on. Patchy drops a fifth of the packets, two at a time. Storm drops them three at a time and adds hiss. On a test set the recognizer got about one word in ten wrong per generation on the patchy line and about one in four in the storm, so the words go much faster.
Your voice
Sound travels both ways over the same encrypted socket that carries the text. The server can also carry it over WebRTC, the way a video call does, but that is switched off for now. You can listen to the recorded run without a microphone. To make your own, press the disc, wait for the count and speak. The microphone is live only from the end of the count until you press the disc again or eight seconds pass. The server keeps your recordings in memory for the visit so you can scrub back through them, and drops them when you leave. Nothing is written to disk.
The copy of your voice can only repeat what the recognizer heard in the recording before it. There is nowhere to type words for it to say, and a take has to sound like one person from start to finish or it is refused.
Each run is repeatable down to the sample for the same recording and the same build of the server. Change the arithmetic slightly, for instance by running on a different processor, and the run follows the same path for about seven generations and then goes somewhere else.
Built with
- Hearing is Whisper, the tiny English model, running in Rust on Candle.
- The voice uses the Sopro v2 turbo weights by Samuel Vitorino, Apache-2.0, run by a Rust implementation written for this page and checked against his reference.
- The recordings are kept as Opus. The optional call is webrtc-rs.
- The readings are from LibriVox’s weekly poem for May 14, 2023, fourteen recordings of “Invictus”, all in the public domain.