In a realistic out-of-dataset setting, say we retrieve a shakuhachi track and a humpback whale track from external libraries:
| Condition | Audio |
|---|---|
| Source S | ![]() |
| Reference R | ![]() |
With Foley or creative intention, we want the shakuhachi to sing like a whale: it should follow the whale's temporal pattern while still sounding like authentic shakuhachi. We compare the following four methods, where the neural generator (pretrained Stable Audio 3) also receives the prompt shakuhachi sound effect.
| Condition | pre-neural synth | neural output |
|---|---|---|
| Raw reference R | ![]() | ![]() |
| Raw R + S mixture | ![]() | ![]() |
| Pointwise concatenative | ![]() | ![]() |
| FO (relational) | ![]() | ![]() |
As we can hear:
There is also the known out-of-dataset robustness issue, prevalent in modern AI generative models, where artifacts and unwanted acoustic-identity modifications become much more common under out-of-dataset conditions, as can be heard across all rows. This is, however, outside the scope of this paper/demo, as we focus on the capabilities of pre-neural synthesis and RAG conditioning.
In this section, we provide samples for the ablation and model-comparison section. We obtain sources and targets from ESC-50 and references from ESC-50-Voice, with the references imitating the targets.
| Condition | CRG example 1: car horn | CRG example 2: cat | FO example 1: sheep | FO example 2: cow | FO example 3: frog | VMO example 1: siren | VMO example 2: hen | VMO example 3: clock alarm | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| pre-neural synth | neural output | pre-neural synth | neural output | pre-neural synth | neural output | pre-neural synth | neural output | pre-neural synth | neural output | pre-neural synth | neural output | pre-neural synth | neural output | pre-neural synth | neural output | |
| Source S | ![]() | — | ![]() | — | ![]() | — | ![]() | — | ![]() | — | ![]() | — | ![]() | — | ![]() | — |
| Reference R | ![]() | — | ![]() | — | ![]() | — | ![]() | — | ![]() | — | ![]() | — | ![]() | — | ![]() | — |
| Target T | ![]() | — | ![]() | — | ![]() | — | ![]() | — | ![]() | — | ![]() | — | ![]() | — | ![]() | — |
| Text-only | — | ![]() | — | ![]() | — | ![]() | — | ![]() | — | ![]() | — | ![]() | — | ![]() | — | ![]() |
| Raw source S | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
| Shuffled source | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
| Pointwise concatenative | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
| CRG, without DAM | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
| FO, without DAM | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
| VMO, without DAM | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
| CRG | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
| FO | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
| VMO | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
Here we hold the humpback-whale reference from the introduction unchanged and use guqin, a Chinese instrument, and another humpback whale recorded under completely different conditions as the sources.
| Condition | Guqin source → fixed humpback-whale reference | Different humpback-whale source → fixed humpback-whale reference | ||
|---|---|---|---|---|
| pre-neural synth | neural output | pre-neural synth | neural output | |
| Source S | ![]() | — | ![]() | — |
| Reference R | ![]() | — | ![]() | — |
| Raw reference R | ![]() | ![]() | ![]() | ![]() |
| Raw R + S mixture | ![]() | ![]() | ![]() | ![]() |
| Pointwise concatenative | ![]() | ![]() | ![]() | ![]() |
| CRG (relational) | ![]() | ![]() | ![]() | ![]() |
| FO (relational) | ![]() | ![]() | ![]() | ![]() |
| VMO (relational) | ![]() | ![]() | ![]() | ![]() |