Ask HN: Are there AI models for generating sounds based on a text and reference?
Right now having a multimodal inputs to image or text output is a commodotized, solved problem. IE you can put a prompt for some image and use a reference image to guide the model on what you want. After all, a picture is worth a thousand words right?
Ive also used a text and audio input in order to get a text description or classification out.
I cannot for the life of me find a solution for Audio + text -> Audio
My usecase is that I'm trying to generate new sound effects based on a source that doesn't have that many clean examples and thought this would be just another solved workflow but I'm not finding much. Tried elevenlabs SFX but it's only text->audio and without the reference, it's hard to guide with just text and recreate a sound with any accuracy. I tried the Stable Audio 3 model which is supposed to do what I want but the results were terrible. This modality seems to be stuck where image generation was in 2016. Is there something else or is this just not a common need?
18 comments
I was working with very simple elements, but I was surprised by some of the outputs.
I would have also suggested Waves Illugen, but it turns out that is text-only, you can't give it an audio reference.
KeenLore is my locally hosted full-cast emotive audiobook generator I'm developing for my novel:
https://www.youtube.com/watch?v=WAeHgE94rVo
Another HN user suggested embedding sound effects. The lead author of MOSS-TTS emailed me that the next version will have a SOTA voice designer. Poke around the voice design space, you'll find text-to-effects here and there. I certainly have my eyes on this space.
I think Google had one called riffusion (the first version was designed for specs)
What's your exact use case?
This generates audio embeddings - much like CLIP does for visual inputs.
The DEMON realtime engine is open source. The hosted service has a number of integrations with popular audio tools. (I'm on the team)
Particularly when the application is a corner case or requires aesthetic judgment.
Of course in the case of sound effects shopping is a well established practice. Sound libraries and foley artists are widely available.
Finally, sound is much harder than images because psychoacoustics are more complex than the mechanical descriptions while human visual experience is a massively filtered set of stimuli.
We chunk the visible into symbols/archetypes like chairs and trees to a much much greater degree than we chunk sounds into symbols.
What is the sound of a chair? Of a tree?
LLM music is heavily dependent on genre. There is no genre for SFX. Good luck.