@seeingwithsound To me the evidence suggests that visual memory is a top-down process that activates neurons in V1 much broadly, while visual percepts are mediated (as we know) from bottom-up inputs to V1. Somehow the brain distinguishes these two modes, which I find quite interesting. I'd think that visual imagery will be similar to visual memory in quality because it will be a top-down influence on V1. I know you're trying to get audio info to generate visual images. If this involves A1 inputs (direct or indirect) to V1, I'd imagine it would be through the similar top-down circuit. Meaning, V1 needs to "learn" how to represent these top-down information and map to visual space.
@hkl 2. Visual recall is typically like retrieving a (low-res) bitmap, whereas voluntary mental imagery is more like re-synthesizing from mentally extracted high-level information, like instructions, e.g. "a square below a rising line" (after having understood a soundscape), much like VR bitmaps on a computer are generated on-the-fly instead of retrieved as images.
Clearly I lack understanding of how this all works, so this is more like a coarse checklist for applicability of theories.
@seeingwithsound You have a good point. Immediate recall will be different than recall of visual memory. I'd think the latter will be similar to visual imagery, since it likely involves prior knowledge (i.e. memory) to generate visual imagery.
It seems thinking in terms of providing information that these top-down circuits normally deliver to V1 would be a productive way to allow learning of auditory information in the visual circuitry.
@hkl Thank you! Yes, I think assuming a similar top-down neural substrate for voluntary visual imagery as for visual memory is a good starting point. However, beyond that there may still be key differences that really matter:
1. Most people can kind of "replay" a spoken phrase that they just heard for up to say 10 seconds in much greater detail than afterwards, so this could play a role in visualizing synthetic soundscapes, comparing what was just heard with what is being visualized.