I've been saying that AI waifus are at most like five years out. Multi-modal AI is already here, you just need a little more model size to get something that maps video and audio to video and audio. GPT-4o already does text+images+audio->text+audio, it's just a question of scale.
Beyond that, a sufficiently autistic group of /g/ people could probably get most of the way there just by stitching existing tech together in clever ways. If the internet had the comradery it did pre-2018, we'd already have open-source AI vtuber gfs.