Various Telegram bots posted to HN do this. The main issue is that you generally have to wait for the OpenAI request to finish before synthesizing audio while the text can stream in immediately. But it's not bad.
Yeah, quite a few out there. As long as you can write an OpenAI API integration and integrate with browser apis for TTS & transcription, you're set. Probably 20-30 hours total for an implementation.
I was reviewing my old projects recently and found one, a voice assistant from 9-10 years ago, where I used voice recognition and TTS plus a bit of NLP to communicate with Wolfram Alpha as a sort of more nerdy Siri. It was pretty trivial, just a bunch of open APIs put together.
The funny part about it (at least to me) was that I called it a *HTML5* voice assistant, because that was the cool and trendy tech then, no mention of AI/ML anywhere.
Aka you speak to it your question and it speaks back its answer while writing to the screen. Or maybe I’ve missed something like that?