Voice to Text / Transcription💡Feature RequestsNew ModelAPIVoiceFeb 5, 2025Voice Input/Voice OutputHave a transcription endpoint with the API like OpenAI Whisper.Status: Completed38 comments436
Log in to comment and vote
Comments38
Amber Doodle
Feb 5, 2025
need this for sure
Sapphire Spaceship
Feb 7, 2025
Must have!
Aquamarine Nucleon
Feb 5, 2025
https://github.com/openai/whisper
Orange Pond
Mar 12, 2025
This is a MUST. The ability to naturally talk to the ai makes having conversations fluid. Makes the process of interaction a real Game Changer!
Harlequin Field
May 19, 2025
This might be cool:
https://huggingface.co/nvidia/parakeet-tdt-0.6b-v2
Green Waffle
Jun 15, 2025
I recommend whisperx -
It uses whisper and a vad model to reduce hallucinations. I’ve been using it in production and makes a big difference in transcript quality vs vanilla whisper or even the new OpenAI voice stuff
https://github.com/m-bain/whisperX
Rose Spoon
Feb 7, 2025
Business users need this
Emerald Sprout
Feb 17, 2025
This and embeddings and I can break ties with OpenAI.
Tan Pine
Feb 26, 2025
This would be a life saver
Red Turtle
Feb 26, 2025
modle lol i suck
Jack
Mar 28, 2025
Voice input is on our list of features we are looking at. On your image analysis issues have you tried Mistral Small 3.1?
Red Turtle
Mar 28, 2025
yes i tried it and its good fast and reliable. thank you for listening to our requests.
Jack
Mar 28, 2025
Brilliant. Merging this to keep track of upvotes.