Skip to main content

Voice to Text / Transcription

Voice Input/Voice Output
Have a transcription endpoint with the API like OpenAI Whisper.

Status: Completed38 comments

Log in to comment and vote

Comments38

  • Amber Doodle

    •

    Feb 5, 2025

    need this for sure

  • Sapphire Spaceship

    •

    Feb 7, 2025

    Must have!

  • Aquamarine Nucleon

    •

    Feb 5, 2025

    https://github.com/openai/whisper

  • Orange Pond

    •

    Mar 12, 2025

    This is a MUST. The ability to naturally talk to the ai makes having conversations fluid. Makes the process of interaction a real Game Changer!

  • Harlequin Field

    •

    May 19, 2025

  • Green Waffle

    •

    Jun 15, 2025

    I recommend whisperx - 

    It uses whisper and a vad model to reduce hallucinations. I’ve been using it in production and makes a big difference in transcript quality vs vanilla whisper or even the new OpenAI voice stuff

    https://github.com/m-bain/whisperX

  • Rose Spoon

    •

    Feb 7, 2025

    Business users need this

  • Emerald Sprout

    •

    Feb 17, 2025

    This and embeddings and I can break ties with OpenAI.

  • Tan Pine

    •

    Feb 26, 2025

    This would be a life saver

  • Red Turtle

    •

    Feb 26, 2025

    modle lol i suck

    • Jack

      Team•

      Mar 28, 2025

      Voice input is on our list of features we are looking at. On your image analysis issues have you tried Mistral Small 3.1?

      • Red Turtle

        •

        Mar 28, 2025

        yes i tried it and its good fast and reliable. thank you for listening to our requests.

        • Jack

          Team•

          Mar 28, 2025

          Brilliant. Merging this to keep track of upvotes.