Filtering of multimodal models

Hello Venice team,

There really needs to be a way to filter by model capability and multimodality when selecting a model (regardless of model type).

The current filtering system is useless, sorry, but all i can filter by is free/not free and encrypted/not encrypted? Really? When you can use 10s of models each with it’s own capabilities for input, output and context length?

Let me filter by model accepted input (image, video, audio). Let me filter by model size and context length. Let me filter by model age (the date the model has been released).

This is CRUCIAL to be able to navigate and use models efficiently.


Another thing that has been bothering me for a while is that models that completely support multimodal inputs on their own site or on other platforms seem to be text only here. Can’t we pass to the model what the user uploads and let the model actually use that instead of being converted to text through your own tools????

What is the point of even having complex multimodal models available here if all they get is the text and a very bad text translation of the user’s uploaded documents?

I am interesting in using AIs for image/ video/ sound analysis. If your platform only sends text to the AI models, even when using filly multimodal PAYED models, then what is the point?

Please authenticate to join the conversation.

Upvoters
Status

New Submission

Board
💡

Feature Requests

Tags

New Model

Date

3 days ago

Subscribe to post

Get notified by email when there are changes.