Async chat completions

SOTA models support up to 128k output tokens. It is impossible to leverage that through the Venice chat completions API endpoint, which times out after 15 minutes. Please create a new async chat completions endpoint, that has no time limit, and ideally, follows the same paradigm used by the batch endpoint of Anthropic and OpenAI: half the price, but may take up to 24 hours to complete.

Please authenticate to join the conversation.

Upvoters
Status

New Submission

Board
πŸ’‘

Feature Requests

Tags

API

Date

3 days ago

Subscribe to post

Get notified by email when there are changes.