SOTA models support up to 128k output tokens. It is impossible to leverage that through the Venice chat completions API endpoint, which times out after 15 minutes. Please create a new async chat completions endpoint, that has no time limit, and ideally, follows the same paradigm used by the batch endpoint of Anthropic and OpenAI: half the price, but may take up to 24 hours to complete.
Please authenticate to join the conversation.
New Submission
Feature Requests
API
3 days ago
Get notified by email when there are changes.
New Submission
Feature Requests
API
3 days ago
Get notified by email when there are changes.