Async chat completions & batch mode

Long text generations timeout after 15 minutes of executions. There are use cases that demand many tens of thousand output tokens. SOTA models support up to 128k output tokens,but it is impossible to get benefit from that in 15 minutes. Please, raise that at least to 60 minutes, and consider adding an async endpoint for a limitless option. Additionally, this async endpoint could embrace the same batch approach both Anthropic and OpenAI support. No quick delivery guaranteed, but half the price

Please authenticate to join the conversation.

Upvoters
Status

New Submission

Board
πŸ’‘

Feature Requests

Tags

API

Date

3 days ago

Subscribe to post

Get notified by email when there are changes.