Long text generations timeout after 15 minutes of executions. There are use cases that demand many tens of thousand output tokens. SOTA models support up to 128k output tokens,but it is impossible to get benefit from that in 15 minutes. Please, raise that at least to 60 minutes, and consider adding an async endpoint for a limitless option. Additionally, this async endpoint could embrace the same batch approach both Anthropic and OpenAI support. No quick delivery guaranteed, but half the price
Please authenticate to join the conversation.
New Submission
Feature Requests
API
3 days ago
Get notified by email when there are changes.
New Submission
Feature Requests
API
3 days ago
Get notified by email when there are changes.