Suggest improvements, new features, or enhanced capabilities for the Venice.ai API. Before submitting, please review our API Docs to understand current capabilities.
Join the conversation in our #api channel on Discord.
ad intellectum infinitum
add embedding model(s)
Apologies if asked/answered, didn’t see it in search results. Having embedding models on API would be great if that’s feasible! thanks
'Responses API' support
Integration with OpenAI Codex CLI doesn’t work with Venice API because the Venice doesn’t support the responses endpoint. https://community.openai.com/t/introducing-the-responses-api/1140929
Add tokens consumed to /image endpoints response
Please consider adding the amount of tokens consumed by the /image endpoints, like you already do for chat completions. This is needed to prevent third parties having to manually update the cost per request, as would be the case for services where end users consume tokens from a pool.
Long-term overloaded model = trust issue
Models that have been overloaded for more than two weeks, or even longer, are a real problem for APIs, as we are forced to have an automated fallback strategy on Openrouter for security reasons. This is a trust issue for venice.ai, as it is not a simple problem of one or two days, but one that persists without being mentioned or fixed.
Access to Diem/USD balance on inference only API's
We're building Dandolo.ai, a decentralized AI inference marketplace that routes requests across multiple Venice.ai providers, and we've hit a critical limitation: inference-only API keys cannot access balance data from the /api/v1/api_keys/rate_limits endpoint, which requires admin keys. Our platform allows providers to register their Venice.ai keys with us to serve AI requests, and we need to programmatically check their VCU/Diem balances to prevent failed requests, calculate fair rewards based on usage, and determine provider availability - but we cannot ask providers to trust us with admin keys that could access their VVV tokens or billing information. Without balance visibility, we cannot scale beyond manual balance entry, fairly distribute rewards, or efficiently route requests based on available credits. We request that Venice.ai either: (1) create a balance-only endpoint accessible to inference keys, (2) include balance information in inference response headers (e.g., X-Venice-Balance-VCU), (3) provide webhook notifications for balance updates, or (4) introduce a new scoped API key type with read-balance and inference permissions but no access to sensitive billing or token management. This enhancement would enable an entire ecosystem of third-party applications to build on Venice.ai while maintaining security best practices, as developers could track usage and balances programmatically without requiring dangerous admin-level access.
Dynamic API rate limits based on VVV stake/VCU allocation
We are hitting daily limits after approx 250 VCU. Is there a plan to introduce dynamic rate limits based on VVV stake or available daily VCU ?
Add function calling on the E2EE GLM 5 e2ee-glm-5
Can Venice enable function/tool calling for e2ee-glm-5? The non-E2EE counterpart zai-org-glm-5 already supports it, so this would bring the E2EE model to feature parity. Currently the only E2EE models with function calling are the Qwen models.
Do not remove deepseek-r1-671b from the API
The reason given is that qwen3-235b is a better model, but I disagree very much. It does not hold a candle to Deepseek. I have used Deepseek for all my work and is essential. qwen3 is not a sufficiently good replacement. I’d be even willing to pay a premium to keep access. Please do not retire DeepSeek. Alternatively, introduce DeepSeek-V3.1.
Edit VCU Limit on API-Keys
Right now it is not possible to edit the VCU Limit once defined. If you buy more VCUs or use more consuming endpoints it would be cool to flexibly change the limit. Right now if you need more you need to change the key.
Display Max Output Tokens
When using Venice AI in VS Code based IDEs, it often asks for the model’s maximum token output count. It would be nice if this was listed on the model page.
Aggregated API consumption limit
Would be great to be able to set a DIEM or USD limit on an API key that does not reset every 24 hours. This would be an aggregated consumption limit regardless of the the age of the key. e.g. Set a limit of 1 DIEM. It does not matter how long it takes to reach the limit but once reached the key is blocked.
API response signing
Sign API responses from Venice to prove that text was generated by Venice and is not being spoofed, enabling one user (API key) to do verifiable work on behalf of another user
API: disable the venice system prompt by default
The venice.ai API injects the venice system prompt by default. That's a bit unexpected and wastes tokens. The API should be raw, it is usually used by agents and other systems that bring their own system prompts. Please disable the venice system prompt by default in the API or at least add a global or per-key option to disable it because it is very inconvenient to add a custom request parameter to each API consumer. For reference: https://docs.venice.ai/api-reference/endpoint/chat/completions#body-venice-parameters
I love love love using ai to create my art work now days I have found my credits get ate up crazy fast I was hoping on finding smarter ways I can go about getting credits for generating images or short vids I don’t know much about VVV or Diem and to be honest my comprehension skills have never been the best
I’d love for any feedback on ways I can create NSFW images easier or with out spending the crazy amounts that I have
Increase /image/multi-edit input limit for GPT Image 2 and Nano Banana 2
Venice currently caps /image/multi-edit at 3 image inputs for gpt-image-2-edit and nano-banana- 2-edit. Please raise the image-input limit to match upstream model capabilities where available. OpenAI’s direct GPT Image API supports up to 16 image references for GPT image edit requests. Google’s Gemini API for Nano Banana 2 supports multi-image prompting through Gemini’s content parts, with much higher general image-input limits, and provider wrappers expose Nano Banana 2 editing with up to 14 references. A 3-image cap blocks practical workflows like character sheet + costume + prop + environment + style reference, product bundles, storyboard panels, and multi-reference consistency testing. A target of at least 16 inputs for GPT Image 2 and 14+ for Nano Banana 2 would better expose the underlying model capability.
Feature request: Programmatic account creation and credit topup API
I'm building https://getbased.health, a lab work dashboard that uses Venice as one of its AI providers. Currently users have to leave the app, create a Venice account manually, get an API key, and come back to paste it in. This creates a lot of friction. Another provider we integrate with, https://ppq.ai, offers API endpoints that enable a seamless in-app flow: POST /accounts/create — instant account creation, returns an API key POST /credits/balance — check remaining balance POST /topup/create/{method} — programmatic credit topup (Lightning, Bitcoin, Monero, Litecoin and Liquid) POST /api/v1/accounts/create?ref={referralId} GET /topup/status/{id} — poll payment status GET /v1/models — dynamic model list with capabilities This lets us offer one-click onboarding: user clicks "Create account", gets an API key instantly, tops up with crypto, and starts using AI — all without leaving the app. Venice already has the model list and balance (via x-venice-balance-diem header). Even a single POST /accounts/create?ref={id} that returns an API key would unlock frictionless in-app onboarding for third-party apps. Would love to see this on the roadmap.
Filter for closed-source models from the API
the GET/models endpoint really needs to have a param where users can have the ability to specify {“privacy: private”} so users have the ability to filter out all of the unethical options if they choose. I don’t want to see the Claudes and the ChatGypts of the world in whatever frontend I use and would like the option to clean up my autogenerated model list without having to manually come up with a hack that can let me filter it after the transmission or even worse manually hiding models that keep showing up. A simple expansion to the API would save you a tiny bit of bandwidth and those like me who refuse to give resources/data/network effects to those who actively scalp the commons without so much as a contribution, a cleaner list of models to choose from. We use the Venice API for Venice and do not wish to give our business elsewhere.
help-me
Hi, how are you? I need some help. I tried to use the venice-uncensored model in the Cursor IDE, but I wasn’t successful. It returns a 400 error. The GLM 5 model works fine, and I can use it without any issues. I’m looking for an uncensored model, and from what I understand, venice-uncensored seems to be the only option available. Could you please guide me on how to configure it properly or what might be causing this error?