Keep Llama-4
I have been playing around with AI characters, kind of fascinated by the difference among the different models. Having been playing with the Venice Medium & Venice large models over the past week, they are absolutely awful for characters, unable to consistently follow the most simple & direct character instructions, while latching onto (generally incorrect) inferences it has made about the character.
The Llama 3.1 & 4 models have their quirks, but they're both much, much better. However, I notice that you are discontinuing them in another week. Is there any chance you could be persuaded to keep one of them (preferably the Llama 4 model, which seems less likely to lapse into stream of consciousness while allowing greater creativity than Llama 3.1)?
Log in to comment and vote
Comments2
Jack
Jun 24, 2025
We've removed Llama-4 due to underwhelming performance and output. The large models still available are:
– 3.1 405B (UI and API)
– DeepSeek R1 (API only)
– Qwen 3.2-35B (UI and API)
Teal Vacuum
Jun 12, 2025
I feel the same way. I’m taking an old Xeon workstation that I wasn’t using and adding a 12GB RTX card in it to see if I can run Ollama, Open WebUI, and Llama 3.2 3B. It looks like Venice was using the smaller Llama models, so it should work reasonably well on the 12GB card. The workstation has two PCIe 3.0 16x slots, so I was able to fit the 12GB card in one and an adapter for an M.2 NVMe drive in the other. Windows 11 Pro now works on the old workstation, so I was able to go into advanced settings and tell it to install apps on the NVME drive instead of the SATA drive where Windows lives. In the next few days, I should know if it’s possible to run Llama effectively for my $400 in graphics card and adapter.