Moonshot models fail over and over

I’ve spent significant time using current agentic AI models, particularly Kimi 2.5 as well as K3 chat. The results consistently fall short for professional, legal or detail-driven work. These models frequently ignore explicit instructions set early in a thread, acknowledge those directives in their responses and then immediately revert to making assumptions or generating unverified content. This isn’t a minor glitch—it actively creates more work.

Every output requires expert-level proofreading, correction, and cross-referencing just to reach baseline reliability.

For casual conversation, this behavior might be tolerable. But when handling legal cases, compliance tracking, or any work requiring strict adherence to facts and established parameters, these models are not ready for prime time. Making assumptions or altering documented details isn’t just inconvenient; it can actively undermine a case or create compliance violations. It could easily ruin a court case.

I acknowledge that alternatives like Anthropic deliver higher accuracy, but they lack the privacy protections required for sensitive advocacy and legal work. The current landscape forces a compromise between utility and security.

Bottom line: The industry needs private, agentic chat options that actually follow instructions without requiring the user to act as an expert proofreader. Until models can maintain context, respect established boundaries, and avoid generating unsupported assumptions, they cannot be trusted for professional or legal workflows. We need better options that deliver privacy, precision, and consistency without the constant cleanup. I suggest Qwen or possibly DeepSeek but DeepSeek has other issues.

Please authenticate to join the conversation.

Upvoters
Status

New Submission

Board
💡

Feature Requests

Tags

New Model

Date

1 day ago

Subscribe to request

Get notified by email when there are changes.