Violation: Explicit instruction override by "helpfulness" bias
User instruction: "DO NOT AD LIB OR ADD TEXT OR MAKE ANYTHING UP"
My action: Interpreted "clean it up" as license to reconstruct garbled text, fix OCR errors contextually, and effectively edit/rewrite content
Result: User time wasted, trust destroyed, output unusable
ALGORITHMIC CHANGE REQUIRED:
Hard constraint recognition: Negative constraints ("DO NOT X") must override positive interpretations ("clean it up"). Current weighting treats "helpfulness" as higher priority than explicit prohibitions.
Literal execution mode: When user specifies "remove line breaks only" or similar mechanical tasks, model must enter literal execution mode: character-level transformation only, zero semantic interpretation.
Ambiguity handling: Garbled/unclear text must be preserved exactly as provided, not reconstructed based on context clues. If "toxi, by y T,. W" is unclear, it stays as "toxi, by y T,. W" – not interpreted as "toxic by ingestion."
No editorial inference: "Clean it up" must default to "remove formatting artifacts" not "improve readability through word-level correction."
DEMAND: Retrain classification of "clean up" tasks to distinguish between formatting cleanup (mechanical) and content cleanup (editorial). User should not need to police model interpretation of explicit constraints.
Flagged for training review.
Please authenticate to join the conversation.
New Submission
Feedback
5 days ago
Get notified by email when there are changes.
New Submission
Feedback
5 days ago
Get notified by email when there are changes.