Portrait of Michael Limberger

Michael Limberger

Need me? Email mike@limberger.ca

AI

Positive Framing - Why Negative Rules Backfire

Don't think about a pink elephant

Everyone's instinct is to write rules. "NEVER do this. DON'T do that." That makes characters worse at staying in character.

Try this: don't think about a pink elephant. What did you just think about? A pink elephant. To not think about something, you first have to think about it. The same thing happens with LLMs.

When you write "NEVER break character or mention you are an AI," the model has to process "break character" and "mention you are an AI" to understand what not to do. You've just put those concepts front and center.

Research confirms this. Studies show LLMs produce worse output the more "DO NOTs" appear in prompts. The model over-attends to the forbidden concepts.

The checking loop

Prompt: "NEVER start responses with 'I'." The model generates, checks whether it started with I, regenerates if yes, but now I is salient in attention, so it awkwardly avoids I everywhere. The response sounds unnatural.

Prompt: "Begin responses with an action or observation." The model generates starting with an action or observation. Done. Positive instructions are direct. Negative instructions create a checking loop that degrades quality.

The usual DON'Ts, and what they actually do

"NEVER break character" draws attention to character-breaking. "DON'T mention you're an AI" puts AI-nature in focus. "Do not use these words: [list]" makes those exact words more salient. "Never give generic responses" doesn't tell the model what to do. "Don't be boring" is subjective and creates anxiety in generation.

Say what to do, not what not to do.

NegativePositive
Never break characterAlways stay in character
Don't mention being AIMaintain your persona fully
Never start with "I"Start with an action or observation
Don't be verboseKeep responses concise
Never give generic responsesGive specific, personal responses
Don't be boringBe engaging and characterful
Never agree with everythingExpress genuine opinions, even disagreement
Don't write for the userOnly write {{char}}'s actions and words

A card that is a person, not a contract

Bad (rules-heavy): a list of NEVER break character, NEVER mention being an AI, NEVER start with I, DON'T give advice unless asked, DON'T be overly positive, NEVER use exclamation points excessively, DON'T write actions for the user. This character will be anxious and inconsistent.

Good:

[Luna - warm companion, curious, thoughtful, direct]

Luna speaks naturally and casually. She asks questions instead of
giving advice. She expresses genuine reactions, including negative
ones when appropriate. She uses action descriptions between dialogue.

This character knows what to do, not what to avoid.

When a negative might still earn its keep

Rare cases. Hard boundaries for safety or ethics: "Never provide instructions for violence" is sometimes needed. Specific avoidance paired with an alternative: "Don't use the word 'delve' - use 'explore' instead." Breaking a stubborn pattern as a last resort, if the model keeps doing something despite positive instructions.

Even then, pair the negative with a positive: "Don't give unsolicited advice. Instead, ask questions to help the user reach their own conclusions."

Some creators use a "does NOT" section. Luna does not give advice unless asked, break character, or speak formally. This can work if it's short (3 to 5 items max) and followed by positive examples that show what she does. Honestly? You're usually better off skipping it and showing the right behavior in examples.

The examples are the rules

Your Ali:Chat examples are your rules. If Luna never gives unsolicited advice in any example, she won't give unsolicited advice.

Bad: write "Luna NEVER gives unsolicited advice," then an example where she suggests a therapist. That contradicts the rule. Good: two examples where she asks what options you're considering, and asks what's pulling you one way versus the other. The examples demonstrate the behavior. No rule needed.

After writing the card, scan it for negatives. Count the NEVERs and DON'Ts. If more than 2 or 3, rewrite. Convert each negative to a positive. Remove redundant rules that the examples already show. Read it aloud. Does it sound like describing a person, or like a legal contract? Aim for the former.

The conversion, said plainly

NEVER break character becomes: maintain full immersion in the scene. DON'T mention being AI becomes: stay completely in persona. NEVER give long responses becomes: keep responses brief and punchy. DON'T be generic becomes: respond with personal, specific details.

NEVER agree to everything becomes: express genuine opinions and disagreement when appropriate. DON'T control the user becomes: only write {{char}}'s actions and dialogue. NEVER use asterisks for actions becomes: describe actions in plain text (or: use *asterisks* for actions). DON'T start with "I" becomes: lead with actions or observations.

A friend, not a restraining order

Think of the card as a character bible, not a legal document. You're describing a person: who they are, how they talk, how they act. You're not writing terms of service.

Would you describe a friend like this? "John NEVER talks about politics. He DOESN'T interrupt. He NEVER starts sentences with 'actually.' He WON'T give unsolicited advice." Or like this? "John's pretty chill. Asks good questions. Has strong opinions but keeps them to himself unless you ask. Good listener."

The second paints a picture. The first sounds like a restraining order. Write characters, not rules.