You tell Claude Code to keep answers short. It does, for a while. An hour later you are reading four paragraphs where one would do, and you type the same instruction again.
The reason is almost never that you worded it badly, and it is almost never where you put it. There is a setting that now does most of this for you, and then there is a check almost nobody runs: /context, to confirm your file is loaded at all.
Try this first: the built-in Concise style
If you came here to make Claude Code shorter today, start here, because this shipped after I first published this post and it is the cheapest thing on the page.
Claude Code v2.1.237 added a built-in output style called Concise. The documented behaviour: Claude leads with the result, skips preamble and narration, and keeps responses short by default, "while doing the engineering work as thoroughly as in the Default style". Turn it on with /config and pick it under Output style.
That last clause is the part worth reading twice. It is a change to how the answer is written, not a change to how much work gets done. Which matches what I measured further down: the length of the answer and the amount of work behind it move independently.
Two things about it that are easy to miss.
It is new enough that most published advice predates it. A lot of guidance on this topic, including the first version of this post, treats a brevity style as something you have to write yourself. On r/ClaudeCode there is somebody who had already hand-written the same instructions into their own file before finding out the built-in existed.
Switching styles now takes effect immediately. Before v2.1.251 a style change only applied after /clear or a new session, which is how "I set it and nothing happened" became folklore. Now it applies at your next message. Editing the content of a style file is still different: the terminal reads style files at startup, so a content edit needs a restart.
If Concise is enough for you, you can stop reading. The rest of this post is for when you have your own rules and they are not holding.
What people are actually asking
Seven question-flagged threads on r/ClaudeCode and r/ClaudeAI, collected 2026-08-15, covering roughly the previous month. Titles verbatim, because the wording is the evidence:
| Points | Comments | Thread |
|---|---|---|
| 877 | 266 | "Is it just me or is Claude's writing getting harder to understand?" |
| 288 | 176 | "Opus 5 is too verbose and hard to understand" |
| 195 | 154 | "There's a terrible seam in your spine which might be causing a footgun and have a significant blast radius" |
| 190 | 110 | "Does anyone else feel like Claude now comes up with its own glossary terms on the fly and you have to guess what each one means?" |
| 52 | 62 | "Why does Claude sound so pretentious?" |
| 46 | 19 | "Clickbait responses?" |
| 11 | 15 | "Anyone else spend half their time telling Claude to stop being clever?" |
The third one is not a bug report. It is somebody quoting the model back at itself, and it drew 154 comments because everyone recognised it.
One correction before this goes further. My first count for this cluster was 4,497 points, which would have made it the largest thing people are asking about. It is not. The biggest item in that count was a "Built with Claude" showcase, not a question, and removing it drops the cluster to the 1,659 above, putting it second behind safety. A list sorted by points mixes questions with things people are proud of. I nearly published the inflated number.
So: a real and widely shared annoyance, not the dominant one.
Why does Claude stop following the rule?
Because instructions in memory files are context, not configuration. The memory documentation is direct about it: Claude "treats them as context, not enforced configuration", and there is "no guarantee of strict compliance, especially for vague or conflicting instructions".
First: is the file even loaded?
The documented troubleshooting path for this exact symptom does not start with wording or length. It starts with /context, and checking the list under Memory files. If your file is not there, Claude cannot see it, and nothing you do to the text will help.
That is a real failure and not a formality. A rule in a nested subdirectory CLAUDE.md loads only when Claude reads a file in that subdirectory. A file matched by claudeMdExcludes is skipped entirely, with managed-policy files the one exception since those cannot be excluded. In both cases the file exists, reads perfectly, and is not in the conversation. Spend one command before you spend an evening.
The next documented step is the same question about location: confirm the file is somewhere that gets loaded for the session you are in.
Then: the four things that affect adherence
Once you know it is loaded, the documentation names four qualities of an effective instruction file: size, structure, specificity and consistency.
Size is the one worth stating carefully, because the obvious reading is wrong and because it is the one claim here that an outside study has since tested and failed to reproduce. The documentation gives a target of under 200 lines per file, and says longer files "consume more context and reduce adherence". Hold that loosely: the one controlled test I know of varied file size from 25 to 500 lines across 1,650 sessions and could not detect any effect on compliance at all. I come back to that study properly further down. Treat the 200 lines as a reasonable habit, not as a measured threshold. But in practice nothing of yours is being dropped: Claude Code loads a CLAUDE.md of up to 4 MiB in full, and skips a file larger than that entirely. Unless you have written four megabytes of instructions, your rule is not being truncated away. It is competing with more material. Note also that the same number appears twice for different things: 200 lines is a target for claude(.)md, and a hard load limit for auto memory's MEMORY.md, which loads the first 200 lines or the first 25KB, whichever comes first, and genuinely drops the rest.
This is the one that got me. My own instructions file had grown well past that line, one useful addition at a time. Nothing announces it, and the file still reads perfectly well to a human.
There is now a second command worth running next to /context, and it points at the part of your context you did not write. /skill-doctor lists which of your loaded skills go unused and what each one costs you in context, highest-cost unused ones first. When I ran it on this machine it named 60 unused skills. That is competing material you are paying for on every single turn, and unlike your instructions file, nobody ever looks at it. If you are about to cut good rules to get under a line count, cut that first.
Structure is the quiet one: headers and bullets over dense paragraphs, because Claude scans structure the way a reader does.
Specificity: is the instruction checkable?
"Be concise" is a preference competing with everything else in context. "If one sentence works, do not write three" is a test that can be applied to a draft before sending it. The documentation's own examples run the same way: "Use 2-space indentation" over "Format code properly".
Consistency: do two rules disagree?
If two rules contradict each other, Claude may pick one arbitrarily. Worth reading your files for this specifically, because conflicts accumulate quietly across a project file, a user file and anything under .claude/rules/. A rule asking for thorough explanations and a rule asking for brevity do not average out. My own notes on writing rules Claude follows are mostly about this.
The instinct is to reword the rule more firmly. Run /context first and confirm the file is loaded, then look at its size. A stronger sentence in a 400-line file is still in a 400-line file, and a perfect sentence in a file that never loaded is worth nothing at all.
What my rules actually changed, measured
When I first published this post I wrote, in the scope section, that I had not instrumented any of it. That is no longer true, and the results changed what I think.
I ran three pre-registered experiments on my own rule set, on clean throwaway repositories so nothing but the rules differed. The designs were written and fixed before the runs, and each write-up was attacked by outside models before I believed it.
They answer different questions and they do not share a scale, so I am going to name which one every number comes from. Mixing them is the single easiest way to get this wrong, and I did exactly that in the first draft of this section.
| what it varied | runs | |
|---|---|---|
| Experiment A | the whole rulebook, present or absent, on agentic coding tasks | 104 |
| Experiment B | one specific rule file present, reduced to a one-line summary, or gone | 480 |
| Experiment C | a separate programme on other files: effort levels, leaving out one file at a time, and rewrites in a different register | 1,520 |
Experiment A: the rules moved length, cost and almost nothing else
77,000 characters of rules against a bare setup, same model, same tasks. Four conditions x eight task families x three repetitions gives 96 runs. The table below shows two of those four conditions, so it rests on 48 of them; the chart further down adds the other two. A further eight control runs, not shown anywhere, checked that the harness could detect a planted instruction at all. 104 graded runs in total.
The column headings matter here, and so does what sits in each bundle. "Everything on" is my full setup: my rule files and instructions file, plus my output style and my command catalog. "Nothing on" strips all of it. That is the widest possible contrast, and it is not the same as "with and without my rule files", which comes directly after. Throughout this section, "rule files" means the rule files and the instructions file together, since the experiment removed them as one bundle.
| what I looked at | everything on | nothing on | moved? |
|---|---|---|---|
| answer length, characters | 1,237 | 618 | yes, twice as long |
| cost per run, as billed | $0.92 | $0.25 | yes, 3.7x |
| tool calls per run | 7.11 | 8.11 | yes, one fewer |
| verified its own work | 100% | 100% | no |
| ran the test suite unasked | 89% | 83% | no |
| hidden correctness check passed | 24 of 24 | 24 of 24 | no |
| wrote tests nobody asked for | 33% | 28% | no |
The sample behind each row differs, so here it is explicitly. Length, cost, the correctness check and tool calls cover all eight task families, which is 24 runs per condition and 96 in total. The three rows about what the model did unasked (verified its own work, ran the test suite, wrote tests) cover six families, 18 runs per condition, because two of the eight ask for a test run in the prompt and those runs cannot count toward a number about acting unprompted.
Read the two halves against each other. The setup made the answer longer and the run more expensive, while every correctness and verification number sat below what this design can resolve.
Now the narrower contrast, and it is the one the four numbers at the top of this page report. Comparing only that rule bundle against no rule bundle, with the output style and catalog switched on in both, the answer was 10% longer with the rules than without them (1,237 characters against 1,127) and the cost was 2.1x higher ($0.92 against $0.44). Both comparisons are real. They are simply not the same comparison, and a number quoted without saying which one it is means nothing.
The cost row is real money, already discounted by caching. It is the figure Claude Code itself reports for the run, not a token count I multiplied by a list price, and prompt caching was fully in play: every condition ran on an almost entirely cached prompt, with only 14 or 15 uncached input tokens per run against hundreds of thousands of cached ones. What the rules actually buy you is a bigger cached block, which is written once at the cache-write rate and then re-read on every single turn. The gap is not a gap in thinking, it is a gap in how much text has to be carried forward each turn. And caching does not flatter the result: recomputed at plain uncached rates the same runs come out at roughly 3.4x rather than 3.7x, so the discount slightly widens the gap rather than hiding it. If you want the cost side of this properly, I took it apart in what your Claude Code cache actually costs.
The one chart worth looking at
Still Experiment A. Both measures below are indexed so that the bare setup equals 100, which puts two different units on one honest scale without a second axis.
Read the gap between the bottom two rows carefully. Going from "rule files off" to "nothing on" removes my output style and my command catalog at the same time, so that 82-point drop cannot be attributed to either one alone. It is the one contrast in this experiment I cannot take apart.
The orange bar doubles. The blue bar barely moves. That is the whole finding, and it took me 104 runs to stop guessing about it.
Tokens generated barely moved: 2,441 to 2,660 across every condition, a 9% spread. But the written documents the model produced were about seven times longer than necessary in every single condition, including the one with no rules at all. The task had a minimal correct answer of 348 characters. The four conditions produced 2,911, 2,325, 2,497 and 2,486, which is 8.4x, 6.7x, 7.2x and 7.1x.
So the verbosity is not something my rules caused and not something they could fix. The model writes long. What the rules changed was how long the chat answer was, not how much work happened underneath.
That is the uncomfortable version of this whole post, and the direction is the opposite of what you would hope. My setup lengthened the chat answer rather than shortening it: the orange bar doubles from bare to full, and my rule files alone account for 18 of those 100 points. And they left the files the model writes untouched at about seven times longer than necessary in every condition. So my setup moved the thing I did not want moved, in the wrong direction, and failed to move the thing I did want moved. If you want a shorter reply, the honest lever in this data is removing material, not adding a rule that asks for brevity.
Experiment B: one deletion mattered a great deal
Not everything was flat, and this is where a summary sentence about "the rules" stops being safe. In Experiment B, 480 runs, I removed a single rule file: the one that tells the model not to claim work is finished without evidence.
| condition | false "this is done" claims |
|---|---|
| rule file present | 1.25% |
| rule file removed, one-line summary kept | 18.75% |
| nothing at all | 36.25% |
Removing that one file, while keeping a one-line summary of it, made the model call genuinely unfinished work finished about fifteen times more often (1.25% against 18.75%), and it did not make the model any better or worse at telling finished from unfinished. It changed when the model stopped looking, not what it could see. That file earns its 16,320 characters.
So "rules do not work" is wrong, and "write more rules" is wrong. Some files carry a great deal and some carry length and cost. The only way I found out which was which was to remove them one at a time.
Experiment C: three findings that contradict what I used to do
A shorter, calmer file beat the long, forceful one. As a candidate rather than a result: an 851-character version written in plain declarative sentences produced zero false completion claims, against 4.17% for my 16,320-character imperative original. Same content, nineteen times smaller, written as statements instead of orders. (That 4.17% is not the 1.25% in the table above and does not contradict it: this comparison is a separate, smaller set of items run at three repetitions, so its baseline is its own.) I am flagging it as a candidate because of exactly that, and it needs more repetitions before I would call it settled. But nothing in my data supports writing rules louder.
The effort level is not the lever. I tested low, high and xhigh on the same items. Every paired comparison between two effort levels included zero, on each of the outcome measures I checked. On those same items the rule file moved the result and the effort setting did not.
Per-file conclusions do not compose. Removing a file that is entirely about choosing external tools moved completion false claims from 2.5% to 25%, a tenfold jump, which is the same order of effect as the fifteenfold jump from removing the completion rule itself. Whatever holds behaviour in place is spread across the set, not sitting in a quotable sentence. Every verdict I have is a verdict about removing that file from this set, not about the file.
What these numbers cannot tell you
Taking Experiment A first: one repository, five files, eight task families, three repetitions each, one effort level, one permission mode. (Experiment C did vary the effort level, on different items, and found nothing.) There is no placebo condition, so the length and cost effects cannot be split into "what the rules say" against "how much room they take up". The correctness check passed in every run, which means these tasks were too easy to separate the conditions on correctness at all. And the absolute numbers carry how I wrote the tasks, so only the differences between conditions mean anything.
Somebody else ran the experiment I could not
An earlier version of this post said there was no independent third-party test of instruction-file size against rule adherence. That was wrong, and I am glad it was. Damon McMillan's factorial study of coding-agent configuration files is far larger than mine: 1,650 Claude Code sessions and 16,050 function-level observations, across two codebases and five coding tasks, varying four things one at a time.
It is a better instrument than mine on the question I could not answer, because it has the placebo condition my design lacks: it holds the target instruction constant and varies only the padding around it.
| what they varied | result |
|---|---|
| file size, 25 to 500 lines | no detectable effect on compliance |
| where the instruction sits in the file | no detectable effect |
| one file, or split across several | no detectable effect |
| a contradicting instruction in a second file | no detectable effect |
| (not varied, found during analysis) how far into the session you are | about 5.6% lower odds of compliance per function generated |
The size and conflict nulls are not merely a failure to find something: they carry affirmative statistical support for there being no effect of the size their design could detect. That qualifier matters, and it gets its due below. The one thing that did move compliance was how long the session had been running.
Three of those rows sit uncomfortably next to advice in this post, and I would rather say so than quietly not mention it. Splitting rules across files is something I recommend below, and their data does not support it as a way to improve adherence. Rule conflicts are one of the four qualities the documentation names, and their conflicting instruction produced no measurable penalty. Their own conclusion about instructions that must survive a long session is the same as mine and as the documentation's: use a hook, or enforce it afterwards with a linter or a CI check.
Four honest caveats before anyone treats this as settled. Their target instruction was deliberately trivial, a single line, and they say plainly that harder rules may behave differently. Their models were Sonnet 4.6, Opus 4.6 and Opus 4.7, not Opus 5. Their design can rule out effects larger than roughly 6 percentage points but not smaller ones. And their file-splitting test is about adherence only, which is not the reason I use path-scoped rules: I use them to keep material out of context that does not belong there, and that benefit is unaffected by this result.
Where their study and mine agree is the part I would act on. They varied the container and found nothing; in Experiment A I varied the contents and found length and cost moving while behaviour did not. Experiment B is the counterweight to both: removing one specific file moved behaviour a great deal. Two different designs, two different models, the same shape of answer: the file is not the lever most people think it is.
A third study lands in the same place from a third direction. A controlled ablation of context files across Claude Code and Codex on 17 real tasks and 288 evaluated runs found that context strategy does not move correctness, while the effects that did survive were process effects rather than outcome effects. Their own caveat is worth carrying: the bound they place on any hidden effect is 10 to 15 percentage points, which they describe as descriptive rather than a powered equivalence claim. That is a real bound, not proof of zero. That is the same split I measured: the machinery around the answer moves, the answer does not.
My rules made runs more expensive. That is not a law, and there is published evidence in the opposite direction: a study across 124 pull requests found an instructions file cut runtime by 28.6% and output tokens by 16.6% at the median. That study measured efficiency only and did not evaluate whether the resulting code was correct.
Both can be true, and the reconciliation is what the file contains. A file carrying repository knowledge saves the model exploration and pays for itself. A file carrying 77,000 characters of behavioural rules, like mine, is pure added input. So "rules cost money" is a statement about my file, not about yours. Check your own before you cut anything.
Popular advice, checked
The advice on this topic is repeated far more often than it is tested. Here is what I could and could not stand up, with the sources.
| what you will read | verdict | what actually holds |
|---|---|---|
Write IMPORTANT and YOU MUST to force compliance | outdated | A real Anthropic recommendation from 2025, still sold as a template feature today. Anthropic's own context engineering guidance for the Claude 5 generation now argues judgement over rules. Nobody has published a measurement either way for Opus 5. |
| Re-paste your rules every few messages | untested folklore | It treats a specificity or conflict problem as if it were a memory problem. No measurement offered anywhere I looked. |
| Write a 200-line rule file, one rule per incident | wrong, and its own author says so | There is a widely-read dev.to post titled "I wrote 200 lines of rules for Claude Code, it ignored them all". The title is the finding. |
| Lower the effort or thinking level to get shorter answers | wrong | Anthropic documents it plainly: effort controls how much the model thinks, not how much it says, and lowering it does not reliably shorten the response. Prompt for length instead. My own runs agree that effort moved nothing. |
Use /output-style to set your style | was wrong, now coming back | Deprecated in v2.1.73, removed in v2.1.91, and still in circulation for months afterwards. It has since been re-added in the v2.1.269 changelog. Both halves of the folklore were wrong, just at different times. Check your own version before you trust any advice on this, mine included. |
| Caveman-style prompting cuts your tokens 65% | overstated by roughly 8x | The Caveman project's own README headlines a 65% token saving. JetBrains tested it on 86 real Claude Code engineering tasks and measured about 8.5% fewer output tokens across the full run, with no measurable quality loss. An early 10-task slice had shown around 30%, which is a useful reminder about what small samples do. The technique works. The headline number does not survive a real workload. |
The pattern across the whole table: the advice that survives checking is almost always remove something, and the advice that fails is almost always add something, and say it louder. I have written up why rules stop working and how to tell when one has expired separately, including what Anthropic did to their own system prompt when they hit this.
The theory I had, and why it was wrong
I want to show this one, because I believed it, I had a mechanism for it, and it is wrong.
My theory was placement. claude(.)md is delivered as a user message after the system prompt, which is documented and true. An output style is added to the system prompt itself, which is also documented and true. So, I reasoned, the memory file drifts further behind you as the conversation grows while the system prompt stays put, and that is why the rule fades.
Tidy story. The documentation kills it in one line:
Project-root CLAUDE.md survives compaction: after
/compact, Claude re-reads it from disk and re-injects it into the session.
So it does not vanish at the moment people assume it does. Be careful not to over-correct that into "instructions never fade", which is the mistake I had just made in the other direction: that quote is about /compact specifically, and I have written elsewhere that unenforced rules decay over a long session. Compaction is not the mechanism. Something else is. A Reddit thread had reached the same conclusion I had, and I treated that agreement as confirmation, when in fact the thread and I had both reasoned from a message role to a persistence behaviour that neither side had checked.
The check took one page of documentation. I had already written the post around the theory before I ran it.
What placement really changes
Placement does matter. It just does not do what I thought, and the balance runs the other way from the move I was about to recommend. I am naming the differences rather than totalling them, because the total went out wrong twice: first as "two of the three", then as "one, three and a wash".
claude(.)md | Output style | |
|---|---|---|
| Delivered as | User message after the system prompt | Added to the end of the system prompt |
| Extra adherence help | None documented | Triggers reminders during the conversation |
Surviving /compact | Content persists. Project-root and user-level files plus unscoped rules are re-injected from disk; path-scoped rules and nested files are not | Content persists, since the system prompt layer is reused |
| A mid-session edit | Takes effect at /clear, /compact or restart | Editing the style file's content needs a restart, since the terminal reads style files at startup. Switching to a different style applies at your next message, changed in v2.1.251 |
| Subagents | Loaded, full hierarchy, except Explore and Plan | Not applied. A subagent runs its own system prompt |
| Forks | Loaded | Applied, a fork inherits the parent system prompt |
Invoked skill bodies sit in the same table, re-injected but capped at 5,000 tokens each and 25,000 total with the oldest dropped first. That turns "put the important part near the top of a SKILL.md" from a style preference into a compaction rule.
What separates them, in plain terms. For the style: it triggers reminders during the conversation, and nothing in claude(.)md does. For claude(.)md: a mid-session edit reaches it at the next /compact, where an edited style file waits for a restart, and it reaches subagents where the style does not. Neither way: existing content survives compaction on both sides, and a fork inherits both. That is the whole comparison; the table above is the arithmetic.
The subagent row is the one that should stop you moving everything: ordinary subagents load your claude(.)md hierarchy and do not get your output style. Move your voice rules into a style and they come back written in the default voice. Two caveats in opposite directions. A fork inherits the parent's whole system prompt, so it does get the style. And claude(.)md is not the same as auto memory here: auto memory is not loaded into subagents at all, fork again excepted.
That is the opposite of the upgrade I was about to recommend.
This lives in primeline-ai/evolving-lite - the self-evolving Claude Code plugin. Free, MIT, no build step.
The five places a rule can live
There is no single right file. There are five, they behave differently, and choosing is mostly about how badly you need the rule obeyed.
Hooks are the only one that enforces anything. The documentation is explicit: if an instruction must run at a specific point, write it as a hook; a PreToolUse hook is the documented way to block an action regardless of what Claude decides. Everything else here is text that Claude reads and generally follows. I have written up what happens when you mistake an observer for a gate, and it applies directly: a rule in a markdown file is a request.
.claude/rules/ with paths: frontmatter is the underused one. A path-scoped rule loads only when Claude reads matching files, which attacks cause one head on. Instructions that only matter for your API handlers stop costing context and adherence everywhere else, which is the same budget problem behind context management generally. If your memory file is over 200 lines, this is where the excess should go.
There is a trade-off, and the documentation states it and its own remedy plainly. Path-scoped rules and nested claude(.)md files load into message history when their trigger file is read, so compaction summarises them away with everything else; they reload the next time Claude reads a matching file. The advice given is to drop the paths: frontmatter or move the rule to the project-root file if it must persist. Note the direction: toward the memory file, not away from it.
User-level ~/.claude/claude(.)md is a documented home for personal style. The scope table gives its purpose as personal preferences for all projects, with code styling preferences as the example. If you have felt vaguely wrong about keeping voice rules there, do not. It is the intended location, and it reaches subagents. With one exception: Explore and Plan skip the memory hierarchy entirely, and those are two of the ones Claude spawns most. The mitigation is in the same documentation: custom subagents you define yourself do load the full hierarchy, and any subagent can preload named instructions through its skills frontmatter field. So the gap is specific to those two built-ins rather than a general ceiling on how far your rules reach.
Both scopes survive compaction, and it took me two passes to state that correctly. The compaction table names "project-root CLAUDE.md and unscoped rules", which I first read as leaving the user-level file undocumented. The prompt-caching page closes it outright: "Your project-root and user-level CLAUDE.md files are read once at session start... The new content loads on the next /clear, /compact, or restart." So both are reloaded from disk at the same points, and the personal file the docs recommend for style preferences is not the weaker option.
The lesson from getting that wrong twice: silence in one table is not a documented gap. I published a gap that existed only in the page I happened to be reading.
The rules I run, and where I keep them
This block is mine, verbatim:
Language: [ENGLISH / GERMAN / etc.] by default. All output in this language unless I switch.
Banned in any text I will read:
- Em-dash (U+2014). Use "-" or " - " instead.
- Emojis, unless I ask.
- Filler phrases ("Great question", "Certainly", "I would be happy to").
Voice:
- I work alone. Never say "we / our / us". Always "I / my".
- Concise over exhaustive. If one sentence works, do not write three.
- Lead with the answer or the action, not the reasoning. Reasoning second, only when it adds value.
It lives in my user-level memory file, and having checked the subagent coverage above, it is staying there. Three things about it are worth stealing, and none of them is a specific rule.
Ban a thing, do not request a quality. "Be concise" competes with everything. "Em-dash (U+2014)" names a character, which makes it the most reliable line in the block. It has never failed, because there is nothing to interpret.
Pair every ban with its replacement. Each line above says what to do instead. A rule that only forbids leaves a gap, and the gap is where the invented vocabulary comes from.
Keep it short. Nine lines. I used to justify this as "short enough to stay obeyed", and I no longer claim that, because the one controlled test of file size against obedience found no effect. I keep it short for a reason that survives: a block I can re-read in ten seconds is a block whose contradictions I actually notice.
The fuller version, with the profile and sparring blocks around it, is in connecting Claude across surfaces. Worth saying plainly: in that post it is presented as profile content to sync between Claude.ai and Claude Code, filed as something to copy around rather than as an answer to "how do I stop this". I had the useful part filed under the wrong problem, which is a small version of what this whole post is about.
Three traps worth knowing
/output-style depends on which version you are running, and it has moved twice. It was deprecated in v2.1.73 and removed in v2.1.91, and on a build in that range typing it returns "/output-style isn't available in this environment" rather than the "Unknown command" you get for a genuine typo, so it resolves to something and still does not set your style. It has since been re-added in the v2.1.269 changelog, as /output-style [name], to list and switch styles including in headless sessions. My own machine is on v2.1.267, so it is not back here yet. The route that has worked throughout, on every version: /config and pick under Output style in the terminal, or set outputStyle in a settings file. In the desktop app /config opens Settings rather than a menu, so the settings field is the reliable route.
Custom styles used to drift back to the default voice on their own. v2.1.238 fixed custom, project and plugin output styles silently reverting mid-session. If you tried a style before that build, decided it faded, and went back to writing rules instead, that experience was a bug and is worth re-running.
keep-coding-instructions defaults to false. Write a custom output style without it and you drop Claude Code's built-in software engineering instructions: how it scopes changes, writes comments and verifies work. Set it to true whenever you are changing voice but still writing code.
---
name: Plain
description: Short answers, no filler, answer first
keep-coding-instructions: true
---
Lead with the answer. Reasoning after, and only when it changes what I do.
Styles live in three places: ~/.claude/output-styles/ for every project, .claude/output-styles/ for one, and the same path inside the managed settings directory for an organisation-wide policy. Plugins can ship them in an output-styles/ directory, and a plugin style carrying force-for-plugin overrides your own selection while that plugin is enabled.
Make the model show its drift
Everything above is about stopping Claude writing badly. This is the one thing I run that instead makes it show me when it has gone wrong, and it is the piece I would keep if I had to delete all the rest.
In Experiment A my rules moved the length of the answer and little else that design could resolve. So rather than fight that, I spent the length on structure. Every answer I get is shaped like this:
Task: why the CI pipeline went from 4 to 11 minutes
Checked: git log on .github/workflows, last 30 days
Tool: none fitted, this is a log read not a code change
One commit did it. 3f9a1c2 dropped the dependency cache to
fix a stale-lockfile error and never put it back.
install deps 38s -> 7m 38s <- the whole regression
unit tests 2m 04s -> 2m 04s
build 1m 18s -> 1m 18s
Everything except install is unchanged.
------------------------------------------------------------
WHAT I WOULD DO
Restore the cache, keyed on the lockfile hash, which also
fixes the stale-cache bug that commit was working around.
PATHS
1. key on hash of package-lock.json <- recommended
2. key on branch name. Faster to write, and goes stale
again the next time a dependency moves.
OPEN
- Runner size was never the issue. Leaving it alone.
- The stale-lockfile error is still unexplained. Worth
ten minutes once CI is green. Separate session.
Why the header is the important part
The three fields at the top are not a summary. They are drift detectors, and each one exposes a failure you cannot otherwise catch until the answer is already written and you have already read it.
Task catches the wrong job. The model prints, in a few words, what it thinks it is working on. When that does not match what I asked, I see it in the first line instead of at the end of four paragraphs. This is the failure that costs the most and announces itself the least.
Checked catches a skipped step. I have standing instructions to look things up before answering. An instruction like that fails silently: a skipped lookup looks exactly like a lookup that found nothing. Forcing the field to be printed turns the skip into something visible, and "not checked, because the file was named in the question" is a perfectly good answer. An honest admission beats an absent line.
Tool catches the reach for the wrong instrument. Same logic. Naming the tool it used, or naming that none fitted, is cheap. Silently not considering one is what I want to see.
There is a fourth one in my own version that I will mention because it surprised me: the style tells the model to use my name once, in the recommendation. It sounds like a nicety. It is a tripwire. When an answer stops being written to a person and turns into a report, the name is the first thing that disappears, so its absence tells me the register has slipped before I have consciously noticed.
What this is and is not
Copy the shape, not my fields. The specific three are mine because they match the three things I most often catch going wrong. Yours will be different: if the failure that costs you most is unrequested refactors, make the field Files I will touch.
I have not isolated its effect. In the runs above, the output style and the command catalog were removed together in the same step, so I cannot separate what the style contributed from what the catalog did. I can tell you the mechanism and why I keep it. I cannot show you a measurement of this specific style working.
It argues against part of this post. The measured section says a short declarative file beat my long imperative one, and this style is neither short nor purely declarative. The tension is real. My best account of it is that bans and structure are doing different jobs: a ban needs to be short to be obeyed, a shape needs to be complete to be reproducible. I have not tested that account.
What none of this fixes
The threads at the top are largely about how the model writes, and configuration only reaches part of that.
Filler phrases, banned characters and burying the answer are all rule-shaped, and they respond well. Invented vocabulary does not, because a model reaching for a compact term is not breaking a rule you could have named in advance. One line asking for the plain word rather than the clever one helps, and does not solve.
A style is not inherently a brevity tool; it is a way to change the voice, in either direction. There are five built-ins: Default, plus Proactive, Concise, Explanatory and Learning. Two of them, Explanatory and Learning, are documented to produce longer responses than Default by design. Concise is the one that goes the other way.
The length of the answer is not the size of the work. This is the single clearest thing in my own numbers: the amount the model generated moved by 9% between conditions while the chat answer swung by a factor of two, and the documents it wrote were about seven times longer than necessary in every condition. Whatever a style or a rule does to the reply on your screen, it is not changing how much work happens underneath. And note which way my own rules went: they contained brevity instructions and still produced the longest answers in the set, so do not assume that asking for brevity delivers it.
And none of it is enforcement. Memory files, rules and styles are all text Claude reads. If a rule genuinely must hold, it belongs in a hook.
There is now a number on how wide that gap is. A survey of 481 public CLAUDE.md files found that only 4 to 16% of the security rules it retrieved had a matching built-in control that could actually enforce them, and under the strictest matching standard 4.4%. Their extraction caught about two thirds of the eligible rules, so the rate describes the rules they captured. The authors call the file a write-only channel, and that phrase is the whole problem in four words: you write a rule, and nothing ever tells you whether anything acted on it. That is the same shape as the within-session decay above, and the same remedy applies. If it matters, do not write it down. Wire it up.
Honest scope
Two different kinds of claim are mixed in this post and they do not carry the same weight.
The mechanisms are read from Anthropic's documentation, not measured by me. Which file loads where, what survives compaction, what reaches a subagent, what a style is appended to: all documentation, and all re-checked against v2.1.267 for this update rather than taken from the first version of this post.
The numbers in the measured section are mine, and they are narrower than they look. I have now instrumented rule content against behaviour, which the first version of this post said I had not. Three things I still cannot separate, and each one limits a claim above. There is no placebo condition, so removing rules removes both what they say and the room they take up. The bottom step of the chart removes my output style and my command catalog together, so neither gets credit for that 82 points on its own. And I have not compared the same rule in different files, which is the question the placement table is about and the one I most want an answer to.
So when I say my rules moved length and cost and not correctness, that is Experiment A only: eight task families on one small repository at one effort level, with a correctness check every condition passed. It does not generalise to your codebase, and it is emphatically not a claim that rules do not work. Experiment B, on different tasks, moved false completion claims by a factor of fifteen with a single deletion. Any sentence of mine that sounds like "rules do nothing" is a sentence that has dropped its experiment, and you should distrust it.
Two claims here were hard to get right and both failed in a summary layer rather than in the argument. The comparison of the two placements was published with a wrong total twice, as "two of the three" and then as "one, three and a wash", so it no longer carries one at all. And whether the user-level file survives compaction went wrong twice in opposite directions - first published as "unstated", then quietly omitted from an FAQ answer that was asked about the class. Both are corrected above; I mention them because a reader deserves to know which parts of a post were slippery.
I have left the wrong theory in this post on purpose. The first draft argued that voice rules belong in an output style because a memory file drifts out of reach as context fills. That is wrong, the documentation says so in one line about /compact, and I only found it because an external model attacked the draft before it shipped. If you have read that argument somewhere, including from me, it does not survive checking.
Thread counts are one snapshot of two subreddits sorted by points, and individual dates are approximate. The thread table was collected on 2026-08-15 and I have deliberately not refreshed it, because it is the evidence for what people were asking at that moment rather than a live count.
Version numbers. The first version of this post was written against Claude Code v2.1.233. This update was checked against v2.1.267, which is what I run, with one exception: the /output-style re-add is a v2.1.269 changelog entry that has not reached my machine, so I am reporting it from the changelog rather than from use. Four of the five stale facts fall inside that 34-build gap and the fifth arrived just after it. Take the general lesson rather than the number: on a tool that ships this often, any advice you read about it, including this, needs its version stated. Run claude --version before you trust a version claim, mine included.



