CCAO-F · module 7 of 7 · 10% of the exam

Troubleshooting and optimization

Weight: 10 percent of the exam (roughly 6 of 60 questions). Tests whether you can diagnose why an output disappointed, work sensibly within context limits, make recurring tasks reliable, and know when to start over rather than push on.

What the exam expects

This is the diagnostic domain. A scenario describes a disappointing result and asks for the root cause, the first thing to check, or the change that most improves it. The wrong answers are rarely absurd. They are the things people actually try: upgrade the model, restart the conversation, ask more forcefully, run it again and hope.

The organizing principle for the whole domain is that the cheap fixes come first, and they are almost always the right ones. Supplying context, specifying the output, and clearing stale conversation state cost nothing and take seconds. Changing plan or model tier costs money and usually addresses a cause that was not the problem.

The three causes of poor output

Nearly every quality complaint traces to one of three causes, and learning to tell them apart from the symptom is most of what this domain rewards.

CauseSymptom in the scenarioFix
Missing contextOutput is competent but generic, or states internal facts that are wrong or inventedSupply the actual source material; ground answers in it and ask for attribution
Ambiguous instructionsOutput is specific but the wrong shape: wrong length, wrong format, wrong deliverableState audience, format, length, and what must be preserved
Contaminated or stale contextA rejected idea keeps resurfacing; earlier decisions get lost; superseded facts get citedStart a fresh conversation carrying only what is still true, or curate the knowledge base

Missing context is the diagnosis when the request supplied nothing that only this situation could have. "Write a launch announcement for our new product" produces something that could describe any product in the category, because that is genuinely all the model was given. The same cause produces fabrication: asked about an internal escalation policy it was never shown, Claude has nothing to ground on. The fix is to supply the policy and require answers to be attributed to it, so unsupported claims become visible. A heavier model tier reduces confident errors somewhat but cannot conjure information it has never seen.

Ambiguous instructions is the diagnosis when the output is specific but the wrong thing. "Shorten this report" names no target length, no output form, and no constraint that all five recommendations must survive, so three separate gaps each produced a mismatch. Watch especially for ambiguous scope in the verb: "review this document" can reasonably mean critique it or rewrite it, and Claude picking the one you did not want is a specification failure rather than a defect.

Contaminated or stale context is the diagnosis when something that should be gone keeps influencing the output. Everything said in a conversation stays in context, so a pricing approach rejected an hour ago keeps exerting pull on later drafts. Repeating the rejection more forcefully usually makes it worse, because it adds still more text about the rejected approach. The equivalent inside a Project is a knowledge base holding both the retired policy and its replacement.

Context windows in practice

The context window is Claude's working memory for one conversation. Paid plans support windows ranging from 200K tokens up to 1M on the newest models, and a portion is always reserved for the response, so the longest usable conversation is slightly smaller than the headline figure.

What matters for the exam is the practical behaviour rather than the numbers:

  • Quality degrades before any limit is announced. Long sessions start losing track of earlier decisions well before an error appears. Waiting for a system message to tell you to start a new conversation is a poor trigger.
  • A conversation approaching the limit can have earlier messages summarized so it can continue, and the chat history itself is preserved. This is designed accommodation, not a fault and not a usage penalty. It is still a good moment to consider a fresh conversation.
  • More context is not uniformly better. Irrelevant material dilutes attention and consumes the window. Carrying an unrelated topic forward is pure overhead.
  • Detail buried in the middle of a very long document is harder to surface. The practical remedy is to direct attention explicitly by naming the section or page range, and to require attribution to specific passages so gaps become visible. Re uploading the file changes nothing, and changing file format changes packaging rather than the underlying difficulty.

When to start a new conversation

Start freshStay put
The topic has changed entirely and none of the accumulated context is relevantThe follow up depends directly on the analysis just completed
An approach was explored and rejected and keeps resurfacingA draft is close and needs targeted correction
A long session has started losing track of earlier decisionsYou want a different tone for the next section

When you do start fresh, seed the new conversation deliberately: a short written summary of the decisions that still stand, plus the current version of the artifact. That preserves what matters and discards the accumulated noise. Re pasting the entire prior conversation defeats the purpose by doubling the history competing for attention.

The mirror image matters just as much. When a draft is broadly right but the tone is too formal and one section is out of scope, the context is exactly what you need, and two sentences of specific correction beat starting over. Discarding useful context is as much a mistake as carrying useless context forward.

Improving reliability

For a task that recurs, the goal is to reduce run to run variance rather than to accept it or to brute force it.

  • Specify the output structure explicitly, and show one example of a good result. Demonstrating the target is among the strongest levers available to a non technical user.
  • Ask Claude to state its assumptions or ask clarifying questions before producing a long deliverable. This turns a silent guess about audience or scope into a visible decision you can correct in one line, before effort is spent.
  • Require attribution on anything a reviewer will spot check. Traceability is what makes review fast, because a reviewer can verify a sample quickly and any claim without a source stands out as a risk.

Three approaches that look like reliability engineering and are not. Running the same prompt three times and picking the best result costs three times the usage and still leaves a human choosing between drafts every cycle. Generating it three times and publishing the version the runs agree on measures consistency rather than correctness, and repeated runs can agree on the same error. Asking Claude to rate its own confidence and skipping review when it is high substitutes self assessment for verification, which is not acceptable on compliance work.

Also note that length is not precision. A longer vague prompt generally performs worse than a short specific one, and elaborate polite phrasing adds tokens without adding specification.

Optimizing usage

When a team repeatedly exhausts its allowance, attack the drivers before spending money:

  • Move recurring reference documents into Projects so they are supplied once instead of re attached daily.
  • Batch related questions into single well specified messages. A long chain of short clarifying turns re sends the accumulated conversation each time and generally consumes more, not less.
  • Reserve heavier model tiers and higher effort settings for work that needs them, and turn off tools that are not relevant to the task.
  • Track consumption in the usage settings so decisions are based on where the allowance is actually going.

Only after those are exhausted does upgrading a seat or buying usage credits make sense. A wasteful pattern on a larger allowance is still wasteful.

How to think through the question

  1. Read the symptom precisely. Generic, wrong shape, or haunted by something that should be gone? That maps directly to missing context, ambiguous instructions, or contaminated context.
  2. Check what the scenario actually establishes. If it says this is the opening request, "the conversation is too long" is not available. If it says the retired policy is in project knowledge, "start a new conversation" cannot help.
  3. Ask what the question wants: root cause, first step, or the change that most improves it. A root cause question rewards diagnosis; a "most improves" question rewards the highest leverage fix, which is not always the first thing you would do.
  4. Eliminate options that spend money before the free fixes are tried, and options that treat a symptom while leaving the cause.
  5. Eliminate mechanisms that do not exist: usage limits silently degrading quality, time of day affecting response length, longer questions preventing fabrication, attachments behaving differently from pasted text.

Worked example. "A manager insists the team upgrade to a more capable tier because the answers are not good enough. On inspection, prompts are one line long, no source documents are ever supplied, and nobody has set up a Project."

Step 1: the symptom is general dissatisfaction, but the scenario hands you the diagnosis in its detail, namely no context and no specification. Step 2: what is established is that the cheap fixes are entirely untried. Step 3: the question asks what the manager should be told, so it wants the recommendation. Step 4: eliminate "upgrade immediately," which spends money on a cause the evidence does not support. Step 5: the interesting distractor is "upgrade and restructure the prompts simultaneously," because doing both sounds prudent. It is wrong for two reasons: it locks in a cost that may be unnecessary, and changing two variables at once makes it impossible to learn which one mattered. The answer is to fix context and specificity first. Notice the general shape: when a scenario carefully lists untried free improvements, those details are the answer.

Exam traps

  • Reaching for a model tier upgrade. It is the most available explanation for "not good enough" and rarely the actual cause. Tier is the expensive lever and seldom the binding constraint when prompts lack context.
  • Doing two things at once because both sound sensible. Upgrading and fixing the prompts together is a favourite distractor. It obscures cause and spends unnecessarily.
  • Repeating a correction more forcefully. It feels like the natural escalation and usually reinforces the very material you want gone.
  • Restarting the conversation as a universal remedy. It is right for topic changes and contamination, and wrong when the context is exactly what makes a targeted correction possible.
  • Believing usage limits degrade quality. They gate access. They do not silently shorten or weaken responses.
  • Treating automatic summarization of a long conversation as a fault or a penalty. It is designed behaviour and the history remains available.
  • Assuming more context is always better. Irrelevant context costs attention and window space.
  • Confusing length with precision. Longer prompts, more polite prompts, and repeated instructions are not more specific prompts.
  • Accepting self reported confidence in place of verification on work that carries consequence.
  • Re running a prompt unchanged and hoping. Nothing about the input changed, so this is luck plus usage.
  • Expecting a preference stated in one conversation to persist into the next. Persistence comes from account wide or project instructions.

Quick reference

  • Generic output means missing context. Wrong shape means ambiguous instruction. Recurring rejected ideas mean contaminated context.
  • Ambiguous verbs ("review", "shorten", "clean up") are a specification failure, not a defect.
  • Fabricated internal facts are fixed by supplying the source and requiring attribution, not by a heavier tier.
  • Contradictory project knowledge is fixed by curation, not by an instruction to prefer newer documents.
  • Start fresh on a topic change, on a rejected direction that keeps returning, and on a long session that is drifting. Seed the new conversation with the decisions that still stand plus the current artifact.
  • Stay put when the follow up depends on the work just done or the draft only needs targeted correction.
  • Reliability levers: explicit output structure, one worked example, assumptions stated up front, attribution for reviewability.
  • Usage levers: Projects for recurring documents, batched well specified messages, right sized tier and effort, unused tools off, then credits or an upgrade.
  • Non mechanisms to reject: quality degrading from usage limits, time of day effects, longer questions preventing fabrication, attachments behaving differently from pasted text.

Ready to test this domain?

Drill mode gives instant feedback: pick a wrong answer and you immediately see why it is wrong.

Start the drill

Not an official source. This is a free, independent study resource from siasola, built by an engineer who sat these exams and wanted better prep material to exist. It is not affiliated with, endorsed by, or sponsored by Anthropic. Claude is a trademark of Anthropic, PBC. Exam facts follow the official exam guides; registration for the real exams happens through the Anthropic Partner Academy and Pearson VUE, not here.