The Definitive Guide to the WBS12 20-Mark Evaluate Question

PART 1: HOW THE 20-MARKER DIFFERS FROM THE 8 AND 10

20 min read

The relationship between the three question types follows a strict hierarchy. Each step up adds exactly one requirement to the level below it:

DISCUSS (8mk)     Two-sided argument with chains
      +
ASSESS (10mk)     + Supported judgement
      +
EVALUATE (20mk)   + Full bilateral development
                  + Effective conclusion with
                    specific recommendation

Most students treat the 20-marker as "a longer assess." It is not. The two additions — full bilateral development and an effective conclusion with a recommendation — are structural requirements that the 8 and 10 markers do not demand. They are not optional extras. They are what Level 4 is gated behind.

The other critical difference is scale. The 20-marker has a 5-band mark range within Level 4 alone (16–20). The difference between 16 and 20 is real, meaningful, and determined by very specific things. Understanding exactly what separates each mark within Level 4 is what this guide is for.


PART 2: THE COMPLETE LEVEL DESCRIPTORS — WHAT THEY MEAN IN PRACTICE

LEVEL 1  (1–5 marks)
─────────────────────────────────────────────────────
What it says:  Isolated knowledge. No real chains.
               Weak or no application to extract.
               Generic assertions.

What it looks like:
"Amazon failed in China because of competition.
External factors like market conditions also played
a role. Internal factors such as poor management
were important too. Overall both internal and
external causes were responsible."

Why it's Level 1: No chain exists anywhere.
No extract data used. Lists categories rather
than explaining mechanisms. Generic assertions
stated as fact without reasoning.

Typical mark: 3–4/20


LEVEL 2  (6–10 marks)
─────────────────────────────────────────────────────
What it says:  Knowledge applied to the extract.
               Chains present but incomplete.
               Generic or superficial evaluation.

What it looks like:
An answer that identifies causes, applies some
extract data, builds partial chains (2-stage:
cause → effect but stops before the specific
business outcome), and attempts a conclusion
that is either generic ("overall both internal
and external factors contributed") or simply
restates the arguments without weighing them.

Why it's Level 2: Chains present but stop one
stage early. Extract data sometimes used but
sometimes just mentioned. Evaluation is
either absent or unconditional. No genuine
bilateral development — one side dominates.

Typical mark: 7–8/20


LEVEL 3  (11–15 marks)
─────────────────────────────────────────────────────
What it says:  Thorough knowledge. Effective use
               of context throughout. Developed
               chains. Some bilateral development.
               Quant/qual data used. Conclusion
               present but not effective or decisive.

What it looks like:
Two sides developed with full chains. Extract
data woven in throughout. A conclusion that
identifies a position but either:
(a) doesn't state a condition under which
    the conclusion would change, OR
(b) doesn't propose a specific recommendation
    tailored to this business, OR
(c) the competing argument is underdeveloped
    relative to the main argument.

Why it's Level 3: The analytical quality is
high but the EVALUATION component is
incomplete. "Unlikely to show significance
of competing arguments" is the key phrase —
one side always dominates or the conclusion
doesn't fully weigh both against each other.

Typical mark: 12–13/20


LEVEL 4  (16–20 marks)
─────────────────────────────────────────────────────
What it says:  Thorough knowledge. Full chains.
               FULLY bilateral development.
               Effective conclusion with
               recommendation. Full awareness
               of significance of competing
               arguments.

What it looks like:
Two sides EQUALLY developed with complete
chains. Both sides use extract data. The
significance of each argument is assessed
within the body — not just presented. The
conclusion weighs both sides, states a
decisive position, names a condition, and
proposes a specific recommendation using
this business's data. The conclusion adds
something new — it does not merely restate
what P1–P3 argued.

Typical mark: 16–18/20 (16 = L4 entry)
              18–19/20 (strong L4)
              20/20 (perfect — rare)

PART 3: WHAT SEPARATES EACH MARK WITHIN LEVEL 4

This is the most important section for a student already performing at Level 3. The jump from 15 to 16 is clearing Level 4 entry. The jump from 16 to 20 is a different problem entirely.

16/20 — LEVEL 4 ENTRY
Both sides present with full chains.
Extract data used on both sides.
Genuine bilateral development.
Conclusion present and commits to a position.
MISSING: condition stated / recommendation
is generic / significance of one argument
not fully addressed in the body.

17/20 — SOLID LEVEL 4
Everything at 16 PLUS:
Condition stated in conclusion
("only if..." / "provided that...").
Significance addressed for both arguments.
MISSING: recommendation is still slightly
generic OR one chain slightly weaker.

18/20 — STRONG LEVEL 4
Everything at 17 PLUS:
Specific recommendation using business
name and extract data.
Both chains equally developed and complete.
MISSING: conclusion adds something new
beyond what was argued in P1–P3 /
very minor chain imprecision somewhere.

19/20 — NEAR PERFECT
Everything at 18 PLUS:
Conclusion adds a new insight — a time-
based condition (short run vs long run),
a secondary recommendation, or a
limitation of the conclusion itself.
Both arguments fully developed with
significance explicitly weighed.

20/20 — CEILING
Every component perfect. Conclusion is
decisive, conditional, extract-anchored,
proposes a specific tailored recommendation,
and adds something beyond the body.
No weak chains anywhere. Extract data
woven naturally throughout — not bolted on.

The difference between 15 and 20 is concentrated in two places: how you evaluate within the body (not just in the conclusion) and the quality of the effective conclusion. Most students who score 13–14 are already writing two decent chains. They are losing marks because their evaluation is superficial and their conclusion is a restatement, not a recommendation.


PART 4: THE FOUR-PARAGRAPH SKELETON — NON-NEGOTIABLE

The 20-marker has a mandatory structure. Deviating from it costs marks because the structure exists to satisfy the Level 4 descriptor requirements in the most efficient order.

TIME BUDGET: 28 minutes HARD CEILING
(Never exceed this. Q1+Q2 are worth 60 marks.
Q3 is worth 20. Over-running Q3 is the most
expensive time mistake on the paper.)

PARAGRAPH 1 — Main argument (7 minutes)
Function: Your strongest, best-evidenced argument.
Content:
  - State the factor/cause/argument clearly
  - Build a 5-stage chain (see Part 6)
  - Weave extract data into stages 2–3
  - Reach a specific, named business outcome
  - State the significance: why does this
    argument matter for this business specifically?

PARAGRAPH 2 — Evaluate P1 (4 minutes)
Function: Challenge your own main argument.
Content:
  - "However, this depends on whether..."
  - Identify the condition that weakens P1
  - Build a 2–3 stage mini-chain showing why
    P1 is limited or contingent
  - Use extract evidence to anchor the challenge
  - End: "This holds only if [condition]."

THIS IS WHAT MOST STUDENTS MISS. Evaluating your
OWN argument before moving to the counter-argument
is what "bilateral development" means in practice.
It shows the examiner you understand the limits
of your position — not just its strengths.

PARAGRAPH 3 — Competing argument (7 minutes)
Function: A DISTINCT second argument, genuinely
competing with P1 — not just a different angle
on the same point.
Content:
  - Same structure as P1
  - Must use DIFFERENT extract data from P1
  - Must reach a different business outcome
  - Must directly challenge whether P1 is
    the more important explanation

PARAGRAPH 4 — Evaluate P3 + Effective Conclusion (6 minutes)
Function: Challenge P3 AND deliver the judgement.
Content part 1 (2 min):
  - Challenge P3 with a mini-chain
  - Show its limitations using extract evidence
Content part 2 — THE EFFECTIVE CONCLUSION (4 min):
  - Weigh P1 vs P3 — state which is stronger
  - Reference specific extract evidence for why
  - State a condition: "only if..." / "unless..."
  - State a counter-condition: what would change
    the conclusion
  - Propose a SPECIFIC RECOMMENDATION using
    this business's name and extract data
  - Add something NEW — not a restatement of
    what you already argued

The emergency rule : If you are running out of time and have not finished P4, write the effective conclusion immediately — even if it means leaving P3 or P4 evaluation incomplete. A partial P4 with an effective conclusion always outscores a complete P4 without one. The conclusion alone can push from Level 3 to Level 4.


PART 5: THE 5-STAGE CHAIN — THE CORE TECHNICAL SKILL

Every paragraph must contain at least one complete 5-stage chain. This is what separates Level 2 (2-stage chains) from Level 3 (3–4 stage chains) from Level 4 (5-stage chains with significance).

STAGE 1 — KNOWLEDGE TRIGGER
State the factor/cause clearly.
Signal words: "One internal cause was..." /
"A key external factor was..."

STAGE 2 — CONTEXT ANCHOR (extract data)
Connect to specific extract evidence.
Signal: "As Extract F confirms..." /
"Given that..." / "Since..."
The data must be WOVEN IN here — not mentioned
as a separate sentence before or after.

STAGE 3 — MECHANISM
Explain HOW the factor causes something.
Signal: "This means that..." /
"As a result..." / "Therefore..."
This is the cause-and-effect link.
Most students stop here. Don't.

STAGE 4 — SPECIFIC BUSINESS OUTCOME
Name the exact consequence for THIS company.
Not "revenue fell" — but "Amazon's Chinese
Kindle revenue fell despite the market growing
18% to $6bn in 2021, meaning Amazon was
losing share even in an expanding market."
The specificity is what earns the mark.

STAGE 5 — SIGNIFICANCE
Why does this matter MORE than the other
argument? Why is this the decisive factor?
Signal: "This is particularly significant
because..." / "This matters most because..."
"Compared to [other argument], this factor..."
This is the evaluation component WITHIN
the analysis. It is what prevents the
competing argument from having equal weight.

Applied to Q3 of this paper (Amazon Kindle China):

"One significant internal cause of Amazon's failure was its inability to adapt the Kindle to Chinese consumer preferences [Stage 1]. As Extract F confirms, the Kindle neglected online fiction titles — including popular series such as Harry Potter — which were either incomplete or missing entirely, while Chinese domestic competitors were offering content libraries tailored to local reading habits [Stage 2]. This meant that even consumers who were aware of the Kindle and willing to purchase an e-reader had limited reason to choose Amazon's device over domestic alternatives such as iFlytek or Huawei [Stage 3]. As a result, Amazon progressively lost market share in China's digital reading market — which, as Extract E shows, generated over $6bn in revenue in 2021, growing 18% from 2020 — meaning Amazon was exiting a rapidly expanding market precisely because its product had failed to meet the content needs of Chinese readers [Stage 4]. This internal failure is particularly significant because it was entirely within Amazon's control — unlike regulatory or competitive pressures — meaning it represents a missed opportunity that competent product management could have avoided [Stage 5]."

That single paragraph earns marks at every stage. Each sentence does specific work. Nothing is padding.


PART 6: THE EFFECTIVE CONCLUSION — THE LEVEL 4 GATEKEEPER

The examiner report for WBS12 states this in every series: "To achieve the top level, an effective conclusion is sought." The mark scheme Level 4 descriptor says the response must lead to "a supported judgement." These are the same requirement stated differently.

An effective conclusion has five non-negotiable elements. Miss any one and it is not an effective conclusion — it is a Level 3 conclusion.

ELEMENT 1 — WEIGHING
State which argument (P1 or P3) is stronger
and why. This is not a restatement. It is a
direct comparison: "P1 is more significant than
P3 because..."

ELEMENT 2 — EXTRACT EVIDENCE
The weighing must reference specific extract data.
Not just naming the business — quoting or
paraphrasing a specific figure or fact.

ELEMENT 3 — CONDITION
State the condition under which the conclusion
holds: "only if..." / "provided that..." /
"unless..." / "this holds only where..."
Without a condition, the conclusion is
unconditional — which is Level 3 maximum.
This is confirmed in every examiner report.

ELEMENT 4 — COUNTER-CONDITION
State what would change the conclusion:
"However, if [condition], then [other argument]
would become the more significant factor."
This demonstrates full awareness of competing
arguments — exactly what the Level 4 descriptor
requires.

ELEMENT 5 — SPECIFIC RECOMMENDATION
Propose a concrete action for THIS business
using its name and extract data. This must
add something new — not restate P1 or P3.
It is the answer to: "given all of this,
what should this business actually do?"

What each element looks like for the Amazon/Kindle Q3:

"On balance, internal causes were more responsible for Amazon's decision to withdraw the Kindle from China than external ones [WEIGHING BEGINS]. Although domestic competition from Xiaomi, iFlytek and Huawei presented a genuine external challenge [acknowledging other side], the Kindle's failure to adapt — neglecting popular fiction titles and failing to develop the product as the market evolved over 10 years — represents internal decisions that Amazon controlled and could have changed [ELEMENT 1 — WEIGHING]. This judgement is supported by the fact that China's digital reading market grew to $6bn in 2021, up 18% from 2020 — demonstrating that external market conditions were actually favourable, meaning the failure must primarily be attributed to Amazon's internal choices [ELEMENT 2 — EXTRACT EVIDENCE]. This conclusion holds only if we accept that Amazon had the resources and market intelligence to adapt its product — which its status as one of the world's largest technology companies suggests it did [ELEMENT 3 — CONDITION]. However, if government regulation had directly restricted the Kindle's operation in China — as it did with other Western platforms — then external causes would have been the decisive factor, since no amount of internal adaptation could overcome regulatory barriers [ELEMENT 4 — COUNTER-CONDITION]. Amazon should therefore have invested in building a localised Chinese content library and partnerships with domestic publishers such as CITIC Press Group, whose good relationship with Amazon is confirmed in Extract F — this internal fix, had it been made earlier, may have prevented withdrawal entirely [ELEMENT 5 — SPECIFIC RECOMMENDATION]."

That conclusion is approximately 230 words. It takes 4 minutes to write if you know what you are doing. It is the difference between 13 and 18.


PART 7: BILATERAL DEVELOPMENT — THE MOST MISUNDERSTOOD REQUIREMENT

Level 4 requires "fully bilateral development." Most students interpret this as "write about both sides." That is Level 3. Bilateral development at Level 4 means something more specific:

Each side must be evaluated as well as argued.

LEVEL 3 BILATERAL:          LEVEL 4 BILATERAL:

Side A argued               Side A argued (P1)
Side B argued               Side A evaluated — its limits
Conclusion                  and conditions (P2)
                            Side B argued (P3)
                            Side B evaluated — its limits
                            and conditions (P4 start)
                            Conclusion that weighs A vs B

The evaluation of each side — P2 and the first part of P4 — is what Pearson means by "bilateral." You are not just presenting two arguments. You are actively testing both arguments against the evidence and finding their limits before concluding which one survives that test better.

This is why the four-paragraph skeleton is structured the way it is. P1 and P3 are your arguments. P2 and P4 are you stress-testing your own arguments. The conclusion is where you state which argument survived better and why.


PART 8: THE TWO-ELEMENT TRAP — UNIQUE TO Q3

Some Q3 questions name two specific factors in the question itself. The Oct 2023 paper does this: "Evaluate whether Amazon's decision to stop selling its Kindle in China was due to internal or external causes of business failure."

This is not two sides of a discussion — it is two categories you must address. The examiner report for series where this occurs states explicitly that students who addressed only one category were limited to Level 2 regardless of the quality of their writing on that one category.

The rule: If the question names two factors, strategies, or categories — you must address both in your answer. One cannot substitute for the other. A perfect analysis of only internal causes cannot reach Level 3 if external causes are absent.

For the Amazon question this means:

  • P1 must cover internal causes (product failure, marketing failure, complacency)
  • P3 must cover external causes (domestic competition, market conditions, other Western companies leaving China)

If your essay only argues one category — however well — you are capped at Level 2.


PART 9: HOW TO USE EXTRACT DATA — THE MOST IMPORTANT TECHNICAL RULE

The same rule from the 6 and 8-marker applies here, but it is even more important at 20 marks because the volume of data in the extract is greater and the expectation of its use is higher.

The three ways students misuse extract data:

MISTAKE 1 — STATING (earns zero App marks):
"According to Extract E, Amazon's digital reading
market generated $6bn in 2021."

This is copying the extract. The data is
mentioned but not used. Zero marks.

MISTAKE 2 — BOLTING ON (earns partial marks):
"Domestic rivals launched e-reader devices.
This reduced Amazon's market share. According
to Extract E, the market grew 18% in 2021."

The data is present but sits outside the chain.
It does not do work in the argument.

MISTAKE 3 — WEAVING IN (earns full marks):
"Although China's digital reading market grew
18% to $6bn in 2021 (Extract E), Amazon's
market share fell — meaning Amazon was losing
ground in an expanding market, which is
characteristic of internal failure rather
than adverse market conditions."

The data is embedded in the chain. It proves
the point being made. It is doing argumentative
work. This is what full application marks look like.

Practical rule: Every piece of extract data you use should be answering the question "how does this data prove my point?" If it is not answering that question, it is decoration — and decoration earns zero.


PART 10: MARKS WITHIN LEVEL 4 — PRECISE MARK PLACEMENT

MARK    WHAT'S PRESENT / WHAT'S MISSING

16/20   L4 entry. Both sides argued. Both chains
        complete. Bilateral present. Extract used
        on both sides. Conclusion commits to a
        position. MISSING: condition in conclusion
        OR recommendation is generic OR one
        evaluation (P2/P4) is weak.

17/20   Everything at 16 PLUS condition stated.
        P2 and P4 both present with mini-chains.
        MISSING: recommendation still generic
        OR conclusion adds nothing new beyond
        the body OR one chain slightly underdeveloped.

18/20   Everything at 17 PLUS specific recommendation
        with business name and extract data.
        Both P2 and P4 challenge the argument
        with extract evidence. MISSING: conclusion
        adds nothing truly new / very minor
        imprecision in one chain.

19/20   Everything at 18 PLUS conclusion adds a
        new dimension — time-based condition
        (short run vs long run), a secondary
        recommendation, or a limitation of the
        conclusion itself. Both arguments equally
        developed. All chains at stage 5.

20/20   Perfect execution of every component.
        Chains complete throughout. Extract woven
        naturally — never bolted on. P2 and P4
        both challenge with extract evidence.
        Conclusion decisive, conditional, specific,
        and adds something beyond the body.
        No weak sentence anywhere.

PART 11: THE COMPLETE 20-MARKER IN 90 SECONDS — EXAM DAY CHECKLIST

BEFORE YOU WRITE (3 minutes planning):
□ What are the TWO distinct arguments the
  question requires? (If the question names
  two categories — address both.)
□ What extract data supports each argument?
  Note 2 specific facts/figures for each side.
□ Which side do I think is stronger? Why?
□ What is the condition under which my
  conclusion would change?
□ What specific recommendation can I make
  for this business using extract data?

PARAGRAPH 1 (7 minutes):
□ State the main argument clearly
□ Weave extract data into the chain at stage 2–3
□ Complete the 5-stage chain to specific outcome
□ State significance: why this matters most

PARAGRAPH 2 (4 minutes):
□ Challenge your own P1 argument
□ Use extract data to anchor the challenge
□ State the condition: "this holds only if..."

PARAGRAPH 3 (7 minutes):
□ State the competing argument — distinct from P1
□ Use DIFFERENT extract data from P1
□ Complete the 5-stage chain
□ State significance relative to P1

PARAGRAPH 4 (6 minutes):
□ Challenge P3 briefly with extract evidence
□ EFFECTIVE CONCLUSION — all 5 elements:
  ① Weigh P1 vs P3 — state which wins and why
  ② Reference specific extract evidence
  ③ State the condition: "only if..."
  ④ State counter-condition: "unless..."
  ⑤ Specific recommendation — business name
     + extract data + something NEW

FINAL CHECKS:
□ Have I addressed both factors/categories
  if the question named two?
□ Is extract data WOVEN IN — not bolted on?
□ Does P2 evaluate P1 (not just add to P1)?
□ Does the conclusion add something new?
□ Is there a condition in the conclusion?
□ Is the recommendation specific to THIS
  business using extract evidence?
□ Have I stayed within 28 minutes?

TIME: 28 minutes HARD CEILING
WORD COUNT: 450–550 words

PART 12: SIDE BY SIDE — LEVEL 2, LEVEL 3, LEVEL 4

Using the Amazon/Kindle question directly:


LEVEL 2 answer (typical: 7–8/20):

"Amazon failed in China due to internal causes. It did not adapt its product to Chinese consumers and did not have the right content. This meant fewer people bought the Kindle. Externally, there were many competitors like Xiaomi and Huawei. These took market share from Amazon. The market was growing, with $6bn in revenue in 2021. Overall both internal and external causes contributed to Amazon's failure in China."

Why Level 2: Two sides present. Some extract data mentioned. But chains are 2-stage (cause → effect with no mechanism or specificity). Conclusion is generic — "both contributed" is not a judgement. No condition. No recommendation.


LEVEL 3 answer (typical: 12–13/20):

"Internally, Amazon failed to adapt the Kindle to Chinese tastes. Extract F confirms it neglected popular fiction titles like Harry Potter, which were incomplete or missing. This meant Chinese readers had little reason to choose Kindle over domestic alternatives, so Amazon lost market share. This is an internal cause because Amazon controlled its content library.

However, the rise of domestic competitors like Xiaomi, iFlytek and Huawei represents an external cause. These businesses launched their own e-reader devices that better met Chinese needs, including colour screens and smaller formats. As Extract E shows, this led to Amazon losing market share to domestic rivals. This is an external cause because Amazon could not control what competitors built.

Overall, both internal and external causes contributed. The internal causes were perhaps more significant because Amazon had a good relationship with publishers, meaning it had the resources to adapt but chose not to."

Why Level 3: Full chains present. Both sides developed. Extract data used on both sides. Conclusion commits to a position and gives a reason. But: P1 is not evaluated before P3 — no bilateral evaluation. Conclusion has no condition. No specific recommendation. The conclusion does not add anything new beyond restating the P1 argument.


LEVEL 4 answer (target: 17–18/20): See the model sentences built throughout this guide. The full Level 4 answer incorporates: P1 chain at all 5 stages → P2 evaluation of P1 with condition → P3 chain at all 5 stages → P4 evaluation of P3 → effective conclusion with all 5 elements.


PART 13: WHY THIS STRUCTURE AND NOT ANYTHING ELSE

Every component of this structure maps directly to a Level 4 descriptor requirement:

DESCRIPTOR REQUIREMENT    HOW SATISFIED
──────────────────────────────────────────────────────
"Thorough knowledge"       5-stage chains with
                           accurate business concepts
                           throughout — no gaps.

"Effective use of          Extract data woven into
context throughout"        stages 2–3 of each chain —
                           not mentioned separately.

"Coherent logical chain"   5-stage structure with
                           signal words at each
                           transition.

"Fully bilateral"          P1 + P2 (evaluate P1) +
                           P3 + P4 (evaluate P3) —
                           both sides argued AND
                           tested.

"Awareness of              Significance stated at
significance of            Stage 5 of each chain AND
competing arguments"       in P2/P4 evaluation
                           paragraphs.

"Effective conclusion      All 5 elements present:
with recommendation"       weighing + evidence +
                           condition + counter-
                           condition + specific
                           business recommendation.

Remove any component and a descriptor requirement goes unmet. Unmet requirement = examiner moves you down within Level 4 or back to Level 3. That is why this structure, in this order, is the best way to answer a 20-mark Evaluate question.

Veridian Legacy · progress saved in this browser · sign in to sync across devices

Up next

The Definitive Guide to the WBS12 4-Mark Questions

PART 1: THE 4-MARKER IS TWO DIFFERENT QUESTIONS

13 min