A dark stepped slate surface partly obscured by a sheet of frosted glass, suggesting a visible decline whose exact magnitude remains uncertain.

Reading Between the Lines: What New Zealand’s PIAAC Results Show — and What the Official Explanation Does Not Yet Establish

Working paper

Author note: I previously worked within Ako Aotearoa’s Manako adult literacy and numeracy team and retain professional relationships across the tertiary education system. This paper is independently written and based on publicly available sources. No current agency or organisation commissioned or endorsed it.

Provenance note: This paper was developed through an orchestrated human–AI research workflow: I directed the inquiry, challenged and tested the analysis, verified the evidence and retain full responsibility for the conclusions.

Evidence boundary: This analysis is limited to the publicly available evidence located at the time of writing. Relevant unpublished or unlocated Ministry or OECD analysis may exist and could alter the finding.

Executive summary

New Zealand PIAAC results show a large measured decline in adult literacy and numeracy compared with the previous comparable cycle: mean literacy fell by 21 points and mean numeracy by 15. The decline was unusually large internationally, broad across population groups, and more pronounced among lower-performing adults. Those facts deserve serious attention. [1][2]

They also require careful interpretation. New Zealand’s response rate was lower than in the previous cycle. The OECD assigned the country an Extended Non-Response Bias Analysis classification of Caution–Low rather than a clean pass. Interviewer recruitment, retention, monitoring and validation created genuine operational concerns. The OECD identified 301 New Zealand cases associated with unusual interviewer-linked response patterns and applied targeted technical treatment. These are not trivial qualifications. They justify caution about reading the raw score differences as an exact one-for-one measure of lost underlying capability. [3][4]

The Ministry of Education was therefore right to resist a simplistic story of national collapse. It also had access to direct operational information about difficult field conditions, lower participation and possible changes in respondent engagement. It had a legitimate responsibility to prevent a contested estimate from hardening immediately into an unquestioned public fact.

But the Ministry went further than caution. Its public report said that “a lot” of the measured change was likely caused by survey difficulty and that equivalent real declines across groups were “not plausible”. Those statements are not merely warnings about uncertainty. “A lot” is a claim about magnitude. “Not plausible” is a strong judgement about the range of credible explanations. The visible technical record does not quantify the share of the decline attributable to non-response, interviewer effects, engagement or other field conditions. [1][5][6]

The most defensible finding is therefore neither that the survey is fully reliable nor that the Ministry acted improperly. It is narrower and stronger:

The Ministry’s caution was substantively legitimate, its presentation was rhetorically asymmetric, and its strongest causal language was more confident than the visible evidence warranted.

I describe this as interpretive overreach, supported by asymmetric emphasis. The current record does not justify a finding of material minimisation, misrepresentation, deception or cover-up.

1. The question beneath the headline

Public evidence rarely travels from data to public understanding without being compressed. Technical reports contain measures, uncertainty, modelling choices, caveats and competing interpretations. Public summaries must select what comes first, what receives explanation, and what becomes memorable.

That selection is unavoidable. It is not automatically manipulation. Every institution has to translate complex evidence into language that ministers, practitioners, journalists and citizens can use.

The important question is not whether framing occurred. It always does. The question is whether the evidence supports the story being told.

The New Zealand PIAAC case is unusually useful because the underlying facts are not seriously disputed. The OECD and the Ministry agree that a large decline was measured. They agree that survey conditions create reasons for caution. They agree that population composition does not explain the decline. The disagreement lies in narrative centre and causal confidence.

The OECD country note foregrounds the result: New Zealand experienced particularly large declines in literacy and numeracy, with deterioration concentrated among lower-performing adults. The Ministry’s public interpretation gives greater weight to the possibility that post-COVID field conditions, participation difficulties and lower engagement inflated the measured change. [1][2]

Both statements can be partly true. The analytical task is to determine how far each one is supported.

2. What the survey measured

PIAAC is the OECD’s international assessment of adult skills. It measures literacy, numeracy and adaptive problem solving through a structured household survey and cognitive assessment. New Zealand participated in the earlier comparable cycle in 2014 and again in the 2023 cycle.

The headline trend is clear:

• Mean literacy declined from 281 to 260 points: a fall of 21 points.

• Mean numeracy declined from 271 to 256 points: a fall of 15 points.

• The declines appeared across most population groups.

• The lower end of the proficiency distribution deteriorated more sharply.

• The share of adults at low proficiency increased. [1][2]

These are estimates, not a national census. They carry sampling and measurement uncertainty. The literacy and numeracy scales are, however, formally linked across cycles. The OECD incorporates linking error into the uncertainty of trend comparisons. The linking-error terms are materially smaller than the observed changes and should not be mechanically subtracted from them. They widen uncertainty; they do not erase direction. [3][7]

The Ministry’s own decomposition also matters. Changes in the age, education, migration and language composition of the population do not explain the measured fall. In the literacy analysis, compositional change would have moved the expected score slightly upward overall. The decline therefore cannot be dismissed as the statistical consequence of New Zealand simply having a different adult population in 2023. [1]

The responsible starting point is consequently firm but bounded:

PIAAC measured a large deterioration in adult literacy and numeracy performance. The exact relationship between that measured deterioration and a change in underlying national capability remains uncertain.

That distinction should govern everything that follows.

3. Why caution is genuinely warranted

The case for caution is substantial. It should be presented in its strongest form before evaluating whether the public explanation went too far.

3.1 Response and representativeness

New Zealand’s overall response rate was approximately 48–49.3 per cent, lower than in the previous cycle. A low response rate does not automatically produce large bias; survey methodology has repeatedly shown that response rate and non-response bias are not the same thing. Bias depends on who did not respond, how respondents and non-respondents differ, and whether weighting variables adequately capture those differences.

The OECD conducted an Extended Non-Response Bias Analysis and classified New Zealand as Caution–Low. That classification means some non-response bias may be present. It is not a pass, and it should not be ignored. [3][4]

The survey design also reduced the effective sample size: for literacy, the published design effects imply an effective sample of roughly 988 rather than the full respondent count. That matters for precision, not direction. A smaller effective sample increases uncertainty around the national estimate; it does not, by itself, explain why the estimate moved downward. The more consequential question is bias: whether non-response, engagement, interviewer effects or other survey conditions systematically depressed the measured result. [4]

The sample also differed from population benchmarks in areas including ethnicity and employment status. Weighting and calibration can correct observed differences when suitable variables exist, but they cannot guarantee correction for unobserved differences, motivational differences or poorly measured characteristics.

3.2 Fieldwork and interviewer conditions

The fieldwork environment was difficult. Interviewer recruitment and retention were problematic. High employment made field staffing harder. Weather events disrupted data collection. Interviewer attrition was high, and the OECD’s New Zealand adjudication recorded Caution for monitoring and validation. [4]

The OECD also detected unusual response patterns clustered around particular interviewers. It identified 301 New Zealand cases requiring targeted handling. These cases were excluded from estimating the population model, although their cognitive responses still contributed to plausible-value estimation. The treatment indicates a real robustness concern without establishing that all 301 assessments were fabricated or invalid. [3][4]

A national agency with direct knowledge of the field operation could reasonably give these conditions significant interpretive weight.

3.3 Engagement and low-stakes performance

PIAAC is a low-stakes assessment for participants. Adults do not receive qualifications, employment benefits or personal results that materially affect their lives. Performance therefore reflects not only maximal cognitive capacity but also willingness to persist, concentrate and cooperate in the assessment context.

If 2023 respondents were less engaged than 2014 respondents, scores could fall without an equivalent loss in latent maximum ability. The Ministry’s public materials identify respondent engagement as a plausible concern.

But engagement is analytically complicated. Lower willingness or ability to mobilise literacy and numeracy under ordinary motivational conditions may itself be substantively relevant. A society does not use capability only when citizens are maximally motivated in a laboratory-like setting. The line between measurement contamination and changed typical performance is not self-evident. [8]

3.4 The broad pattern is genuinely surprising

A decline across many groups over less than a decade would be historically significant. The Ministry could reasonably consider a literal, equal capability decline across nearly every group unlikely. Broad simultaneous change increases the need to test common survey or contextual mechanisms.

That is the charitable core of the Ministry’s position: the survey probably did not measure pure skill loss, and the raw point differences should not be read literally.

This case explains why caution was necessary. It does not yet establish how much correction is warranted.

4. What the Ministry communicated

The Ministry published the headline scores. Its national report and Education Counts material placed methodological caution near the front. The pre-release briefing records concern that commentators and the OECD would foreground the declines while cautions would receive less attention. The planned communication approach was largely reactive, with ministers and selected adult-literacy sector figures pre-briefed. [1][5]

None of this is inherently improper. Communication planning, risk identification and stakeholder briefing are ordinary institutional practices.

The key issue lies in two high-leverage formulations.

First, the Ministry said that “a lot” of the measured change was likely due to difficulties conducting the survey and achieving comparable respondent engagement.

Second, it said that actual declines of this breadth across groups were “not plausible”.

The linked Foundation Education Update repeated the same interpretive position and referred readers to the release briefing for further detail. It did not supply a quantitative decomposition. [6]

These phrases changed the epistemic status of the public message. The Ministry moved from a well-supported statement—

survey limitations create uncertainty about the exact magnitude—

to a stronger conclusion—

survey limitations probably explain a large share of the measured decline.

The first statement is demonstrated. The second remains plausible but unquantified.

5. The strongest charitable case

A fair analysis should state the Ministry’s defence in language its strongest informed advocate could accept.

The Ministry had access to more operational detail about the New Zealand field operation than a reader of the country note. The public record documents lower response, interviewer instability, weather disruption, unusual interviewer-linked patterns, sample discrepancies and a decline that appeared implausibly broad. The OECD did not give New Zealand a clean non-response pass. Monitoring and validation received Caution. Residual bias can remain even when alternative weighting does not materially move published means.

The pre-release material also reflects a predictable communications problem. A 21-point literacy decline could easily become shorthand for a precise, proven collapse in national skill. Once embedded in public memory, that compressed story would be difficult to correct. The Ministry had a responsibility to communicate that the measured magnitude was uncertain.

On the strongest charitable reading, “not plausible” was not intended as a statistical theorem but as a professional judgement based on the total field evidence: a genuine decline may have occurred, but the survey probably overstated it. On that reading, “a lot” can be understood as a synthesis of operational evidence that was not reduced to a single published model.

This defence is coherent. It explains:

• why the Ministry’s framing differed from the OECD’s;

• why caution was prominent;

• why the Ministry resisted literal interpretation of the score changes;

• why communications planning began before publication;

• why the operational evidence could reasonably be given more weight than a public reader might give it.

The strongest defensible version of this position is:

The survey measured a large decline, but unusual response, fieldwork and engagement conditions mean the raw magnitude should not be treated as a one-for-one decline in underlying adult capability. A meaningful decline may be genuine, but the survey probably overstates its scale.

That formulation remains plausible. The problem is not that it can be imagined. The problem is that the public record does not show how the size of the overstatement was determined.

6. The strongest evidence-based challenge

The critical case begins by accepting legitimate caution.

It does not require us to believe that the raw scores are perfect. It asks whether the known evidence supports the Ministry’s stronger causal language.

Several findings constrain the answer.

First, the declines remain large relative to formal linking uncertainty. Cross-cycle linking does not dissolve the result. [3][7]

Second, population composition does not explain the decline. The Ministry’s own analysis indicates that compositional change would have raised expected literacy slightly rather than lowered it. [1]

Third, the OECD adjudication recorded Pass for sample weighting and Pass for assessment data. New Zealand’s alternative raking and external sample comparisons did not produce notable differences in mean proficiency. This does not prove the absence of all residual bias, but it weakens the proposition that known weighting problems explain a large share of the fall. [4]

Fourth, Caution–Low does not estimate direction or magnitude. It indicates that some non-response bias may exist. It cannot be converted into an implied downward correction of ten, fifteen or twenty points. [3][4]

Fifth, the 301 cases support caution but do not supply the missing magnitude. They were not 301 proven-invalid assessments removed from the national result. No public sensitivity analysis shows what New Zealand’s mean would have been if their cognitive responses had been excluded or modelled differently. [3][4]

Sixth, engagement cannot simply be labelled error. Reduced effort may arise from survey conditions, but it may also reflect changed typical performance. Treating engagement as contamination alone assumes the conclusion the analysis needs to establish. [8]

Finally, the broad subgroup pattern is ambiguous. It may indicate a common engagement effect. It may also indicate genuine deterioration concentrated at the lower end. The Ministry gave greater explanatory weight to the first interpretation without publicly showing why it was more likely.

The critical conclusion is therefore narrow:

The Ministry accurately identified genuine survey limitations, but communicated a stronger causal conclusion than the visible technical evidence supports. The exact magnitude of decline is uncertain; the claim that much of it was survey-generated remains unquantified.

7. The narrative finding

The two cases do not carry equal evidential burdens.

The charitable case requires an unobserved proposition: the combined effects of non-response, interviewer conditions and engagement were large enough to explain a substantial share of the decline.

The critical case requires a narrower proposition: in the absence of a visible magnitude estimate, the stronger causal wording has not been demonstrated.

The critical case therefore has the lower evidential burden.

This does not make the Ministry’s underlying concern wrong. It means the concern was communicated at a higher confidence level than the public analysis could support.

The conclusion is:

Interpretive overreach, supported by asymmetric emphasis.

Asymmetric emphasis applies because the Ministry gave greater rhetorical and structural weight to the survey-effect explanation than to the possibility of a materially genuine decline.

Interpretive overreach applies because a plausible but insufficiently tested explanation was presented as likely and substantial.

The evidence does not establish material minimisation. The Ministry published the declines, acknowledged that some genuine change may exist, and raised real methodological concerns. The public effect of the framing has not been demonstrated strongly enough.

The evidence does not establish misrepresentation. The headline results were not falsified, and the survey limitations were real.

The evidence does not establish deception or cover-up. It does not show concealment, manipulation, knowledge of falsity or improper intent.

The compressed finding is:

Substantively legitimate caution; rhetorically asymmetric; causally overconfident.

8. What can responsibly be said

The following public claim is supported:

New Zealand’s PIAAC results measured a large decline. Survey-quality concerns justify caution about treating the raw score difference as an exact measure of capability loss. However, the Ministry’s claim that “a lot” of the decline was survey-generated carries more causal confidence than the visible technical record quantifies.

A stronger but still defensible formulation is:

The Ministry’s presentation moved from legitimate methodological caution into interpretive overreach. It treated a plausible survey-effect explanation as more established than the visible evidence supports.

The following claims are not supported:

• intentional deception or dishonesty;

• suppression of the measured decline;

• deliberate misleading of the public;

• that the decline is wholly genuine;

• that the survey limitations are trivial;

• that no additional analysis exists;

• that intent can be inferred from communications planning.

This boundary is not rhetorical timidity. It is what makes the central criticism difficult to dismiss.

9. Evidence that could change the finding

The current finding should remain revisable.

It would weaken if a credible New Zealand-specific sensitivity analysis showed that alternative handling of non-response, interviewer-linked cases or engagement raised national scores by a large amount.

It would also weaken if release-era OECD correspondence explicitly endorsed the Ministry’s “a lot” formulation or if a pre-release Ministry decomposition quantified a substantial survey effect.

The finding could strengthen if records showed that no quantitative analysis existed, that the interpretation relied primarily on plausibility judgement, or that contrary internal evidence was known and disregarded.

The most discriminating remaining evidence is therefore:

• a counterfactual estimate for the 301 interviewer-linked cases;

• a decomposition of non-response, calibration and engagement effects;

• release-era analytical notes supporting “a lot”;

• OECD–Ministry correspondence about interpretation;

• release-day Q&A and reactive communications;

• external expert advice provided before publication.

The present analysis does not depend on assuming what those records contain.

10. Reading between the lines

The wider lesson is not that official reports are untrustworthy. Nor is it that technical caution is a disguise.

The lesson is that a public narrative can become more confident than its evidential base without implying bad faith.

A measured result may be real and uncertain. An institutional explanation may be plausible and overconfident. Communication may be responsible in purpose and asymmetric in effect.

Reading public evidence well requires holding these statements together.

It requires separating:

• what was observed;

• how it was measured;

• what uncertainty remains;

• what explanation is inferred;

• how confidently that explanation is communicated;

• what the public is likely to remember.

That is the practical meaning of interpretive sovereignty in this case. It is not reflexive distrust of institutions. It is the capacity to respect evidence without surrendering judgement to the first authoritative narrative placed around it.

Conclusion

New Zealand’s adult skills results should not be read as a precise, uncontested measurement of national capability loss. The survey carried genuine quality concerns, and the Ministry was right to foreground them.

But caution has a boundary.

The public record supports uncertainty about the exact size of the decline. It does not quantify the conclusion that a large share was survey-generated. The Ministry’s strongest wording therefore travelled further than its visible evidence.

This is a finding about evidential proportionality, not motive. It is a finding of interpretive overreach.

The distinction matters. It allows us to take the measured decline seriously without pretending the estimate is exact, and to challenge the official explanation without attributing motives the evidence cannot establish.

Invitation to correction

This paper is not a closed verdict. The Ministry of Education, OECD and other informed readers are invited to identify any overlooked quantitative analysis, sensitivity estimate or technical correspondence that supports—or materially changes—the interpretation set out here. Any such evidence should be incorporated and the finding revised without embarrassment.

Version 1.0. Material new evidence will be incorporated and the analysis revised where warranted.

Sources and notes

[1] Ministry of Education — Skills in New Zealand: Survey of Adult Skills 2023 (PIAAC).

[2] OECD — Survey of Adult Skills 2023: New Zealand country note.

[3] OECD — Survey of Adult Skills 2023: Reader’s Companion and methodology chapter.

[4] OECD — Survey of Adult Skills 2023 Technical Report and New Zealand adjudication annex.

[5] Ministry of Education — METIS 1338898, release briefing for the New Zealand report.

[6] Ministry of Education — METIS 1338879, Foundation Education Update.

[7] OECD — Education at a Glance 2025 PIAAC trend methodology note.

[8] OECD — Interviewers, test-taking conditions and the quality of the PIAAC assessment.


Discover more from THISISGRAEME

Subscribe to get the latest posts sent to your email.


Comments

Kia ora! Hey, I'd love to know what you think.

Discover more from THISISGRAEME

Subscribe now to keep reading and get access to the full archive.

Continue reading