HomeSampling MethodsJudgment Sampling: Bias and Agreement Calculator

Judgment Sampling: Bias and Agreement Calculator

Judgment Sampling: Bias and Agreement Calculator

Judgment Sampling Calculator

In judgment sampling the researcher is the instrument, so the sample is only as good as the judgement behind it. This tool tests that judgement directly: it scores your exposure to eight cognitive biases known to distort expert selection, then measures whether a second or third judge would have chosen the same cases, using percent agreement, Cohen kappa, Fleiss kappa, PABAK and Jaccard overlap.

8 cognitive bias checks Cohen and Fleiss kappa PABAK for rare selection Selection overlap matrix Free, no sign up

0Quick Answer

Judgment sampling selects cases on expert judgement rather than chance, so its weakness is the judge, not the arithmetic. Test it by having a second expert select independently from the same pool and computing Cohen kappa = (Pₒ − Pₑ) / (1 − Pₑ). Example: two auditors agreeing on 92 of 100 files, with 78 percent agreement expected by chance, gives kappa = 0.636, substantial agreement.

Key Takeaways

  • Judgment sampling is another name for purposive or judgmental sampling, but the practical question it raises is different: how reliable is the judge?
  • A judgement nobody else would reproduce is not a sampling strategy, it is a preference. Inter-rater agreement is how you tell the two apart.
  • Cohen kappa corrects agreement for chance; use Fleiss kappa for three or more judges.
  • When selection is rare, kappa collapses even at high agreement. Report PABAK alongside it or you will misread your own data.
  • Eight cognitive biases distort expert selection predictably. Naming which ones apply is more useful than claiming none do.

1What Is Judgment Sampling?

On terminology. Judgment sampling, judgmental sampling, purposive sampling and selective sampling name the same family of non probability designs. Methodology journals prefer purposive; audit, inspection and quality control practice prefer judgment. If you want the strategy chooser, coverage matrix and saturation tools, use our purposive sampling tool. This page handles the question that the audit and expert selection literature cares about and the qualitative literature largely does not: how good is the judgement, and would anyone else have made the same one?

Judgment sampling selects units because an experienced person believes those units are the right ones to examine. An auditor picks the ledger entries that look irregular. An inspector picks the machines most likely to be failing. A supervisor picks the three sites that best represent the region. In each case a human expert, not a random mechanism, decides who is in the sample.

That makes the researcher the measuring instrument, and instruments have to be calibrated. In quantitative sampling you validate the procedure; in judgment sampling there is no procedure to validate, only a person. So the two questions that matter are whether the judgement is reliable, meaning another competent expert would reach a similar selection, and whether it is systematically distorted, meaning known cognitive biases pushed it in a predictable direction. Neither question is answered by the sample size.

Both are answerable. Reliability is measured by having two or more judges select independently from the same candidate pool and computing chance corrected agreement. Bias exposure is assessed by working through the specific heuristics that are known to distort expert selection: anchoring on the first cases seen, availability of memorable cases, halo effects from reputation, confirmation of an existing hypothesis, and simple accessibility. This tool does both.

Two judges selecting independently from the same candidate pool Would a second expert have chosen the same cases? Candidate pool of 20 units AAA BB A only: 3 cases B only: 2 cases Both agreed: 2 cases Neither: 13 cases Percent agreement = (2 + 13) / 20 = 75 percent, but much of that is agreeing on what to exclude. Cohen kappa corrects for that chance agreement. Here kappa = 0.29, only fair. Two experts using the same brief disagreed on 5 of the 7 cases either of them picked.

Figure 1.1 High percent agreement can hide poor reliability, because most of it comes from agreeing on exclusions.

2Setup: The Pool, the Judges and the Bias Check

Before you start. Judgment sampling is a non probability design. This tool computes no margin of error, because none exists. It measures two things that can be measured: how reliable the judgement was, and which cognitive biases it was exposed to.
Every unit the judges could have chosen from. Agreement is measured against this pool.

Cognitive bias exposure scorecard

Eight heuristics known to distort expert selection. Each answer scores 0 for low exposure, 1 for moderate and 2 for high. A higher total means the judgement was more likely to be pushed in a predictable direction.

Judges: one card per judge

Enter each judge and the case IDs they selected, comma separated. IDs can be anything consistent: numbers, file references, plot codes. The tool compares selections pairwise across the pool. Judge names are editable.

3Results

Nothing has been assessed yet. Set the pool size, score the biases and list your judges above, then click Score Bias and Measure Agreement.

4Interpretation of Results in Detail

Run the tool above. This section then fills in with your own agreement and bias numbers.

How to read each number

Percent agreement. The share of the candidate pool on which two judges made the same decision, counting both joint selections and joint rejections. It is intuitive and it is almost always misleadingly high, because when you select 8 cases from 200 the two judges agree on 192 exclusions before they consider a single inclusion. Never report percent agreement on its own for a selective task.

Expected agreement. How much of that agreement two judges would have reached by chance alone, given how often each of them selects. This is the quantity percent agreement ignores and kappa subtracts. When both judges select rarely, expected agreement is very high, which is precisely why raw agreement looks so good.

Cohen kappa. Agreement above chance, scaled so that 0 means no better than chance and 1 means perfect. The Landis and Koch bands are conventional: below 0 poor, 0 to 0.20 slight, 0.21 to 0.40 fair, 0.41 to 0.60 moderate, 0.61 to 0.80 substantial, above 0.80 almost perfect. For a selection task, moderate is a reasonable target and substantial is strong. A kappa near zero with high percent agreement means the two judges agreed on almost nothing that mattered.

The kappa paradox and PABAK. Kappa has a well documented failure mode: when the thing being judged is rare, kappa can be low or even negative while agreement is genuinely high. This is not a quirk to be explained away, it is a reason to report a second statistic. PABAK, the prevalence adjusted bias adjusted kappa, equals 2 times percent agreement minus 1, and it removes the dependence on how often selection occurs. When kappa and PABAK diverge sharply, the divergence itself is your finding: it says agreement is real but selection is rare, and you should report both figures with that explanation rather than quoting whichever number flatters you.

Fleiss kappa. When three or more judges each rate every case, Fleiss kappa gives one overall chance corrected agreement figure across the whole panel rather than a set of pairwise numbers. It answers a subtly different question: not whether two specific people agree, but whether the panel as a whole is applying a shared standard. Report Fleiss for the panel and the pairwise table underneath it, because a good overall figure can hide one judge who disagrees with everyone.

Jaccard overlap. The share of cases that either judge selected on which both agreed. Unlike percent agreement it ignores the exclusions entirely, so it answers the question a practitioner actually asks: of everything anyone flagged, how much did we both flag? It is often the most honest single number in this table, and it is always lower than percent agreement.

The selection rate per judge. How many of the pool each judge picked. Very different rates between judges are informative in themselves: one person applying a much looser threshold is not the same problem as two people disagreeing about which cases matter, and kappa alone cannot distinguish them. Compare the rates before you interpret the kappa.

Times each case was selected. The case by case map shows which units every judge flagged, which only one flagged, and which none did. Unanimous cases are your defensible core. Cases picked by exactly one judge are where the judgement was doing the work, and those are the ones a reviewer will ask about, so have a reason ready for each.

The bias exposure index. Eight heuristics, scored 0 to 2, rescaled to 0 to 100. This is not a statistic and it corrects nothing. It is a structured way of writing down what you already suspect about your own selection process, so that the limitations section names specific mechanisms instead of gesturing at subjectivity. Below 30 means the process had reasonable safeguards; above 60 means the selection was probably driven as much by the heuristics as by the criteria.

Independence of judges. This modifies everything else. If the judges discussed the cases first, or one saw the other choices, then high agreement measures conformity rather than reliability. The tool flags this because a kappa of 0.9 from two people in the same room is worth less than a kappa of 0.5 from two people working blind.

What none of this can tell you. Two judges can agree perfectly and both be wrong in the same direction. Agreement measures reliability, not validity. If both experts share the same training, the same institutional assumptions and the same blind spot, your kappa will be excellent and your sample will still miss what matters. High agreement plus a high bias score is the combination to worry about most, because it looks reassuring and is not.

5How to Write Your Results in Research

Use the templates below. Each one names the judge, the basis, the agreement statistics and the bias exposure, which is what reviewers and audit reviewers look for in a judgmental selection.

Run the tool to auto-fill these templates with your own values.

Rules that make a judgement defensible in writing

  1. Say who judged and what qualified them. A judgment sample rests on the judge, so anonymising the expertise removes the only warrant the design has.
  2. State the judgement basis as criteria, not as instinct. "Entries irregular on amount, timing or vendor" is defensible; "entries that looked wrong" is not.
  3. Report the candidate pool size. Selecting 8 of 200 and 8 of 12 are entirely different acts, and without N the reader cannot tell which you did.
  4. Give the selection rate. It sets the context for every agreement figure that follows.
  5. Report kappa and PABAK together when selection is rare, and explain the divergence rather than choosing the flattering number.
  6. Confirm the judges were independent, or say plainly that they were not. Agreement between people who conferred is conformity.
  7. Name the biases you were exposed to. Specific mechanisms read as competence; a blanket claim of objectivity reads as naivety.
  8. List the cases only one judge selected and give the reason each was kept or dropped. That list is where the judgement is visible.
  9. Do not report a margin of error. There is no sampling distribution behind a judgement.
  10. Distinguish reliability from validity explicitly. Say that agreement shows consistency, not correctness.

Common wording mistakes and the fix

Wrong wordingWhy it failsCorrect wording
"A judgment sample of 20 items was selected."No judge, no basis, no pool size."Two auditors independently selected 20 of 200 ledger entries against three stated irregularity criteria."
"The two reviewers agreed on 94 percent of cases."Almost all of that is agreement on exclusions."Raw agreement was 94 percent, but Cohen kappa was 0.41 and Jaccard overlap 0.34, indicating only moderate reliability."
"Kappa was low, so we report percent agreement instead."Choosing the flattering statistic."Kappa was 0.28 against 94 percent agreement; because selection was rare (4 percent) we also report PABAK = 0.88 and interpret both."
"The expert selection was objective."Unfalsifiable and undermines credibility."Selection was expert dependent; exposure to availability and accessibility bias is reported as a limitation."
"The reviewers discussed and agreed the final sample."Describes consensus, not measured reliability."Reviewers selected independently, then reconciled; agreement is reported from the independent stage."

6Formulas Used

The agreement matrix for two judges
a = both selected  b = J1 only  c = J2 only  d = neither  (a+b+c+d = N)
NSize of the candidate pool both judges worked from
aCases both judges selected; your defensible core
b, cCases only one judge selected; where the judgement did the work
dCases neither selected; usually the largest cell and the source of inflated agreement
Observed and expected agreement
Pₒ = (a + d) ÷ N   Pₑ = [(a+b)(a+c) + (c+d)(b+d)] ÷ N²
PₒObserved agreement, the plain percentage of matching decisions
PₑAgreement expected by chance, given how often each judge selects
WarningWhen selection is rare, Pₑ approaches Pₒ and raw agreement becomes meaningless
Cohen kappa
κ = (Pₒ − Pₑ) ÷ (1 − Pₑ)
κAgreement above chance; 0 is chance level and 1 is perfect
BandsLandis and Koch: ≤0 poor, .01-.20 slight, .21-.40 fair, .41-.60 moderate, .61-.80 substantial, >.80 almost perfect
NegativeWorse than chance; the judges are systematically picking different cases
PABAK, prevalence adjusted bias adjusted kappa
PABAK = 2Pₒ − 1
PABAKKappa recomputed as if selection were equally likely and both judges equally strict
UseReport alongside kappa whenever the selection rate is far from 50 percent
DivergenceA large gap between kappa and PABAK is a finding: agreement is real but selection is rare
Prevalence and bias indices
PI = |a − d| ÷ N   BI = |b − c| ÷ N
PIPrevalence index; high values mean selection is rare or near universal, which depresses kappa
BIBias index; high values mean one judge selects far more freely than the other
ReadThese two explain most cases where kappa looks wrong relative to raw agreement
Jaccard overlap of the selections
J = a ÷ (a + b + c)
JOf every case either judge selected, the share both selected
WhyIgnores the joint exclusions entirely, so it cannot be inflated by a large pool
NoteAlways lower than percent agreement, and usually the most honest single figure
Fleiss kappa for three or more judges
κ₣ = (P̄ − P̄ₑ) ÷ (1 − P̄ₑ)
PᵢAgreement on case i = (Σ nᵢⱼ² − m) ÷ (m(m − 1))
mNumber of judges rating every case
P̄ₑΣ pⱼ² where pⱼ is the overall share of decisions in category j
UseOne panel level figure; still report the pairwise table beneath it
Selection rate and unanimity
rate = 100 × (cases selected) ÷ N   unanimity = 100 × (cases all judges picked) ÷ (cases any judge picked)
rateHow selective each judge was; compare rates before interpreting kappa
unanimityShare of the flagged pool that every judge agreed on; your defensible core
SingletonsCases exactly one judge chose; each needs a stated reason for keeping or dropping
Bias exposure index
B = 100 × (Σ scoreᵢ) ÷ (2 × 8)
scoreᵢ0 low, 1 moderate, 2 high exposure on each of eight heuristics
BIndex 0 to 100; below 30 reasonable safeguards, above 60 heuristic driven
NoteA structured self assessment, not a statistic; it corrects nothing
Reliability against validity
high κ & high B → shared blind spot
ReliabilityJudges agree with each other; measured by kappa
ValidityJudges are right; not measured by anything on this page
TrapTwo judges with the same training can agree perfectly and be wrong together

7How to Use This Tool

  1. Type the engagement name and, in one line, the basis on which cases were judged.
  2. Enter the candidate pool size N, meaning every unit the judges could have chosen from.
  3. Record the expertise level of the judges and whether they worked independently, because both modify how the agreement figures should be read.
  4. Work through the eight cognitive bias questions honestly; scoring yourself low here only misleads you.
  5. Add one card per judge and paste the case IDs that judge selected, comma separated.
  6. Use consistent IDs across judges. The tool matches on the ID text, ignoring case and spacing.
  7. With two judges you get Cohen kappa and PABAK; with three or more you also get Fleiss kappa for the whole panel.
  8. Or switch to the upload tab and click the columns that should each become a cluster, one judge per column.
  9. Click Score Bias and Measure Agreement, then read the kappa card and the singleton cases first.
  10. Copy the methods paragraph into your report and download the judgement report for your file or working papers.

8Detailed Reference Tables

Table 8.1 The eight cognitive biases in expert selection

HeuristicHow it distorts selectionLow exposure looks like
AnchoringThe first few cases examined set the template for everything afterCriteria written before any case was seen
AvailabilityMemorable or recent cases are over selectedSelection driven by records, not recall
ConfirmationCases that support an existing hypothesis are preferredDisconfirming cases deliberately sought
Halo and reputationKnown names or well regarded units are judged differentlyIdentifying details masked during selection
AccessibilityEasy to reach cases are chosen over relevant onesSelection made before access was considered
RepresentativenessCases that look typical are assumed to be typicalTypicality checked against data, not intuition
Salience of extremesDramatic cases crowd out ordinary onesExplicit quota for unremarkable cases
Stakeholder pressureA manager or gatekeeper steers the selectionSelection made and recorded before consultation

Table 8.2 Interpreting kappa, the Landis and Koch bands

KappaInterpretationVerdict for a selection task
Below 0.00Poor, worse than chanceJudges are picking different things
0.00 to 0.20SlightNot a reproducible judgement
0.21 to 0.40FairWeak; tighten the criteria
0.41 to 0.60ModerateAcceptable, report openly
0.61 to 0.80SubstantialStrong for expert selection
0.81 to 1.00Almost perfectCheck the judges were truly independent

Table 8.3 Why kappa can look wrong: the prevalence paradox

Scenarioa / b / c / dPercent agreementCohen kappaPABAKWhat to report
Balanced selection40 / 10 / 10 / 4080.0%0.6000.600Kappa alone is fine
Rare selection, high agreement4 / 2 / 2 / 9296.0%0.6450.920Both, with the gap explained
Very rare selection2 / 3 / 3 / 9294.0%0.3680.880Both; kappa understates
One judge much stricter5 / 20 / 1 / 7479.0%0.2500.580Both, plus the bias index
Near universal selection92 / 2 / 2 / 496.0%0.6450.920Both, with the gap explained

Table 8.4 Judgment sampling against the alternatives

MethodWho selectsReliability testable?Margin of error?Best use
Judgment samplingAn expertYes, with a second judgeNoTargeting known risk or interest
Purposive samplingThe researcher, by strategyRarely testedNoQualitative depth and variation
Haphazard or convenienceNobody, effectivelyNoNoAvoid where possible
Quota samplingFieldworker within quotasPartlyNoControlling composition cheaply
Monetary unit samplingAn algorithm, weighted by valueNot neededYesAudit of monetary populations
Simple randomChanceNot neededYesUnbiased population estimates
Stratified randomChance within strataNot neededYesPrecision with known subgroups

9Example Results (8 Worked Cards)

Example 1. Two auditors, balanced selection from 100 files

N = 100, a=40 b=10 c=10 d=40

confidence interval plot: Two auditors, balanced selection from 100 filesPoint estimate with interval central value lowerupper rangespread

Figure 9.1 Agreement statistics with their spread from chance level to perfect.

Percent agreementCohen kappaPABAKJaccardVerdict
80.0%0.6000.6000.67Substantial

What it means: With selection near 50 percent, kappa and PABAK coincide at 0.600 and either can be reported alone. This is the only situation in which raw agreement, kappa and PABAK tell the same story.

How to write it: Two auditors independently selected from 100 files against three stated criteria. Raw agreement was 80.0 percent, Cohen kappa 0.600 (substantial), Jaccard overlap 0.67.

Example 2. Rare selection: the kappa paradox in action

N = 100, a=4 b=2 c=2 d=92

vertical bar chart: Rare selection: the kappa paradox in actionValue by case, vertical bars mean cases in selection order

Figure 9.2 Times each case was selected, drawn as vertical bars against the panel average.

Percent agreementCohen kappaPABAKJaccardVerdict
96.0%0.6450.9200.50Substantial

What it means: Agreement of 96 percent looks excellent but 92 of those cases are joint exclusions. Kappa of 0.645 is the honest figure; PABAK of 0.920 shows how much of the gap is caused by rare selection rather than poor judgement.

How to write it: Raw agreement was 96.0 percent but, because only 6 percent of the pool was selected, we report Cohen kappa 0.645 alongside PABAK 0.920 and Jaccard overlap 0.50.

Example 3. Very rare selection depresses kappa to fair

N = 100, a=2 b=3 c=3 d=92

horizontal bar chart: Very rare selection depresses kappa to fairGroup comparison, horizontal barsA25.8B37.3C16.5D22.4E29.0F33.2value

Figure 9.3 Selection counts per judge compared side by side as horizontal bars.

Percent agreementCohen kappaPABAKJaccardVerdict
94.0%0.3680.8800.25Fair

What it means: Only two cases were picked by both judges out of eight either flagged, so Jaccard is 0.25. Kappa of 0.368 understates agreement because of the prevalence index, but Jaccard confirms the reliability genuinely is weak here.

How to write it: Cohen kappa was 0.368 against 94.0 percent raw agreement; the prevalence index of 0.90 explains the divergence, and Jaccard overlap of 0.25 confirms limited reliability.

Example 4. One judge far stricter than the other

N = 100, a=5 b=20 c=1 d=74

dot strip plot: One judge far stricter than the otherOne dot per case, strip plot centre each dot is one selected case

Figure 9.4 Every case in the pool shown as one dot, positioned by how many judges picked it.

Percent agreementCohen kappaPABAKJaccardVerdict
79.0%0.2500.5800.19Fair

What it means: Judge 1 selected 25 cases and Judge 2 only 6. The bias index of 0.19 flags this directly: the problem is not that the judges disagree about which cases matter, it is that they are applying different thresholds.

How to write it: Selection rates differed markedly (25 versus 6 percent), giving a bias index of 0.19; kappa of 0.250 reflects threshold disagreement rather than disagreement about individual cases.

Example 5. Four judge panel, Fleiss kappa

N = 60, m = 4 judges, 17 cases flagged

line trend plot: Four judge panel, Fleiss kappaTrend across cases, line plot case number

Figure 9.5 Agreement plotted across judge pairs to reveal an outlying judge.

Percent agreementCohen kappaPABAKJaccardVerdict
93.3%0.6850.8670.45Substantial

What it means: Fleiss kappa of 0.685 across the whole panel is substantial, and the six pairwise values run from 0.555 to 0.777, so no single judge sits far from the rest. Always check that spread, because a good panel figure can hide one outlying judge.

How to write it: Four assessors rated all 60 candidates; Fleiss kappa was 0.685. Pairwise Cohen kappa ranged from 0.555 to 0.777, indicating a consistently applied shared standard.

Example 6. Judges conferred: agreement measures conformity

N = 80, judged together, a=12 b=1 c=1 d=66

histogram: Judges conferred: agreement measures conformityDistribution across cases, histogram bins count

Figure 9.6 Distribution of cases by the number of judges who selected them.

Percent agreementCohen kappaPABAKJaccardVerdict
97.5%0.9080.9500.86Almost perfect

What it means: A kappa of 0.908 would be excellent from independent judges. These two discussed the cases first, so the figure measures how readily one deferred to the other. The tool flags this and the number should not be reported as reliability.

How to write it: Because assessors reviewed the cases jointly, the observed kappa of 0.908 reflects consensus rather than independent reliability, and no inter-rater reliability claim is made.

Example 7. Negative kappa: judges picking different things

N = 50, a=1 b=8 c=9 d=32

donut share chart: Negative kappa: judges picking different thingsShare of cases in each categoryshare 42 percent of casesshare 27 percent of casesshare 19 percent of casesshare 12 percent of cases100%

Figure 9.7 Share of the flagged pool that was unanimous, split or singleton.

Percent agreementCohen kappaPABAKJaccardVerdict
66.0%-0.1040.3200.06Poor

What it means: A negative kappa means the two judges agreed less than chance would predict. Of the 18 cases either flagged, they concurred on one. The criteria are not functioning as criteria and the selection cannot be defended as a shared standard.

How to write it: Cohen kappa was negative (-0.104) with Jaccard overlap of 0.06, indicating the two assessors were not applying a shared standard; the criteria were rewritten and selection repeated.

Example 8. Single judge, bias check only

N = 200, one expert, 15 selected

box and whisker plot: Single judge, bias check onlySpread by group, box and whiskergroup 1group 2group 3group 4value

Figure 9.8 Spread of selection rates across the judge panel.

Percent agreementCohen kappaPABAKJaccardVerdict
n/an/an/an/aNot testable

What it means: With one judge there is no agreement to measure, so the only available evidence is the bias scorecard, which came out at 56 out of 100. That is the common situation and the weakest one; half a day of a colleague re-selecting would change what this study can claim.

How to write it: Selection was made by a single expert, so inter-rater reliability could not be assessed. Bias exposure was scored at 56 of 100, with availability and accessibility bias the dominant concerns.

10How to Collect Raw Data in the Field

10a. The 18 point protocol for a defensible judgement

Plan

  1. Write the selection criteria before you look at a single case. This is the only defence against anchoring, and it costs ten minutes.
  2. Define the candidate pool explicitly and count it. Without N there is no denominator, no selection rate and no agreement statistic.
  3. Recruit a second judge before you start, not afterwards. A retrospective second opinion on cases already chosen measures nothing.
  4. Decide how disagreements will be resolved in advance: unanimous only, either judge, or reconciliation. Deciding afterwards invites the choice that flatters the result.
  5. Mask identifying details where you can. Removing names, units and known reputations is the cheapest available defence against halo bias.
Independent parallel judgingIndependent parallel judging, then compareCandidate pool, NJudge A selectsJudge B selectsno contactCompare only after both have finished. Any contact before that turns reliability into conformity.

Figure 10.1 The two judges must not see each other's selections until both are complete.

Kit

  1. Give each judge the identical pool list in the same order, or in a deliberately randomised order if order effects are a concern.
  2. Give each judge the written criteria and nothing else. Verbal briefing drifts between judges.
  3. Use a simple in or out recording sheet with a reason column, so the basis of each decision is captured while it is fresh.
  4. Keep the two sheets physically separate until both are complete.

Judge

  1. Have each judge decide on every case in the pool, not just mark the ones they like. Kappa needs a decision on all N.
  2. Record a one line reason for every selection. These reasons are what turn a judgement into a criterion applied.
  3. Do not let either judge revise after seeing the other sheet. Revisions belong in a separate reconciliation stage, recorded separately.
Decide on every case, not just the interesting onesCase IDIn / OutReasonCriterion metF-0142INAmount 4x the vendor medianC1 amountF-0143OUTRoutine, within all thresholdsnoneEvery case in the pool gets a row, including the ones you reject. Kappa needs a decision on all N.Marking only the interesting cases makes agreement impossible to compute.

Figure 10.2 A decision on every case, with a reason, is what makes the agreement statistic computable at all.

Compare

  1. Compute agreement before reconciling. Once the judges discuss, the independent figure is gone and cannot be recovered.
  2. List the singleton cases, those exactly one judge picked, and record the reason each was ultimately kept or dropped.
  3. Compare the two selection rates before interpreting kappa, because a threshold difference and a case level disagreement look identical in kappa but need different fixes.
Unanimous, split and singleton casesWhere the judgement actually did the workUnanimous2Defensible coreSingleton5Needs a stated reason eachNeither13Inflates raw agreementReport all three counts. The middle column is the one a reviewer will ask about, case by case.

Figure 10.3 Singleton cases are where the judgement is visible, so each needs a recorded reason.

Reliability versus validity versus consensusReliabilityJudges agree witheach otherMeasured by kappaValidityJudges are actuallyrightMeasured by nothing hereConsensusJudges talked andconvergedNot evidence of either

Figure 10.4 Three things that are easily confused; only the first is measured on this page.

Check

  1. If kappa comes out below 0.40, rewrite the criteria rather than the sample. Low agreement means the criteria are not doing the work; a bigger sample will not help.
  2. If kappa comes out above 0.85, verify the judges really were independent. Very high agreement on a subjective task usually means contamination.
  3. File both original sheets, the agreement calculation and the reconciliation record. In judgment sampling the audit trail of who decided what, and when, is the whole evidence base.
Action by kappa levelWhat the kappa figure tells you to do nextBelow 0.40 → rewrite the criteria; the judgement is not reproducible0.40 to 0.60 → acceptable; report openly and justify the singleton cases0.61 to 0.85 → strong; report kappa, PABAK and the selection ratesAbove 0.85 → verify independence before celebrating

Figure 10.5 The kappa value points to a specific next action, not just a label.

10b. Recording sheet column specification

ColumnFormatExampleWhy it matters
EngagementText, header onceQ3 expense ledger reviewLinks the sheet to the judgement basis.
Judgement basisText, header onceIrregular on amount, timing or vendorThe criteria as written before any case was seen.
JudgeName or initials, header onceJudge A, RPThe judge is the instrument, so they must be identified.
Pool size NInteger, header once200The denominator for every rate and agreement figure.
Case IDConsistent across judgesF-0142Agreement is matched on this text, so the format must not drift.
DecisionIN / OUTINEvery case needs one; marking only inclusions makes kappa impossible.
ReasonOne lineAmount 4x the vendor medianTurns a judgement into a criterion applied.
Criterion metC1 / C2 / C3 / noneC1 amountShows which written rule drove the decision.
ConfidenceHigh / Medium / LowMediumLow confidence selections are where disagreement concentrates.
Time decidedHH:MM10:42Reveals fatigue and anchoring effects across a long session.
Independent?Yes / No, header onceYesDetermines whether agreement means reliability or conformity.
NotesFree text, shortBorderline against C2Explains anything the decision column cannot.

Completeness rule: every case in the pool needs a row from every judge. A judge who records only their selections cannot be included in any agreement calculation. Independence rule: sheets stay separate until both are finished; the agreement figure is computed from the independent stage only, even if you reconcile afterwards.

10c. Filled worked recording sheet

#Case IDJudge AA reasonJudge BB reasonOutcome
1F-0142INAmount 4x vendor medianINAmount outlierUnanimous
2F-0143OUTWithin all thresholdsOUTRoutineNeither
3F-0147INPosted after period closeOUTTiming explained by accrualSingleton A
4F-0151OUTKnown recurring vendorINVendor not on approved listSingleton B
5F-0158INRound sum, no supporting docINNo documentation attachedUnanimous
6F-0163OUTRoutineOUTRoutineNeither
7F-0169INSplit just under approval limitOUTTwo genuine invoicesSingleton A
8F-0174OUTWithin thresholdsOUTRoutineNeither
9F-0180INNew vendor, first payment largeINNew vendor riskUnanimous
10F-0186OUTRoutineINManual journal, no approverSingleton B

What this sheet shows:

  • Every case has a decision from both judges, which is what makes the agreement calculation possible.
  • Rows 1, 5 and 9 are unanimous and form the defensible core of the sample.
  • Rows 3, 4, 7 and 10 are singletons, and the paired reasons show the judges applied different criteria rather than simply disagreeing.
  • Row 3 versus row 4 is instructive: A weights timing, B weights vendor approval. That is a criteria problem, not a competence problem.
  • Recording both reasons side by side is what lets you diagnose the disagreement rather than merely counting it.

10d. Blank print ready recording sheet

Download the blank sheet as a .csv file, open it in Excel, Google Sheets or LibreOffice, then print one copy per judge. The Decision and Reason columns for every case in the pool are what make the agreement statistics computable, so they are pre-labelled here even though most templates only capture the selections.

11Which Sampling Method Should You Use?

Decision tree. Answer these in order and stop at the first yes.

  1. Do you need to project a total or a rate to the whole population with a stated precision? Use a probability design; judgment sampling cannot support projection.
  2. Are you auditing a monetary population and need both targeting and projection? Use monetary unit sampling, which weights by value while remaining a probability method.
  3. Do you want to examine the cases most likely to contain the problem, and can you accept that findings do not project? Judgment sampling is the right tool; measure the judgement with this page.
  4. Do you have two or more experts available? Use judgment sampling with independent parallel judging so reliability is measurable.
  5. Are you doing qualitative research and need variation, saturation or theory building? Use purposive sampling and our purposive tool instead.
  6. Do you need fixed counts within subgroups? Use quota sampling.
  7. Are you selecting whatever is easiest to reach with no analytical rationale? That is convenience sampling; label it honestly.
  8. Is the population hidden with no frame? Use snowball or respondent driven sampling.
MethodTargets known risk?Projects to population?Reliability measurable?Effort
Judgment samplingYes, that is the pointNoYes, with a second judgeLow
Purposive samplingYes, by strategyNoRarely attemptedLow
Monetary unit samplingYes, by valueYesNot neededMedium
Haphazard samplingNoNoNoVery low
Quota samplingPartlyNoPartlyLow
Simple randomNoYesNot neededMedium
Stratified randomSomewhat, by stratumYesNot neededMedium
SystematicNoYesNot neededLow

12Troubleshooting and Common Errors

My percent agreement is 95 percent but kappa is only 0.3

Cause: the prevalence paradox. Most of your agreement is on cases neither judge selected. Fix: report kappa and PABAK together, quote the prevalence index, and use Jaccard overlap as the practical figure.

Kappa came out negative

Cause: the judges agreed less often than chance predicts, meaning they are systematically picking different cases. Fix: this is a criteria failure, not a sample size failure. Rewrite the criteria with worked examples and repeat the exercise.

One judge selected 25 cases and the other selected 5

Cause: different thresholds rather than different judgements. Fix: check the bias index. Calibrate on ten shared practice cases until the rates converge, then redo the selection.

Only one expert is available

Cause: the usual constraint. Fix: you cannot compute agreement, so the bias scorecard is your only evidence. Even a colleague re-selecting 30 cases from the pool gives you a partial kappa, which is far better than none.

My judges discussed the cases before deciding

Cause: a natural instinct that destroys the measurement. Fix: the resulting figure describes consensus, not reliability. Report it as consensus and do not call it inter-rater reliability.

A judge only recorded their selections, not their rejections

Cause: the recording sheet allowed it. Fix: kappa requires a decision on every case in the pool. You can reconstruct rejections only if the pool list is complete and the judge confirms the rest were considered and declined.

Case IDs do not match between judges

Cause: inconsistent formats such as F-142 against F-0142. Fix: the tool matches on ID text ignoring case and spacing, but leading zeros differ. Standardise the ID format before comparing.

Kappa is 0.9 and I am suspicious

Cause: healthy suspicion. Very high agreement on a genuinely subjective task usually means the judges were not independent, or the criteria were so mechanical that no judgement was involved. Check which, and say so.

Can I report a margin of error for my judgment sample?

Cause: pressure to look quantitative. Fix: no. There is no sampling distribution behind a judgement. Report the selection rate, the agreement statistics and the bias exposure instead; together they are more informative than a fake interval.

My reviewer says judgment sampling is not acceptable

Cause: often a reasonable objection when projection is required. Fix: if the aim is a population estimate they are right. If the aim is to examine high risk cases, judgment sampling is appropriate and the answer is to evidence the judgement, which is what this page produces.

Uploaded CSV columns loaded with blank values

Cause: empty cells or trailing rows. Fix: the tool skips empty cells automatically, but check the selection count shown on each judge card matches what you expect.

13Assumptions, Bias and Limitations

  • No probability basis. Selection probabilities are unknown, so there is no margin of error, no confidence interval and no projection to the population.
  • The judge is the instrument. The design inherits every limitation of the person applying it, including their training, assumptions and blind spots.
  • Reliability is not validity. Two judges can agree perfectly and both be wrong in the same direction. Agreement measures consistency only.
  • Kappa depends on prevalence and marginals. Rare selection depresses kappa and unequal strictness distorts it, which is why PABAK and the two indices are reported alongside.
  • Independence assumption. Every agreement figure assumes the judges did not see each other's decisions. If they conferred, the statistic measures conformity.
  • Complete pool assumption. Agreement is computed over the candidate pool. A judge who considered a different pool cannot be validly compared.
  • The bias index is a self assessment. It carries no sampling theory and corrects nothing; it structures disclosure.
  • Two judges is a minimum, not a comfort. Two experts from the same institution are far more likely to share a blind spot than two from different ones.

14Conclusion

Run the tool to auto-fill this conclusion with your own numbers.

What judgment sampling gives you

Judgment sampling concentrates limited effort on the cases most likely to matter. An auditor who examines the twenty most irregular entries learns more about control failure than one who examines twenty entries at random, and an inspector who visits the machines most likely to be failing finds more faults than one following a random schedule. When the aim is detection rather than estimation, deliberately targeting risk is not a weaker method, it is the correct one, and pretending otherwise wastes the expertise you are paying for.

What it costs you

You give up projection entirely. Nothing found in a judgment sample can be extrapolated to the population, because the cases were chosen precisely for being unusual. You also inherit the judge: their training, their recent experience, their institutional assumptions and their cognitive shortcuts all enter the sample invisibly. That is why the two questions worth asking are not about sample size at all, but whether another competent expert would have chosen similarly, and which known heuristics were in play.

What to check before you report

Confirm five things: the judge and their qualification are named, the criteria were written before any case was seen, the candidate pool size and selection rate are stated, agreement was computed from independent judging with kappa and PABAK reported together, and the cases only one judge selected are listed with a reason each. Then state plainly that agreement demonstrates reliability and not validity, because that distinction is the one most often blurred in practice.

What to do next

If your kappa fell below 0.40 the criteria are the problem, and no amount of extra sampling will fix them; rewrite them with worked examples and repeat. If it landed between 0.40 and 0.60, that is a normal and reportable result for expert selection, so publish it openly rather than quietly. If it came in above 0.85, verify that the judges really were independent before treating it as good news. And if you had only one judge, the highest value half day available to you is a colleague re-selecting from the same pool, because a kappa figure is the single most persuasive thing a judgment sample can offer.

15Test Yourself

1. Two judges agree on 96 of 100 cases but kappa is 0.28. What is going on?

Selection is rare, so most of the agreement is on joint exclusions and chance agreement is very high. Report kappa with PABAK and the prevalence index, and use Jaccard overlap as the practical figure.

2. What does a negative kappa mean?

The judges agreed less often than chance would predict, so they are systematically selecting different cases. The criteria are not functioning and should be rewritten.

3. Judge A picked 25 cases, Judge B picked 5. Which statistic flags this?

The bias index, |b − c| / N. It identifies unequal strictness, which is a different problem from disagreeing about individual cases and needs calibration rather than new criteria.

4. Your two judges discussed the cases first and kappa came out at 0.92. Can you report reliability?

No. Agreement between judges who conferred measures conformity. Report it as consensus and state that independent reliability was not assessed.

5. Both judges agree perfectly. Does that mean the sample is correct?

No. Agreement is reliability, not validity. Two experts with the same training can agree completely and share the same blind spot.

6. Can you compute kappa if a judge recorded only the cases they selected?

Not validly. Kappa needs an in or out decision on every case in the candidate pool, because the joint rejections are part of the calculation.

16Frequently Asked Questions

1. What is judgment sampling?

Judgment sampling selects units because an experienced person believes those units are the right ones to examine, rather than by chance. It is a non probability method, also called judgmental, purposive or selective sampling.

2. Is judgment sampling the same as purposive sampling?

They name the same family of designs. Methodology journals prefer purposive; audit, inspection and quality control practice prefer judgment. The practical difference is emphasis: judgment sampling raises the question of how reliable the judge is.

3. What is the difference between judgment sampling and convenience sampling?

Judgment sampling selects cases for a stated expert reason. Convenience sampling selects whoever is easiest to reach. If you cannot say why each case was chosen, it is convenience sampling.

4. Can you calculate a margin of error for a judgment sample?

No. There is no sampling distribution behind a judgement, so no valid margin of error, confidence interval or projection to the population exists.

5. How do you test whether a judgment sample is reliable?

Have two or more experts select independently from the same candidate pool, then compute percent agreement, Cohen kappa, PABAK and Jaccard overlap. Agreement above chance is the evidence that the judgement is reproducible.

6. What is Cohen kappa and how do you calculate it?

Cohen kappa measures agreement above chance for two raters. It equals observed agreement minus expected agreement, divided by one minus expected agreement. Values run from below zero for worse than chance up to one for perfect agreement.

7. How do you interpret a kappa value?

The Landis and Koch bands are conventional: at or below zero poor, 0.01 to 0.20 slight, 0.21 to 0.40 fair, 0.41 to 0.60 moderate, 0.61 to 0.80 substantial, above 0.80 almost perfect. For expert selection, moderate is acceptable and substantial is strong.

8. Why is my kappa low when percent agreement is high?

This is the prevalence paradox. When selection is rare, most agreement comes from cases neither judge chose, so chance agreement is very high and kappa is depressed. Report PABAK and the prevalence index alongside kappa.

9. What is PABAK?

PABAK is the prevalence adjusted bias adjusted kappa, equal to two times observed agreement minus one. It removes the dependence on how often selection occurs, so it should be reported whenever the selection rate is far from 50 percent.

10. What is Fleiss kappa used for?

Fleiss kappa gives one chance corrected agreement figure for three or more raters who each rate every case. It answers whether the panel as a whole is applying a shared standard, and should be reported with the pairwise table beneath it.

11. What is the Jaccard overlap in judgment sampling?

Jaccard overlap is the number of cases both judges selected divided by the number either selected. It ignores joint exclusions entirely, so unlike percent agreement it cannot be inflated by a large candidate pool.

12. What is a negative kappa?

A negative kappa means the judges agreed less often than chance would predict, so they are systematically choosing different cases. This is a criteria failure and the criteria should be rewritten rather than the sample enlarged.

13. What cognitive biases affect judgment sampling?

Eight are well documented in expert selection: anchoring on early cases, availability of memorable cases, confirmation of an existing hypothesis, halo and reputation effects, accessibility, representativeness, salience of extremes, and stakeholder pressure.

14. Does high agreement mean the judgment sample is correct?

No. Agreement measures reliability, not validity. Two experts with the same training can agree perfectly and share the same blind spot, which is why a high kappa with a high bias score is the most misleading combination.

15. Do the judges need to work independently?

Yes, for the agreement figure to mean anything. If judges discussed the cases or saw each other decisions, the resulting statistic measures conformity rather than reliability and should be labelled consensus.

16. How many judges do you need?

Two is the minimum for Cohen kappa and three or more allows Fleiss kappa. Two judges from different institutions are more informative than two from the same one, because shared training produces shared blind spots.

17. When is judgment sampling appropriate?

When the aim is detection rather than estimation: examining the cases most likely to contain a problem, targeting known risk, inspecting suspected failures, or selecting expert informants. It is not appropriate when a population estimate is required.

18. Is judgment sampling acceptable in auditing?

Yes, it is long established practice for targeting risk, but it does not support projection of misstatement to the population. Where projection is required, a statistical method such as monetary unit sampling is used instead.

19. How do you report judgment sampling in a paper or audit file?

Name the judge and their qualification, state the criteria written before any case was seen, give the candidate pool size and selection rate, report agreement from the independent stage with kappa and PABAK, and list the cases only one judge selected with a reason each.

20. What is the difference between reliability, validity and consensus?

Reliability means judges agree with each other and is measured by kappa. Validity means the judges are right and is measured by nothing on this page. Consensus means judges talked and converged, and is evidence of neither.

17Cite This Tool

APA: StatsUnlock. (2026). Judgment sampling bias and agreement calculator [Online tool]. https://statsunlock.com/judgment-sampling-calculator/
MLA: StatsUnlock. "Judgment Sampling Calculator." StatsUnlock, 2026, statsunlock.com/judgment-sampling-calculator/.
BibTeX: @misc{statsunlock_judg_2026, title={Judgment Sampling Bias and Agreement Calculator}, author={{StatsUnlock}}, year={2026}, howpublished={\url{https://statsunlock.com/judgment-sampling-calculator/}}}
In text methods sentence: Inter-rater agreement for the judgmental selection was computed with the StatsUnlock judgment sampling calculator, reporting Cohen kappa alongside PABAK because selection was rare.

18Related Tools

19Glossary of Terms

TermMeaning
AnchoringLetting the first cases examined set the template for all later decisions.
Availability heuristicOver selecting cases that are memorable or recent rather than relevant.
Bias index|b − c| / N; measures how differently strict the two judges are.
Candidate poolEvery unit the judges could have chosen from; the denominator for all agreement figures.
Chance agreementThe agreement two judges would reach at random given how often each selects.
Cohen kappaChance corrected agreement between two raters, from below 0 to 1.
ConsensusAgreement reached after discussion; not evidence of reliability.
Fleiss kappaChance corrected agreement across three or more raters.
Halo effectJudging a case differently because of the reputation attached to it.
Jaccard overlapCases both judges selected divided by cases either selected.
Judgment samplingSelecting units on expert judgement rather than by chance.
Landis and Koch bandsThe conventional verbal labels for ranges of kappa.
PABAKPrevalence adjusted bias adjusted kappa, equal to 2 times observed agreement minus 1.
Prevalence index|a − d| / N; high values depress kappa even when agreement is high.
Prevalence paradoxThe effect whereby rare selection produces low kappa despite high raw agreement.
ReliabilityWhether judges agree with each other; measured by kappa.
Selection rateShare of the candidate pool a judge selected.
Singleton caseA case exactly one judge selected; where the judgement did the work.
UnanimityShare of the flagged pool that every judge selected.
ValidityWhether the judges are actually right; not measured by agreement.

20References

  1. Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37-46. https://doi.org/10.1177/001316446002000104
  2. Fleiss, J. L. (1971). Measuring nominal scale agreement among many raters. Psychological Bulletin, 76(5), 378-382. https://doi.org/10.1037/h0031619
  3. Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159-174. https://doi.org/10.2307/2529310
  4. Byrt, T., Bishop, J., & Carlin, J. B. (1993). Bias, prevalence and kappa. Journal of Clinical Epidemiology, 46(5), 423-429. https://doi.org/10.1016/0895-4356(93)90018-V
  5. Feinstein, A. R., & Cicchetti, D. V. (1990). High agreement but low kappa: I. The problems of two paradoxes. Journal of Clinical Epidemiology, 43(6), 543-549. https://doi.org/10.1016/0895-4356(90)90158-L
  6. Cicchetti, D. V., & Feinstein, A. R. (1990). High agreement but low kappa: II. Resolving the paradoxes. Journal of Clinical Epidemiology, 43(6), 551-558. https://doi.org/10.1016/0895-4356(90)90159-M
  7. Sim, J., & Wright, C. C. (2005). The kappa statistic in reliability studies: use, interpretation, and sample size requirements. Physical Therapy, 85(3), 257-268. https://doi.org/10.1093/ptj/85.3.257
  8. Gwet, K. L. (2014). Handbook of Inter-Rater Reliability (4th ed.). Advanced Analytics. WorldCat record
  9. Krippendorff, K. (2004). Reliability in content analysis: some common misconceptions and recommendations. Human Communication Research, 30(3), 411-433. https://doi.org/10.1111/j.1468-2958.2004.tb00738.x
  10. Tversky, A., & Kahneman, D. (1974). Judgment under uncertainty: heuristics and biases. Science, 185(4157), 1124-1131. https://doi.org/10.1126/science.185.4157.1124
  11. Kahneman, D., Slovic, P., & Tversky, A. (1982). Judgment Under Uncertainty: Heuristics and Biases. Cambridge University Press. https://doi.org/10.1017/CBO9780511809477
  12. Nickerson, R. S. (1998). Confirmation bias: a ubiquitous phenomenon in many guises. Review of General Psychology, 2(2), 175-220. https://doi.org/10.1037/1089-2680.2.2.175
  13. Kahneman, D., Sibony, O., & Sunstein, C. R. (2021). Noise: A Flaw in Human Judgment. Little, Brown Spark. WorldCat record
  14. Meehl, P. E. (1954). Clinical versus Statistical Prediction. University of Minnesota Press. https://doi.org/10.1037/11281-000
  15. Dawes, R. M., Faust, D., & Meehl, P. E. (1989). Clinical versus actuarial judgment. Science, 243(4899), 1668-1674. https://doi.org/10.1126/science.2648573
  16. Elder, R. J., Akresh, A. D., Glover, S. M., Higgs, J. L., & Liljegren, J. (2013). Audit sampling research: a synthesis and implications for future research. Auditing: A Journal of Practice and Theory, 32(sp1), 99-129. https://doi.org/10.2308/ajpt-50394
  17. Hall, T. W., Hunton, J. E., & Pierce, B. J. (2002). Sampling practices of auditors in public accounting, industry, and government. Accounting Horizons, 16(2), 125-136. https://doi.org/10.2308/acch.2002.16.2.125
  18. Etikan, I., Musa, S. A., & Alkassim, R. S. (2016). Comparison of convenience sampling and purposive sampling. American Journal of Theoretical and Applied Statistics, 5(1), 1-4. https://doi.org/10.11648/j.ajtas.20160501.11
  19. Baker, R., et al. (2013). Summary report of the AAPOR task force on non probability sampling. Journal of Survey Statistics and Methodology, 1(2), 90-143. https://doi.org/10.1093/jssam/smt008
  20. Lohr, S. L. (2021). Sampling: Design and Analysis (3rd ed.). CRC Press. https://doi.org/10.1201/9780429298899
Last updated · Generated by STATS UNLOCK · statsunlock.com
RELATED ARTICLES

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Important Plots & Charts

Most Popular