Dual‑screening conflicts and Cohen’s κ at title/abstract
How to read Cohen’s kappa during dual title/abstract screening, what counts as a conflict vs a Maybe, and how to run the resolution queue without gaming the metric.

Dual independent screening in practice
At title/abstract, I run two independent decisions per record using three labels: Include, Maybe, Exclude. "Maybe" is a first‑class state — not an error — and it should flow into the same resolution queue as disagreements.
Conflict vs Maybe
- Conflict: the two screeners choose different labels (e.g., Include vs Exclude, Include vs Maybe, Maybe vs Exclude).
- Same: both choose the same label (Include/Include, Maybe/Maybe, Exclude/Exclude).
- A Maybe is not “half include.” Treat it as a disagreement that needs resolution, not a soft pass.
The conflict matrix
Think of a 3×3 grid with Screener A down the left (Include / Maybe / Exclude) and Screener B across the top. The diagonal cells are agreements; everything off‑diagonal is a conflict to resolve. Your queue is simply “all non‑diagonal rows.”
Cohen’s κ (kappa) during title/abstract screening
Cohen’s κ is chance‑corrected agreement. It answers: given each screener’s base rates, how much better is their agreement than chance? κ can look “low” even when raw agreement looks “high” if most records are Exclude. That’s normal for title/abstract screening and it doesn’t mean your team is sloppy.
Quick numeric intuition (binary for simplicity — treating Maybe as a disagreement that goes to resolution):
- Suppose each screener Includes 5% and Excludes 95%.
- Observed raw agreement is 92% (e.g., they match on most Excludes and some Includes).
- Expected agreement by chance from the marginals is 0.05×0.05 + 0.95×0.95 = 0.0025 + 0.9025 = 0.905 (90.5%).
- κ = (0.92 − 0.905) / (1 − 0.905) ≈ 0.015 / 0.095 ≈ 0.16.
So: 92% raw agreement but κ ≈ 0.16 (“slight/fair”). With imbalanced includes, κ penalizes agreement that comes mostly from both people Excluding. Flip the situation (balanced include rates) and κ can jump even if raw agreement barely moves. That’s why κ is a context signal, not a quality badge. Cohen's kappa helps you spot when agreement is driven by a dominant class.
Resolution queue: how the work actually moves
- Put all disagreements and all Maybes into a single queue.
- Resolution path: discuss → if unresolved, bring in a third screener → record the resolved label.
- Don’t “fix” κ by relabeling after the fact. If you want κ computed on a recoded binary scheme, write that rule in the protocol up front and stick to it.
Protocol decisions to make before you start
- Dual vs single screening at title/abstract (many teams do dual on a sample, then decide).
- Exactly which labels are used and how Maybe flows.
- Who resolves conflicts and when to pull a third screener.
- How κ will be reported (scope, unit, and any recoding rules).
PRISMA and honest reporting
PRISMA expects you to report counts for identification, screening, eligibility, and inclusion honestly. Conflicts and Maybes are part of the screening process — resolve them, record the final decisions, and keep your audit trail clean. See the PRISMA 2020 statement and the Cochrane Handbook for terminology and examples.
Where this fits in your workflow
If you’re new to screening workflows, this post fits between search/export and full‑text. For a broader walkthrough, see our primer: Getting started with systematic reviews. Tools can show κ next to the work as a concept cue, but κ never replaces the conflict queue.
My rule of thumb: treat Cohen’s κ as a live diagnostic that you read alongside the queue. High percent agreement with a low κ during imbalanced screening is expected; it’s a prompt to check whether disagreements are concentrated in Includes/Maybes, not a reason to game labels. Keep the queue moving and record the resolution. The statistics should follow the method, not lead it.
