Back to Blog

Dual‑screening conflicts and Cohen’s κ at title/abstract

How to read Cohen’s kappa during dual title/abstract screening, what counts as a conflict vs a Maybe, and how to run the resolution queue without gaming the metric.

George Burchell
September 3, 2026
4 min read
Dual screening conflict matrix with Cohen’s kappa note about imbalance and the resolution paths (discuss or third screener) highlighted

Dual independent screening in practice

At title/abstract, I run two independent decisions per record using three labels: Include, Maybe, Exclude. "Maybe" is a first‑class state — not an error — and it should flow into the same resolution queue as disagreements.

Conflict vs Maybe

  • Conflict: the two screeners choose different labels (e.g., Include vs Exclude, Include vs Maybe, Maybe vs Exclude).
  • Same: both choose the same label (Include/Include, Maybe/Maybe, Exclude/Exclude).
  • A Maybe is not “half include.” Treat it as a disagreement that needs resolution, not a soft pass.

The conflict matrix

Think of a 3×3 grid with Screener A down the left (Include / Maybe / Exclude) and Screener B across the top. The diagonal cells are agreements; everything off‑diagonal is a conflict to resolve. Your queue is simply “all non‑diagonal rows.”

Cohen’s κ (kappa) during title/abstract screening

Cohen’s κ is chance‑corrected agreement. It answers: given each screener’s base rates, how much better is their agreement than chance? κ can look “low” even when raw agreement looks “high” if most records are Exclude. That’s normal for title/abstract screening and it doesn’t mean your team is sloppy.

Quick numeric intuition (binary for simplicity — treating Maybe as a disagreement that goes to resolution):

  • Suppose each screener Includes 5% and Excludes 95%.
  • Observed raw agreement is 92% (e.g., they match on most Excludes and some Includes).
  • Expected agreement by chance from the marginals is 0.05×0.05 + 0.95×0.95 = 0.0025 + 0.9025 = 0.905 (90.5%).
  • κ = (0.92 − 0.905) / (1 − 0.905) ≈ 0.015 / 0.095 ≈ 0.16.

So: 92% raw agreement but κ ≈ 0.16 (“slight/fair”). With imbalanced includes, κ penalizes agreement that comes mostly from both people Excluding. Flip the situation (balanced include rates) and κ can jump even if raw agreement barely moves. That’s why κ is a context signal, not a quality badge. Cohen's kappa helps you spot when agreement is driven by a dominant class.

Resolution queue: how the work actually moves

  • Put all disagreements and all Maybes into a single queue.
  • Resolution path: discuss → if unresolved, bring in a third screener → record the resolved label.
  • Don’t “fix” κ by relabeling after the fact. If you want κ computed on a recoded binary scheme, write that rule in the protocol up front and stick to it.

Protocol decisions to make before you start

  • Dual vs single screening at title/abstract (many teams do dual on a sample, then decide).
  • Exactly which labels are used and how Maybe flows.
  • Who resolves conflicts and when to pull a third screener.
  • How κ will be reported (scope, unit, and any recoding rules).

PRISMA and honest reporting

PRISMA expects you to report counts for identification, screening, eligibility, and inclusion honestly. Conflicts and Maybes are part of the screening process — resolve them, record the final decisions, and keep your audit trail clean. See the PRISMA 2020 statement and the Cochrane Handbook for terminology and examples.

Where this fits in your workflow

If you’re new to screening workflows, this post fits between search/export and full‑text. For a broader walkthrough, see our primer: Getting started with systematic reviews. Tools can show κ next to the work as a concept cue, but κ never replaces the conflict queue.


My rule of thumb: treat Cohen’s κ as a live diagnostic that you read alongside the queue. High percent agreement with a low κ during imbalanced screening is expected; it’s a prompt to check whether disagreements are concentrated in Includes/Maybes, not a reason to game labels. Keep the queue moving and record the resolution. The statistics should follow the method, not lead it.

Related Articles

systematic review
screening

Why 'Just Read the Abstracts' Is the Biggest Misconception in Evidence Screening

Why "Just Read the Abstracts" Is the Biggest Misconception in Evidence Screening TL;DR: Screening abstracts isn't a simple yes/no process—it's complex decision-...

4 min read
systematic review
AI automation

The One Research Task I'd Hand Over to AI Tomorrow

The One Research Task I'd Hand Over to AI Tomorrow TL;DR: Screening is the most time-consuming bottleneck in systematic reviews, and AI is perfectly suited to h...

7 min read
systematic review
screening

What Is Screening in a Systematic Literature Review?

What Is Screening in a Systematic Literature Review? Most people outside research imagine systematic reviews are about reading papers and having insights. Anyon...

9 min read

Ready to Streamline Your Systematic Review?

Experience the power of AI-assisted screening and cut your review time by up to 80%. Join thousands of researchers who trust our platform for their systematic reviews.

George Burchell - Systematic Review Expert

About the Author

Connect on LinkedIn

George Burchell

George Burchell is a specialist in systematic literature reviews and scientific evidence synthesis with significant expertise in integrating advanced AI technologies and automation tools into the research process. With over four years of consulting and practical experience, he has developed and led multiple projects focused on accelerating and refining the workflow for systematic reviews within medical and scientific research.

Systematic Reviews
Evidence Synthesis
AI Research Tools
Research Automation