til/social-sciences/psychology/inter-rater-agreement-kappa-alpha
inter-rater-agreement-kappa-alpha.mdupdated 2026-07-162597 words
ダブルクリックで英日反転
Social Sciences · Psychology

Cohen's κ / Krippendorff's α — Measuring Whether Raters Agree

EN

When multiple raters label the same data, raw agreement inflates results because some matches happen by chance. κ and α correct for that, giving a trustworthy reliability figure.

Why raw agreement is not enough

  • Manual annotation results shift when the rater changes.
  • Raw agreement rate counts chance matches as real agreement.
  • A chance-corrected metric is required before trusting any accuracy score.

Cohen's κ — two raters, nominal scale

  • Formula: κ = (P_o − P_e) ÷ (1 − P_e), where P_o = observed rate, P_e = chance-expected rate.
  • Conventional benchmarks (Landis & Koch 1977): 0.61–0.80 substantial; 0.81+ almost perfect.
  • Limitation: two raters only; does not handle ordinal or interval data natively.

Krippendorff's α — three+ raters, any scale

  • Handles missing data and nominal, ordinal, interval, and ratio scales in one formula.
  • Rule of thumb: α ≥ 0.80 = trustworthy conclusions; ≥ 0.667 = preliminary discussion only.
  • Preferred over κ when rater count or scale type varies across tasks.

Application in AI-search evaluation

  • DoD category E-1 (Reliability) sets minimum κ ≥ 0.70 or α ≥ 0.667.
  • 2–3 raters score factual, brand-mention, and freshness judgements in parallel each week.
  • Example: observed 0.90, chance 0.50 → κ = 0.80, confirming stable scoring logic.
Use κ for two-rater nominal tasks; switch to α when raters, scales, or missing data complicate things — and set a minimum threshold before the annotation campaign starts.
社会科学 · 心理学

Cohen's κ / Krippendorff's α — 評価者間一致度の​測り方

JP

複数の​評価者が​同じ​データに​ラベルを​付けると、​偶然の​一致が​生まれ、​生の​一致率は​信頼性を​過大評価する。​κ(カッパ)と​α​(アルファ)は​偶然一致を​補正し、​正確な​信頼性指標を​提供する。

な​ぜ生の​一致率では​不十分か

  • 評価者が​変わると​手動アノテーション​(=データへの​ラベル付け)の​結果が​変動する。
  • 生の​一致率は​偶然の​一致も​「真の​合意」と​して​計上してしまう。
  • 精度スコアを​信頼するには、​偶然一致を​補正した​指標が​不可欠。

Cohen's κ — 2名・名義尺度向け

  • 算式: κ = (P_o − P_e) ÷ (1 − P_e)。​P_o=実測一致率、​P_e=偶然期待一致率。
  • 判断基準​(Landis & Koch 1977)​: 0.61–0.80=実質的一致、​0.81以上​=ほぼ完全一致。
  • 制約: 評価者は​2名のみ、​順序尺度や​間隔尺度には​対応していない。

Krippendorff's α — 3名以上​・​あらゆる​尺度向け

  • 欠損データ、​名義・順序・間隔・​比率​(レシオ)​尺度を​単一の​算式で​扱える​汎用指標。
  • 目安: α ≥ 0.80 で​「結論を​信頼可」、​≥ 0.667 で​「暫定議論のみ」​(Krippendorff 2004)。
  • 評価者数や​尺度の​種類が​混在する​タスクでは​κより​αを​使う​ほうが​安全。

AI検索評価への​適用例

  • DoD​(完了の​定義)​カテゴリ E-1​(​信頼性)は​ κ ≥ 0.70 または​ α ≥ 0.667 を​最低基準に​設定。
  • 事実確認・ブランド言及・鮮度の​判定を​2〜3名が​並行採点し、​毎週​κ/αを​報告。
  • 実例: 実測一致率0.90、​偶然期待値0.50 → κ = 0.80 で​採点ロジックの​安定を​確認。
2名・名義尺度ならκ、​評価者数・尺度・欠損が​複雑ならαを​選び、​アノテーション開始前に​最低閾値を​決めて​おく。
148 notestil