til/social-sciences/psychology/ceiling-floor-effect
ceiling-floor-effect.mdupdated 2026-07-162391 words
ダブルクリックで英日反転
Social Sciences · Psychology

Ceiling Effect / Floor Effect

EN

When every participant scores perfectly (ceiling) or fails completely (floor), the measuring instrument jams — differences between subjects vanish and the test item carries zero information.

Why it matters

  • Ceiling item: all respondents answer correctly → no differentiation possible.
  • Floor item: all respondents fail → same problem, opposite end.
  • Both scenarios make the item useless for ranking or comparing participants.

Detection rules of thumb

  • p ≈ 1.0 → ceiling effect (item too easy); flag for exclusion.
  • p ≈ 0.0 → floor effect (item too hard); flag for exclusion.
  • p = 0.4–0.6: maximum information — keep these as the main metric.
  • IRT (Item Response Theory) adds the discrimination parameter 'a' for finer filtering.

AI search evaluation context

  • Applies to accuracy scoring across query sets with many factual questions.
  • Items stuck at p ≈ 1 or p ≈ 0 must be flagged and excluded from aggregation.
  • Example (77 factual-static questions): 12 ceiling + 3 floor = 15 invalid → valid set 62.
  • Removing invalid items requires redoing the sample-size calculation.

Related concepts

  • Cohen's d: effect size measure — ceiling/floor collapse variance before d is even computable.
  • Nomological network: construct validity framework; difficulty design sits one level below it.
Always check the proportion correct before aggregating: items at p ≈ 0 or 1 add no signal and must be excluded.
社会科学 · 心理学

天井効果・床効果

JP

テスト項目で​全員が​正解​(天井)または​全員が​不正解​(床)に​なると、​測定尺度が​飽和して​個人差を​捉えられなくなる​現象。​その​項目は​比較情報を​ゼロしか​持たない。

なぜ問題か

  • 天井効果:全回​答者が​正解 → ランキング・比較に​使えない。
  • 床効果:全回​答者が​不正解 → 同様に​弁別不能。
  • どちらも​その​項目の​情報量が​ゼロに​なり、​評価指標と​して​機能しない。

検出の​目安

  • 正答率 p ≈ 1.0 → 天井効果​(易しすぎ)​:除外候補と​して​フラグを​立てる。
  • p ≈ 0.0 → 床効果​(難しすぎ)​:同様に​除外候補。
  • p = 0.4–0.6 の​項目が​情報量最大 → メイン指標と​して​採用。
  • IRT​(項目反応理論)では​識別力パラメータ​「a」も​併用して​さらに​絞り込む。

AI検索評価への​応用

  • 多数の​事実質問を​含むクエリセットで​正答率を​分布で​確認する​ことが​前提。
  • p ≈ 1 または​ p ≈ 0 の​項目は​集計から​除外し、​必要なら​設問を​再設計する。
  • 例​(77問の​事実静的質問)​:天井12問+床3問=15問無効 → 有効62問に​縮小。
  • 除外後は​必要サンプルサイズの​再計算が​必要。

関連概念

  • Cohen's d​(効果​量)​:天井・床では​分散が​ほぼゼロに​なり、​d を​計算する​前に​崩壊する。
  • ノモロジカルネットワーク​(構成概念​妥当性の​上位概念)​:難易度設計は​その​下位の​信頼性の​問題。
集計前に​必ず正答率分布を​確認し、​p ≈ 0 または​ 1 の​項目は​除外してから​比較する。
148 notestil