10-citation-consistency.mdupdated 2026-07-163056 words
ダブルクリックで英日反転
Applied Sciences · Engineering

AI Search Evaluation ⑩ Citation Consistency

EN

Citation Consistency measures whether an AI search engine cites the same domains for the same query week over week, using Jaccard similarity. It captures ranking volatility and determines how many observation weeks are needed to verify an intervention's effect.

Definition & Formula

  • Jaccard similarity of cited domains: |Domains_weekN ∩ Domains_weekN+1| ÷ |Domains_weekN ∪ Domains_weekN+1|
  • Computed per query, then averaged; per-engine × query breakdowns are also examined.
  • Temporal extension: build a weekly Jaccard time series and inspect its variance.

Why It Matters

  • Business: high consistency → stable SERP-like signal; easy to measure intervention effect.
  • Low consistency → high noise; harder to attribute domain changes to actual actions.
  • Academic grounding: test-retest reliability (Cohen 1988); Krippendorff's α can supplement.
  • Helps judge whether an engine's retrieval policy is a stable trait or a shifting state.

Worked Example (hypothetical, 77 factual queries × 2 weeks)

  • ChatGPT Search: 0.78 — highly stable.
  • Copilot: 0.71 · Gemini: 0.65.
  • AI Overview: 0.42 — ~60% of cited domains turn over each week.
  • AI Overview requires ~4 consecutive weeks of data to correct for variance when verifying a client intervention.

Use in the AI-Search Project

  • Applied to all queries in daily longitudinal tracking, rolled up weekly.
  • Aggregated by engine × intent category × week; visualised as a volatility trend chart.
  • Represents the temporal-stability aspect of C2 Source Trustworthiness Bias.
  • Also used to distinguish state vs. trait in engine retrieval behaviour.
A score of 0.42 means 60% domain turnover per week — you need 4+ weeks of data before an AI Overview intervention can be reliably evaluated.
Applied Sciences · Engineering

AI検索評価 ⑩ 引用​一貫​性​(Citation Consistency)

JP

引用​一貫​性​(Citation Consistency)とは、​同じ​クエリに​対して​AI検索エンジンが​週を​またいで​同じ​ドメインを​引用するかを​Jaccard類似度​(=集合の​重複率)で​測る​指標。​ランキング変動の​激しさを​捉え、​施策効果の​検証に​必要な​観測週数の​目安を​与える。

定義と​計算式

  • 一貫性 = |第N週の​引用ドメイン ∩ 第N+1週| ÷ |第N週 ∪ 第N+1週|​(Jaccard類似度)
  • クエリごとに​計算して​平均。​エンジン×クエリ単位の​内訳も​分析する。
  • 時系列延長:週次Jaccard値の​系列を​作り、​分散を​観察する。

なぜ重要か

  • ビジネス観点:​一貫性が​高い​=安定した​SERP​(検索結果​ページ)に​近く、​施策効果を​読みやすい。
  • 一貫性が​低い=ノイズが​大きく、​ドメイン変動を​施策に​帰属しにくい。
  • 学術根拠:再検査​信頼性​(Cohen 1988)​・Krippendorffの​α​(=評価者間一致度の​指標)で​補完可能。
  • エンジンの​検索方​針が​安定した​特性​(trait)か​変動する​状態​(state)かを​判断する​材料に​なる。

計算例​(仮想シナリオ:77問×2週)

  • ChatGPT Search: 0.78​(高安定)
  • Copilot: 0.71 · Gemini: 0.65
  • AI Overview: 0.42 ── 先週引用した​ドメインの​約60%が​今週は​入れ替わる。
  • AI Overview で​施策効果を​検証するには、​分散補正の​ため連続4週以上の​データが​必要。

AIサーチプロジェクトでの​活用

  • 日次縦断トラッキングの​全ク​エリを​対象に​週次集計。
  • エンジン×意図カテゴリ×週で​集計し、​変動トレンドチャートで​可視化。
  • C2​「情報源​信頼性バイアス」の​時間的​安定性​(temporal-stability)​側面を​構成。
  • エンジンの​検索ポリシーが​traitか​stateかを​区別する​指標と​しても​使用。
スコア0.42は​週ごとに​60%の​ドメインが​入れ替わる​ことを​意味し、​AI Overviewの​施策評価には​最低4週の​連続データが​必要に​なる。
148 notestil