ダブルクリックで英日反転
Applied Sciences · Engineering
AI Search Evaluation: The 12 Metrics
The GMO AI Search Lab compares six engines (ChatGPT Search, Gemini, Google AI Overview, AI Mode, Copilot, Claude) across 600 balanced questions using 12 observational metrics mapped to five latent constructs.
The 12 Metrics at a Glance
- Metrics 1–8 are established in prior literature (Princeton GEO benchmark, HELM, etc.)
- Metrics 9–12 are new: brand mention rate, citation consistency, freshness, DA correlation
- Each metric is an observable that moves in the shadow of one or more latent constructs (C1–C5)
Construct Map
- C1 Citation Fidelity ← citation rate (1) + accuracy score (3)
- C2 Source Trustworthiness ← diversity (2) + JP domain ratio (5) + TLD distribution (6) + consistency (10) + DA correlation (12)
- C3 Brand Visibility ← brand mention rate (9); C4 Agent Task Completion ← answer length (4) + latency (8); C5 Freshness ← freshness (11)
How Results Are Used
- Weekly report: all 12 metrics aggregated across 6 engines × sector
- Client deliverable: brand mention rate + citation overlap + DA correlation
- Academic preprint: nomological network validation via hypotheses H1–H6
Key Supporting Concepts
- Jaccard similarity — used in citation overlap (7) and citation consistency (10)
- Cohen's d / κ / Krippendorff's α — effect size and inter-rater reliability for annotation
- Nomological network — the validity argument linking constructs C1–C5 to the 12 metrics
→ Separate the 12 observable metrics from the 5 latent constructs — metrics are the instruments; constructs are the capabilities they measure.
応用科学 · エンジニアリング
AI検索評価:12指標の全体像
GMO AIサーチラボでは、ChatGPT Search・Gemini・Google AI Overview・AI Mode・Copilot・Claudeの6エンジンを600問のバランスクエリで比較するため、12の観測指標を使用する。各指標は5つの潜在構成概念(ラテント・コンストラクト)と対応している。
12指標の概要
- 指標1〜8:Princeton GEOベンチマーク・HELMなど先行研究で確立済みの標準指標
- 指標9〜12:GMO独自追加。ブランド言及率・引用一貫性・鮮度・DA相関(ドメイン権威と引用頻度の相関)
- 各指標は潜在構成概念C1〜C5の「観測可能な影」として機能する
構成概念との対応マップ
- C1 引用忠実性 ← 引用率(1) + 正確性スコア(3)
- C2 情報源信頼性バイアス ← 多様性(2) + 日本ドメイン比率(5) + TLD分布(6) + 一貫性(10) + DA相関(12)
- C3 ブランド可視性 ← ブランド言及率(9);C4 タスク完遂 ← 回答長(4) + レイテンシ(8);C5 鮮度 ← 鮮度指標(11)
成果物ごとの活用方法
- 週次レポート:12指標を全6エンジン×セクターで集計
- クライアント向け:ブランド言及率 + 引用重複率 + DA相関の3点セット
- 学術プレプリント:仮説H1〜H6を通じた名法ネットワーク(ノモロジカル・ネットワーク)の妥当性検証
理解に必要な周辺概念
- Jaccard類似度(=集合の重なり度合いの指標):引用重複(7)・引用一貫性(10)で使用
- Cohen's d・κ・Krippendorff's α:エンジン間の効果量とアノテーション評価者間信頼性の測定
- ノモロジカル・ネットワーク:C1〜C5と12指標を結ぶ構成概念妥当性の論拠
→ 「12の観測指標」と「5つの潜在構成概念」を混同しない——指標は計測器、構成概念はそれが測る能力そのもの。
Applied Sciences · Engineering
AI Search Evaluation: The 12 Metrics
The GMO AI Search Lab compares six engines (ChatGPT Search, Gemini, Google AI Overview, AI Mode, Copilot, Claude) across 600 balanced questions using 12 observational metrics mapped to five latent constructs.
The 12 Metrics at a Glance
- Metrics 1–8 are established in prior literature (Princeton GEO benchmark, HELM, etc.)
- Metrics 9–12 are new: brand mention rate, citation consistency, freshness, DA correlation
- Each metric is an observable that moves in the shadow of one or more latent constructs (C1–C5)
Construct Map
- C1 Citation Fidelity ← citation rate (1) + accuracy score (3)
- C2 Source Trustworthiness ← diversity (2) + JP domain ratio (5) + TLD distribution (6) + consistency (10) + DA correlation (12)
- C3 Brand Visibility ← brand mention rate (9); C4 Agent Task Completion ← answer length (4) + latency (8); C5 Freshness ← freshness (11)
How Results Are Used
- Weekly report: all 12 metrics aggregated across 6 engines × sector
- Client deliverable: brand mention rate + citation overlap + DA correlation
- Academic preprint: nomological network validation via hypotheses H1–H6
Key Supporting Concepts
- Jaccard similarity — used in citation overlap (7) and citation consistency (10)
- Cohen's d / κ / Krippendorff's α — effect size and inter-rater reliability for annotation
- Nomological network — the validity argument linking constructs C1–C5 to the 12 metrics
→ Separate the 12 observable metrics from the 5 latent constructs — metrics are the instruments; constructs are the capabilities they measure.
応用科学 · エンジニアリング
AI検索評価:12指標の全体像
GMO AIサーチラボでは、ChatGPT Search・Gemini・Google AI Overview・AI Mode・Copilot・Claudeの6エンジンを600問のバランスクエリで比較するため、12の観測指標を使用する。各指標は5つの潜在構成概念(ラテント・コンストラクト)と対応している。
12指標の概要
- 指標1〜8:Princeton GEOベンチマーク・HELMなど先行研究で確立済みの標準指標
- 指標9〜12:GMO独自追加。ブランド言及率・引用一貫性・鮮度・DA相関(ドメイン権威と引用頻度の相関)
- 各指標は潜在構成概念C1〜C5の「観測可能な影」として機能する
構成概念との対応マップ
- C1 引用忠実性 ← 引用率(1) + 正確性スコア(3)
- C2 情報源信頼性バイアス ← 多様性(2) + 日本ドメイン比率(5) + TLD分布(6) + 一貫性(10) + DA相関(12)
- C3 ブランド可視性 ← ブランド言及率(9);C4 タスク完遂 ← 回答長(4) + レイテンシ(8);C5 鮮度 ← 鮮度指標(11)
成果物ごとの活用方法
- 週次レポート:12指標を全6エンジン×セクターで集計
- クライアント向け:ブランド言及率 + 引用重複率 + DA相関の3点セット
- 学術プレプリント:仮説H1〜H6を通じた名法ネットワーク(ノモロジカル・ネットワーク)の妥当性検証
理解に必要な周辺概念
- Jaccard類似度(=集合の重なり度合いの指標):引用重複(7)・引用一貫性(10)で使用
- Cohen's d・κ・Krippendorff's α:エンジン間の効果量とアノテーション評価者間信頼性の測定
- ノモロジカル・ネットワーク:C1〜C5と12指標を結ぶ構成概念妥当性の論拠
→ 「12の観測指標」と「5つの潜在構成概念」を混同しない——指標は計測器、構成概念はそれが測る能力そのもの。
Related notes
- Agentic Commerce — ACP and Visibility into Being 'Bought by AI'
- AI Search Evaluation ①Citation Rate — How Many URLs Are Pulled In Per Answer
- AI Search Evaluation ②Source Diversity — How Unskewed the Cited Sources Are
- AI Search Evaluation ③Accuracy Score — How Often It Answers Factual Questions Correctly
- AI Search Evaluation ④Answer Length — How Many Characters It Returns to the User on Average
- AI Search Evaluation ⑤Japanese Domain Ratio — How Often It Pulls In .jp-Family Domains