til/applied-sciences/engineering/ai-search-metrics
ai-search-metrics.mdupdated 2026-07-162666 words
ダブルクリックで英日反転
Applied Sciences · Engineering

AI Search Evaluation: The 12 Metrics

EN

The GMO AI Search Lab compares six engines (ChatGPT Search, Gemini, Google AI Overview, AI Mode, Copilot, Claude) across 600 balanced questions using 12 observational metrics mapped to five latent constructs.

The 12 Metrics at a Glance

  • Metrics 1–8 are established in prior literature (Princeton GEO benchmark, HELM, etc.)
  • Metrics 9–12 are new: brand mention rate, citation consistency, freshness, DA correlation
  • Each metric is an observable that moves in the shadow of one or more latent constructs (C1–C5)

Construct Map

  • C1 Citation Fidelity ← citation rate (1) + accuracy score (3)
  • C2 Source Trustworthiness ← diversity (2) + JP domain ratio (5) + TLD distribution (6) + consistency (10) + DA correlation (12)
  • C3 Brand Visibility ← brand mention rate (9); C4 Agent Task Completion ← answer length (4) + latency (8); C5 Freshness ← freshness (11)

How Results Are Used

  • Weekly report: all 12 metrics aggregated across 6 engines × sector
  • Client deliverable: brand mention rate + citation overlap + DA correlation
  • Academic preprint: nomological network validation via hypotheses H1–H6

Key Supporting Concepts

  • Jaccard similarity — used in citation overlap (7) and citation consistency (10)
  • Cohen's d / κ / Krippendorff's α — effect size and inter-rater reliability for annotation
  • Nomological network — the validity argument linking constructs C1–C5 to the 12 metrics
Separate the 12 observable metrics from the 5 latent constructs — metrics are the instruments; constructs are the capabilities they measure.
応用科学 · エンジニアリング

AI検索評価:12指標の​全体​像

JP

GMO AIサーチラボでは、​ChatGPT Search・Gemini・Google AI Overview・AI Mode・Copilot・Claudeの​6エンジンを​600問の​バランスクエリで​比較する​ため、​12の​観測指標を​使用する。​各指標は​5つの​潜在構成概念​(ラテント・コンストラクト)と​対応している。

12指標の​概要

  • 指標1〜8:Princeton GEOベンチマーク・HELMなど先行研究で​確立済みの​標準指標
  • 指標9〜​12:GMO独自追加。​ブランド言及率・引用​一貫性・鮮度・DA相関​(ドメイン権威と​引用頻度の​相関)
  • 各指標は​潜在構成概念C1〜C5の​「観測可能な影」と​して​機能する

構成概念との​対応マップ

  • C1 引用​忠実性 ← 引用率(1) + ​正確性スコア(3)
  • C2 情報源​信頼性バイアス ← ​多様性(2) + 日本ドメイン比率(5) + TLD分布(6) + ​一貫性(10) + DA相関(12)
  • C3 ブランド可視性 ← ブランド言及率(9);C4 タスク完遂 ← 回答長(4) + レイテンシ(8);C5 鮮度 ← 鮮度指標(11)

成果物ごとの​活用方​法

  • 週次レポート:12指標を​全6​エンジン×セクターで​集計
  • クライアント向け:ブランド言及率 + 引用重複率 + DA相関の​3点セット
  • 学術プレプリント:仮説H1〜H6を​通じた​名法ネットワーク​(ノモロジカル・ネットワーク)の​妥当性検証

理解に​必要な​周辺概念

  • Jaccard類似度​(=集合の​重なり度合いの​指標)​:引用重複(7)・引用​一貫​性(10)で​使用
  • Cohen's d・κ・Krippendorff's α:エンジン間の​効果量と​アノテーション評価者間​信頼性の​測定
  • ノモロジカル・ネットワーク:C1〜C5と​12指標を​結ぶ​構成概念​妥当性の​論拠
「12の​観測指標」と​「5つの​潜在構成概念」を​混同しない​——指標は​計測器、​構成概念は​それが​測る能力​そのもの。
148 notestil