ダブルクリックで英日反転
Applied Sciences · Engineering
Takuya Kataiwa's Domain — LLM Security & Human–AI Coexistence
A learning map of Takuya Kataiwa's research. One unifying axis: placing AI safely and trustworthily within human society — with specialisms branching into offense (security) and understanding (cognition, interpretability).
LLM Security
- Covers both attacking and defending large language models (LLMs).
- Key topics: prompt injection (hijacking AI instructions), jailbreaking (bypassing safety), and output safety evaluation.
- Prompt injection = slipping malicious commands into an AI's instruction text to override its intended behavior.
Usable Security & Interpretability
- Usable Security: reconciling robustness with ease of use — authentication design, consent/privacy UX, human cognition.
- Concrete examples: VR avatar facial-expression authentication; LLM agent browser-consent design.
- Interpretability: making an AI's reasoning legible — token embeddings, intrinsic dimensionality, n-gram models.
Computational Cognitive Science & Symbol Emergence
- Models the mind — individual and collective — computationally.
- Reproducing emotional contagion via graph neural networks (GNNs); simulating how social norms emerge.
- Symbol emergence = meaning and language arising naturally from group interaction.
Catch-Up Order
- 1 → Raw ability: competitive programming (AtCoder) + CTF (security capture-the-flag).
- 2 → AI fundamentals: ML, NLP, how LLMs work.
- 3–5 → LLM security → usable security / cognitive science → applied research (GNN, VR auth).
- Hold both wheels: offensive technical skill AND understanding of people.
→ All of Kataiwa's fields share one spine — safe, trustworthy human–AI coexistence — so master raw ability and AI fundamentals first, then climb toward the human side.
Applied Sciences · Engineering
片岩拓也の研究領域 — LLMセキュリティと「人とAIの共存」
片岩拓也氏の研究を「学習マップ」として整理したノート。全分野をつなぐ一本の軸は「AIを人間社会に安全・信頼可能な形で組み込む方法」。攻防(セキュリティ)と理解(認知・解釈可能性)へと枝分かれする。
LLMセキュリティ
- LLM(大規模言語モデル)を「騙す攻撃」と「守る防御」の両面を研究。
- 主要トピック:プロンプトインジェクション(命令文への悪意ある指示の混入)、ジェイルブレーク(安全機構の回避)、出力の安全評価。
- プロンプトインジェクション=AIへの指示文に不正コマンドを紛れ込ませ、意図した制限を突破する攻撃手法。
ユーザブルセキュリティ・解釈可能性
- ユーザブルセキュリティ(安全性と使いやすさを両立させる分野):認証設計・同意/プライバシーUX・人間認知の特性。
- 具体例:VRアバターの表情を用いた本人認証、LLMエージェントがブラウザを操作する際の同意・プライバシー設計。
- 解釈可能性(Interpretability)=「なぜそう答えたか」をAI自身が人間に説明できる度合い。トークン埋め込みや内在次元数の分析が対象。
計算認知科学・記号創発
- 個人・集団の「心の働き」を計算論的にモデル化する分野。
- グラフニューラルネットワーク(GNN=ノードとエッジの網構造を学習するAI)で感情伝染を再現、社会規範の発生過程をシミュレーション。
- 記号創発=集団内のやり取りから意味や言語が自然に生まれる現象。
追いつく順序
- ①基礎体力:競技プログラミング(AtCoder)+CTF(セキュリティ競技。脆弱なシステムに侵入して「フラグ」を奪う実践形式)。
- ②AI基礎:機械学習・NLP・LLMの仕組み。
- ③〜⑤:LLMセキュリティ → ユーザブルセキュリティ/認知科学 → 応用研究(GNN・VR認証等)。
- 「技術の攻撃側」と「人間理解」の両輪を片方に偏らず同時に育てるのが片岩氏の強みの源泉。
→ 片岩氏の全分野は「安全・信頼ある人×AI共存」という一本の軸でつながる。基礎体力とAI基礎を先に固め、そこから人間側へ登るのが正しい順序。
Applied Sciences · Engineering
Takuya Kataiwa's Domain — LLM Security & Human–AI Coexistence
A learning map of Takuya Kataiwa's research. One unifying axis: placing AI safely and trustworthily within human society — with specialisms branching into offense (security) and understanding (cognition, interpretability).
LLM Security
- Covers both attacking and defending large language models (LLMs).
- Key topics: prompt injection (hijacking AI instructions), jailbreaking (bypassing safety), and output safety evaluation.
- Prompt injection = slipping malicious commands into an AI's instruction text to override its intended behavior.
Usable Security & Interpretability
- Usable Security: reconciling robustness with ease of use — authentication design, consent/privacy UX, human cognition.
- Concrete examples: VR avatar facial-expression authentication; LLM agent browser-consent design.
- Interpretability: making an AI's reasoning legible — token embeddings, intrinsic dimensionality, n-gram models.
Computational Cognitive Science & Symbol Emergence
- Models the mind — individual and collective — computationally.
- Reproducing emotional contagion via graph neural networks (GNNs); simulating how social norms emerge.
- Symbol emergence = meaning and language arising naturally from group interaction.
Catch-Up Order
- 1 → Raw ability: competitive programming (AtCoder) + CTF (security capture-the-flag).
- 2 → AI fundamentals: ML, NLP, how LLMs work.
- 3–5 → LLM security → usable security / cognitive science → applied research (GNN, VR auth).
- Hold both wheels: offensive technical skill AND understanding of people.
→ All of Kataiwa's fields share one spine — safe, trustworthy human–AI coexistence — so master raw ability and AI fundamentals first, then climb toward the human side.
Applied Sciences · Engineering
片岩拓也の研究領域 — LLMセキュリティと「人とAIの共存」
片岩拓也氏の研究を「学習マップ」として整理したノート。全分野をつなぐ一本の軸は「AIを人間社会に安全・信頼可能な形で組み込む方法」。攻防(セキュリティ)と理解(認知・解釈可能性)へと枝分かれする。
LLMセキュリティ
- LLM(大規模言語モデル)を「騙す攻撃」と「守る防御」の両面を研究。
- 主要トピック:プロンプトインジェクション(命令文への悪意ある指示の混入)、ジェイルブレーク(安全機構の回避)、出力の安全評価。
- プロンプトインジェクション=AIへの指示文に不正コマンドを紛れ込ませ、意図した制限を突破する攻撃手法。
ユーザブルセキュリティ・解釈可能性
- ユーザブルセキュリティ(安全性と使いやすさを両立させる分野):認証設計・同意/プライバシーUX・人間認知の特性。
- 具体例:VRアバターの表情を用いた本人認証、LLMエージェントがブラウザを操作する際の同意・プライバシー設計。
- 解釈可能性(Interpretability)=「なぜそう答えたか」をAI自身が人間に説明できる度合い。トークン埋め込みや内在次元数の分析が対象。
計算認知科学・記号創発
- 個人・集団の「心の働き」を計算論的にモデル化する分野。
- グラフニューラルネットワーク(GNN=ノードとエッジの網構造を学習するAI)で感情伝染を再現、社会規範の発生過程をシミュレーション。
- 記号創発=集団内のやり取りから意味や言語が自然に生まれる現象。
追いつく順序
- ①基礎体力:競技プログラミング(AtCoder)+CTF(セキュリティ競技。脆弱なシステムに侵入して「フラグ」を奪う実践形式)。
- ②AI基礎:機械学習・NLP・LLMの仕組み。
- ③〜⑤:LLMセキュリティ → ユーザブルセキュリティ/認知科学 → 応用研究(GNN・VR認証等)。
- 「技術の攻撃側」と「人間理解」の両輪を片方に偏らず同時に育てるのが片岩氏の強みの源泉。
→ 片岩氏の全分野は「安全・信頼ある人×AI共存」という一本の軸でつながる。基礎体力とAI基礎を先に固め、そこから人間側へ登るのが正しい順序。
出典・参考 / Sources
Related notes
- Agentic Commerce — ACP and Visibility into Being 'Bought by AI'
- AI Search Evaluation: The 12 Metrics — Gateway
- AI Search Evaluation ①Citation Rate — How Many URLs Are Pulled In Per Answer
- AI Search Evaluation ②Source Diversity — How Unskewed the Cited Sources Are
- AI Search Evaluation ③Accuracy Score — How Often It Answers Factual Questions Correctly
- AI Search Evaluation ④Answer Length — How Many Characters It Returns to the User on Average