til/applied-sciences/engineering/kataiwa-takuya-domain
kataiwa-takuya-domain.mdupdated 2026-07-163089 words
ダブルクリックで英日反転
Applied Sciences · Engineering

Takuya Kataiwa's Domain — LLM Security & Human–AI Coexistence

EN

A learning map of Takuya Kataiwa's research. One unifying axis: placing AI safely and trustworthily within human society — with specialisms branching into offense (security) and understanding (cognition, interpretability).

LLM Security

  • Covers both attacking and defending large language models (LLMs).
  • Key topics: prompt injection (hijacking AI instructions), jailbreaking (bypassing safety), and output safety evaluation.
  • Prompt injection = slipping malicious commands into an AI's instruction text to override its intended behavior.

Usable Security & Interpretability

  • Usable Security: reconciling robustness with ease of use — authentication design, consent/privacy UX, human cognition.
  • Concrete examples: VR avatar facial-expression authentication; LLM agent browser-consent design.
  • Interpretability: making an AI's reasoning legible — token embeddings, intrinsic dimensionality, n-gram models.

Computational Cognitive Science & Symbol Emergence

  • Models the mind — individual and collective — computationally.
  • Reproducing emotional contagion via graph neural networks (GNNs); simulating how social norms emerge.
  • Symbol emergence = meaning and language arising naturally from group interaction.

Catch-Up Order

  • 1 → Raw ability: competitive programming (AtCoder) + CTF (security capture-the-flag).
  • 2 → AI fundamentals: ML, NLP, how LLMs work.
  • 3–5 → LLM security → usable security / cognitive science → applied research (GNN, VR auth).
  • Hold both wheels: offensive technical skill AND understanding of people.
All of Kataiwa's fields share one spine — safe, trustworthy human–AI coexistence — so master raw ability and AI fundamentals first, then climb toward the human side.
Applied Sciences · Engineering

片岩拓也の​研究領域 — LLMセキュリティと​「人と​AIの​共存」

JP

片岩拓也氏の​研究を​「学習マップ」と​して​整理した​ノート。​全分​野を​つなぐ​一本の​軸は​「AIを​人間社会に​安全・​信頼可能な形で​組み込む方​法」。​攻防​(セキュリティ)と​理解​(認知・解釈​可能性)​へと​枝分かれする。

LLMセキュリティ

  • LLM​(大規模言語モデル)を​「騙す攻撃」と​「守る​防御」の​両面を​研究。
  • 主要トピック:プロンプトインジェクション​(命令文への​悪意ある​指示の​混入)、​ジェイルブレーク​(安全機構の​回避)、​出力の​安全評価。
  • プロンプトインジェクション=AIへの​指示文に​不正コマンドを​紛れ込ませ、​意図した​制限を​突破する​攻撃手法。

ユーザブルセキュリティ・​解釈​可能性

  • ユーザブルセキュリティ​(​安全性と​使いやすさを​両立させる​分​野)​:認証設計・​同意/プライバシーUX・​人間認知の​特性。
  • 具体例:VRアバターの​表情を​用いた​本人認証、​LLMエージェントが​ブラウザを​操作する​際の​同意・プライバシー設計。
  • 解釈​可能性​(Interpretability)​=​「なぜそう​答えたか」を​AI自身が​人間に​説明できる​度合い。​トークン埋め込みや​内在次元数の​分析が​対象。

計算認知科学・記号創発

  • 個人・集団の​「心の​働き」を​計算論的に​モデル化する​分野。
  • グラフニューラルネットワーク​(GNN=ノードと​エッジの​網構造を​学習する​AI)で​感情伝染を​再現、​社会規範の​発生過程を​シミュレーション。
  • 記号創発=集団内の​やり​取りから​意味や​言語が​自然に​生まれる​現象。

追いつく​順序

  • ①基礎体力:競技プログラミング​(AtCoder)​+CTF​(セキュリティ競技。​脆弱な​システムに​侵入して​「フラグ」を​奪う​実践形式)。
  • ②AI基礎:機械学習・NLP・LLMの​仕組み。
  • ③〜⑤:LLMセキュリティ → ユーザブルセキュリティ/認知科学 → 応用研究​(GNN・VR認証等)。
  • 「技術の​攻撃側」と​「人間理解」の​両輪を​片方に​偏らず​同時に​育てるのが​片岩氏の​強みの​源泉。
片岩氏の​全分​野は​「安全・​信頼ある​人×AI共存」と​いう​一本の​軸で​つながる。​基礎体力と​AI基礎を​先に​固め、​そこから​人間側へ​登るのが​正しい​順序。

出典・参考 / Sources

148 notestil