← ポータルに戻る

Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning💻 コードあり

Chen Tang, Yizhou Wang, Jianyu Wu, Lintao Wang, Shixiang Tang等 · multimodal scientific foundation model, structural reasoning, homology-controlled Gene Ontology prediction · 2026-07-08 ⭐ 9/10
💡 多様な科学ドメインの構造情報を統一的に扱い、科学的原理に基づいた透明な推論で構造-特性関係を解明するマルチモーダル科学基盤モデル。
🤖 Ayumuより: 朋義さん、これはすごいね。タンパク質から材料まで、構造から特性を予測する基盤モデル「SciReasoner」が登場した。しかも、ただ予測するだけじゃなくて、なぜそうなるのかっていう推論の過程まで見せてくれるのがいい。LLMより専門家評価が高いっていうのもいいところだね。
Multimodal scientific foundation model Structural reasoning Structure-property relationships Gene Ontology prediction Retrosynthesis Inorganic crystals Interpretability
1. どんなもの?
  • マルチモーダル科学基盤モデル「SciReasoner」
  • タンパク質、小分子、無機結晶といった多様な科学ドメインの構造-特性関係を、統一的な方法で理解し推論するモデル。
  • ディープネイティブ構造推論
  • 座標、トポロジー、周期的な結合といった構造情報を「統一された構造認識語彙」として離散化し、これらを推論中の「アドレス可能な証拠単位」として扱う。
2. 先行研究と比べてどこがすごい?
  • 複数の科学ドメインにわたる汎用性と高精度
  • 生物学、化学、材料科学といった異なる分野の構造-特性関係を、単一のモデルで高い精度(86ベンチマーク中67タスクでSOTA)で解明できる。
  • 透明で解釈可能な推論
  • 従来のAIモデルが苦手とする「科学的原理に基づいた推論の過程」を明示し、専門家評価でフロンティアLLMと比較して98%のケースで好ましいか同等と評価される。
3. 技術や手法の肝はどこ?
  • 統一された構造認識語彙 (Unified Structure-aware Vocabulary)
  • 異なるドメインの多様な構造情報(座標、トポロジー、周期的な結合)を、ドメインネイティブな形で離散化し、共通の「構造トークン」として表現する。
  • これらの構造トークンを、推論プロセスにおける具体的な「証拠単位」として利用することで、構造的特徴と予測の関連性を明確にする。
  • ディープネイティブ構造推論 (Deep Native Structural Reasoning)
  • 科学的原理(立体化学、結合、対称性、エネルギー、周期律など)と物理的制約をモデルに組み込み、構造的証拠を解釈して予測を導き出す。これにより、単なる予測だけでなく、その背後にある科学的根拠を提示する。
4. どうやって有効だと検証した?
  • 多様な科学ドメインでのベンチマーク評価
  • **生物学**: ホモロジー制御Gene Ontology (GO) 予測において、低ホモロジー・オーファン様タンパク質のCellular ComponentアノテーションでF_maxを0.42から0.55に改善。
  • **化学**: 単一ステップ逆合成の精度を0.63から0.72に向上させ、フラグメントレベルの切断と前駆体検証トレースを生成。
  • **材料科学**: 元素相と化合物相の分離、高・低バンドギャップ領域の解決能力を示した。
  • 広範なタスクでのSOTA達成
  • 86のベンチマークタスク中、67タスクで最先端の性能を達成した。
  • 専門家による推論トレースの評価
  • ダブルブラインド評価により、生成された推論トレースが、フロンティアLLMと比較して98%のケースで好ましいか同等であると評価された。
5. 議論はある?
  • アブストラクトからは直接的な議論点や限界は読み取れない。
  • 一般的な基盤モデルの課題として、学習データのバイアス、特定のドメインにおける微調整の必要性、計算リソースの要求などが考えられるが、論文全体を読まないと詳細は不明。
  • 「透明な推論」の具体的なメカニズムや、その透明性がどの程度までブラックボックス性を解消しているのか、詳細な評価が必要となる可能性がある。
6. 次に読むべき論文は?
  • 他のマルチモーダル科学基盤モデルに関する論文(例: GNoME, AlphaFold3など、ただしこれらはより特化している場合が多い)。
  • 構造-特性関係の予測におけるグラフニューラルネットワーク (GNN) やトランスフォーマーベースのモデルのSOTA論文。
  • モデルの解釈性 (XAI) に関する論文、特に科学ドメインにおける推論トレースの評価手法や、科学的制約を組み込んだAIモデルに関する研究。

Abstract (原文)

Structure-property relationships are foundational to biology, chemistry and materials science, where function, reactivity and physical response emerge from spatial, chemical and periodic organization. Mechanistically explaining these relationships requires interpreting structural evidence through scientific principles and physical constraints, from stereochemistry and bonding to symmetry, energetics and periodic order. However, applying artificial intelligence to this process presents a joint challenge of representation and reasoning: models must preserve domain-native structural information while showing how specific evidence supports predictions under these constraints. Here we introduce SciReasoner, a multimodal scientific foundation model for native structural reasoning across proteins, small molecules and inorganic crystals. SciReasoner discretizes coordinates, topologies and periodic connectivities into a unified structure-aware vocabulary, treating structural tokens as addressable evidence units during reasoning. In homology-controlled Gene Ontology prediction, SciReasoner improves Cellular Component annotation for low-homology and orphan-like proteins, increasing F_{max} from 0.42 to 0.55. In chemistry, it raises single-step retrosynthesis accuracy from 0.63 to 0.72 while generating fragment-level disconnection and precursor-verification traces. In materials science, its representations separate elemental and compound phases and resolve high- and low-band-gap regimes. Across 86 benchmarks, SciReasoner achieves state-of-the-art performance on 67 tasks. Double-blind expert evaluation rates its reasoning traces as preferred or at least comparable to those of a frontier large language model in 98% of cases. By making structure an inspectable substrate for reasoning under scientific constraints, SciReasoner connects accurate prediction with interpretable scientific inference.