GSK、Relation Therapeutics と AI 創薬で提携し生物データ活用へ
本文の状態
日本語全文を表示中
詳細モードで約7分の本文を読めます。
同じ出来事の情報源
この情報源を基点に整理
AI News
GSK は英国の Relation Therapeutics と AI 創薬協力を拡大し、1億1000万ドル規模で遺伝子変化への細胞反応データ生成を強化する。
AI深層分析を開く2026年8月3日 19:01
AI深層分析
キーポイント
大規模資金と協力体制の確立
GSK は Relation Therapeutics と最大1億1000万ドルの協定を結び、既存の AI 支援創薬研究を拡大する。
生体データ生成と AI モデル開発の統合
Relation の Lab-in-the-Loop プラットフォームを用いて遺伝子変化や薬剤介入に対するヒト細胞の反応データを大規模に生成し、MORGAN などの AI モデル訓練に活用する。
既存疾患研究からの知見の継承
線維症と変形性関節炎を対象とした先行プロジェクトで得た遺伝子および単細胞オミクスデータの分析手法を、今回の拡大協力の基盤としている。
公開データセットの活用課題
CZ CELLxGENE などの公開リポジトリは重要な学習素材だが、実験プロトコルの違いや重複によるデータリークリスクといった技術的課題が残る。
データ規模とモデル性能の非線形関係
単細胞ファウンデーションモデルは、大規模言語モデルとは異なり、トレーニングデータの増加が常に性能向上につながるとは限らない。モデル容量、データセットサイズ、計算リソースのバランスが重要であり、単純なデータ量の拡大では限界に達する傾向がある。
重要な引用
Relation will generate large-scale datasets measuring how human cells respond to genetic changes and drug interventions.
The agreement places biological data generation alongside AI model development.
A 2025 review in Experimental & Molecular Medicine noted that repositories including CZ CELLxGENE, the Human Cell Atlas, and NCBI Gene Expression Omnibus give researchers access to large volumes of single-cell data.
Bigger biological datasets do not guarantee better models
編集コメントを表示
編集コメント
創薬における AI の活用は、アルゴリズムの改良だけでなく、高品質な生体データの確保が鍵となることを示す重要な事例である。特に公開データセットの利用には技術的課題が残るため、企業独自のデータ生成基盤の重要性が改めて浮き彫りになった。
Source Article
元記事を日本語で読む
本文に関係しない購読案内、埋め込み通知、サイト内プロモーションは除いています。
How Relation generates biological data
Relation は、実験室での検証と計算分析を組み合わせる「Lab-in-the-Loop」アプローチを採用しています。同社の研究には組織プロファイリングや単一細胞・空間トランスクリプトミクス、シーケンシング、標的の妥当性確認が含まれ、機械学習は標的の特定、優先順位付け、妥当性確認、実験設計に活用されています。
また、遺伝子の変化が疾患に関連する細胞特性にどう影響するかを測定する撹乱実験も実施しており、その結果は遺伝子データや患者由来の生物学的データとともに分析されます。
生物学的基盤モデルの学習素材として、公開リポジトリは今なお重要な役割を果たしています。ただし、異なる研究で得られた情報を組み合わせる際には技術的な課題が生じる可能性があります。
2025 年に『Experimental & Molecular Medicine』誌に掲載されたレビューによると、CZ CELLxGENE やヒト細胞アトラス(Human Cell Atlas)、NCBI Gene Expression Omnibus などのリポジトリを利用すれば、研究者は大量の単一細胞データにアクセスできます。同レビューでは、CZ CELLxGENE だけでも標準化された細胞が 1 億個以上提供されていると指摘されています。
ただし、サンプリング手法やシーケンシングプロトコル、実験手順、処理パイプラインは研究によって異なります。単一細胞データには技術的なノイズやアーティファクトが含まれることも多く、基盤モデルのトレーニングにおいては、データの慎重な選定・フィルタリング、構成のバランス調整、品質管理が不可欠です。
データセットの重複も課題の一つです。レビューでは、同じまたは類似した細胞が複数の公開リソースにまたがって存在する可能性が指摘されており、これが学習時に不均衡な影響を与えたり、訓練用とテスト用のデータセットが重複することでデータリークリスクを生じさせたりするとされています。
このレビューは、堅牢なシングルセル基盤モデルを構築する際、高品質で冗長性のないデータセットを整備することが、モデルアーキテクチャの設計と同様に重要であると結論付けています。
生物学的データの規模が大きいほど、必ずしも優れたモデルになるとは限らない
今年 6 月に『Nature Methods』に掲載された研究では、2,220 万個の細胞からなるコーパスを用いて、事前学習データの規模と多様性がシングルセル基盤モデルにどう影響するかを検証しました。研究者らは 400 のモデルを訓練し、6,400 の実験を通じて評価を行いました。
その結果、現在のシングルセル基盤モデルは、利用可能なコーパスの一部のみで訓練しても性能が頭打ちになる傾向があることが分かりました。大規模言語モデルとは異なり、評価されたシステムでは、訓練データを継続的に増やすことで結果が一貫して向上するという明確なデータスケーリング法則は見られませんでした。
研究者らは、モデルの容量、データセットの規模、計算リソースは単に増やせばよいのではなく、バランスよく調整する必要があると指摘しています。この研究で示されたのは、小規模または独自データセットが本質的に優れているという事実ではなく、生物学的な訓練データを追加しても、必ずしも性能向上につながるとは限らないという点です。
2025 年に『Genome Biology』に掲載された別の研究では、Geneformer と scGPT という 2 つのシングルセル基盤モデルが、いくつかのゼロショット評価タスクで検証されました。その結果、これらのモデルは単純な手法を一貫して上回ることはなく、バッチ効果に関する課題も明らかになりました。研究者らは、事前学習済みモデルが大きければ自動的に生物学的表現が向上すると安易に考えるべきではないと注意を促しています。
製薬企業は専門的なデータセットを追求する
Relation 社はすでに、同社が独自機能を持つシングルセル骨アトラスとして位置づける「Osteomics」に、自社のデータ生成アプローチを適用済みです。このプロジェクトでは患者由来のサンプルを用い、シングルセル・空間オミクスとイメージング、ゲノム、プロテオーム、臨床表現型データを統合しています。
同社によると、Osteomics は骨粗鬆症における疾患メカニズムの解明、治療ターゲットやバイオマーカーの特定、患者サブグループの分類に活用されています。英国および豪国の病院と研究パートナーが参加する観察研究の一環として進められています。
先月『Nature Genetics』に掲載された研究でも、シングルセル解析、遺伝子データ、機能検証を用いて骨疾患の細胞・遺伝的要因が調査されました。この論文の著者には Relation 社の研究者も複数含まれています。
2025 年に『Nature Biotechnology』が実施した、AI を中核に据えたバイオファーマの取引分析では、専門的なデータセットを提供する企業が、最近の提携から浮上したトレンドの一つとして挙げられました。その他のトレンドには、高額な初期支払い、新たな治療モダリティの登場、そして大手バイオ企業からの参画拡大が含まれています。
この分析は、高品質で疾患特異的なデータセットが、因果推論や生成型の機械学習モデルにとって不可欠な入力要素へと変わりつつあると指摘しています。具体例として、GSK が Ochre Bio と結んだ別件の契約が挙げられました。これはヒトの肝臓における単一細胞データや灌流された臓器データを対象としたライセンス契約で、金額は 3,750 万ドルに上ります。
もう一つの事例では、アストラゼネカと Pathos AI が 2025 年に Tempus と 2 億ドルの契約を結んでいます。この提携のもと、Pathos は 15 万人以上の患者を対象とした非特定化された臨床データ、ゲノムデータ、画像データを用いて、腫瘍学向けの基盤モデルの開発を進めることになっています。
AI を活用した創薬において、十分な高品質データへのアクセスは依然として大きな制約となっています。製薬研究における連合学習(federated learning)を取り上げた『Nature』の研究ハイライトでは、適切なトレーニングデータの入手制限が AI アプリケーションの主要なボトルネックであると指摘されています。同時に、企業が自社の機密情報を共有することにも規制がかかる場合があることも注記されています。
このように、AI とバイオファーマの間で結ばれる契約は、企業がデータや計算能力をどのように確保するかによって様々です。中には AI プラットフォームへのアクセスに焦点を当てたものもあれば、共同開発、データライセンス、あるいは新たな生物学的データセットの構築を柱としたものもあります。
GSK と Relation の提携には、データ生成とモデル開発の両方が含まれています。Relation はこの協力の一部としてヒト細胞由来のデータセットを生成し、それを用いて潜在的な創薬ターゲットを特定するための AI モデルを学習させます。
(写真:CDC 提供)
関連記事:中国における AI がどのように創薬期間を短縮しているか

原文を表示
GSK has entered into a research collaboration with British biotechnology company Relation Therapeutics worth up to $110 million, expanding the companies’ existing work in AI-assisted drug discovery.
Under the agreement, Relation will generate large-scale datasets measuring how human cells respond to genetic changes and drug interventions. The data will be used to train AI models designed to identify potential drug targets, including models within Relation’s MORGAN platform.
The agreement places biological data generation alongside AI model development. Relation’s research approach links computational analysis with experiments that generate new information on human cells.
The collaboration builds on earlier agreements between GSK and Relation focused on fibrotic diseases and osteoarthritis. Those projects involved observational studies designed to create two functional disease datasets for analysis using Relation’s Lab-in-the-Loop platform.
The earlier work combined human genetics, single-cell multi-omics generated from human tissue, functional assays, and machine learning to identify and validate potential disease targets.
How Relation generates biological data
Relation describes its Lab-in-the-Loop approach as a combination of laboratory experimentation and computational analysis. Its work includes tissue profiling, single-cell and spatial transcriptomics, sequencing, and target validation, while machine learning is used for target identification, prioritisation, validation, and experimental design.
The company also conducts perturbation experiments that measure how genetic changes affect cellular characteristics associated with disease. Those results can then be analysed alongside genetic and patient-derived biological data.
Public repositories remain an important source of training material for biological foundation models, although combining information produced across different studies can introduce technical challenges.
A 2025 review in Experimental & Molecular Medicine noted that repositories including CZ CELLxGENE, the Human Cell Atlas, and NCBI Gene Expression Omnibus give researchers access to large volumes of single-cell data. CZ CELLxGENE alone provides access to more than 100 million standardised cells, according to the review.
Sampling methods, sequencing protocols, experimental procedures, and processing pipelines can differ between studies. Single-cell data can also contain technical noise and other artefacts, requiring careful dataset selection, filtering, composition balancing, and quality control during foundation-model training.
Dataset overlap presents another issue. The review noted that the same or similar cells can appear across multiple public resources, potentially giving them disproportionate influence during training and creating data-leakage risks when training and test datasets overlap.
The review found that assembling a high-quality, non-redundant dataset is as important as model architecture when building robust single-cell foundation models.
Bigger biological datasets do not guarantee better models
Research published in Nature Methods in June this year examined how the size and diversity of pretraining data affected single-cell foundation models using a corpus of 22.2 million cells. Researchers trained 400 models and evaluated them across 6,400 experiments.
The study found that current single-cell foundation models tended to reach performance plateaus after training on only a fraction of the available corpus. Unlike large language models, the systems assessed did not display clear data-scaling laws in which continually increasing training data consistently produced better results.
The researchers found that model capacity, dataset size, and computational resources need to be balanced rather than simply increased together. The study did not establish that smaller or proprietary datasets are inherently better, but it found that adding more biological training data did not consistently lead to further performance gains.
A separate study published in Genome Biology in 2025 assessed two single-cell foundation models, Geneformer and scGPT, across several zero-shot evaluation tasks. The models did not consistently outperform simpler approaches, while the researchers also identified challenges involving batch effects and cautioned against assuming that larger pretrained models automatically produce better biological representations.
Pharma companies pursue specialised datasets
Relation has already applied its data-generation approach to Osteomics, which it describes as a proprietary functional single-cell bone atlas. The project uses patient-derived samples and combines single-cell and spatial omics with imaging, genomics, proteomics, and clinical phenotype data.
According to the company, Osteomics is being used to investigate disease biology, therapeutic targets, biomarkers, and patient subgroups in osteoporosis. Hospitals and research partners in the UK and Australia are involved in the observational study.
Research published in Nature Genetics last month also examined the cellular and genetic determinants of skeletal disease using single-cell analysis, genetic data, and functional validation. Several Relation researchers were among the study’s authors.
A 2025 Nature Biotechnology analysis of AI-focused biopharma deals identified specialised dataset providers as one of several trends emerging from recent partnerships. Other trends included larger upfront payments, new therapeutic modalities, and greater participation from larger biotechnology companies.
The analysis said high-quality, disease-specific datasets are becoming an important input for causal and generative machine-learning models. It cited GSK’s separate agreement with Ochre Bio, worth $37.5 million for data licensing involving human liver single-cell and perfused-organ data.
Another example involved AstraZeneca and Pathos AI entering a $200 million agreement with Tempus in 2025. Under the arrangement, Pathos was to develop oncology foundation models using de-identified clinical, genomic, and imaging data covering more than 150,000 patients.
Access to sufficient high-quality data remains a constraint in AI drug discovery. A Nature research highlight on federated learning in pharmaceutical research identified limited access to suitable training data as a major bottleneck for AI applications, while noting that companies can also face restrictions on sharing proprietary information.
AI-biopharma agreements therefore vary in how companies obtain data and computational capabilities. Some centre on access to AI platforms, while others cover joint development, data licensing, or the creation of new biological datasets.
The GSK–Relation agreement includes both data generation and model development. Relation will produce human cellular datasets as part of the collaboration and use them to train AI models for identifying potential drug targets.
(Photo by CDC)
See also: How AI is shortening drug discovery timelines in China

Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.
The post Why biological data matters more in AI drug discovery appeared first on AI News.
AI算出
主要ニュースainew評価標準
本記事は単なる業界動向の紹介ではなく、特定の企業間での大規模提携(1.1 億ドル)と、Lab-in-the-Loop やシングルセル基盤モデルに関する具体的な技術的知見を報じており、AI 創薬分野における実質的な進展を示す。
6つの評価軸を見る
- AI関連度
- 75
- 情報源の信頼性
- 25
- 新規性
- 75
- 調べる価値
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
今日のまとめ
AIデイリーブリーフで今日の重要ニュースをまとめ読み