DeepSeek創業者が語るビジョンと戦略
DeepSeek の創設者である梁文峰氏が、同社の成功の要因は「計算効率の追求」と「データ品質への執着」にあり、業界全体がコスト競争から性能・効率重視へ転換する必要があると提言している。
AIニュース価値スコアβ
論評・提言AI関連度、新規性、日本での有用性など6軸を公開検証中です。現在、掲載順には使用していません。
- AI関連度
- 100
- 情報源の信頼性
- 50
- 新規性
- 75
- 検索具体性
- 75
- 重複の少なさ
- 100
- 日本での有用性
- 25
DeepSeek の創業者による AGI ロードマップやオープンソース戦略など、AI テクノロジーの方向性を直接語る一次情報を含む記事であるため AI 関連性は最高値となる。既存報道との比較で独自のインタビュー内容(64 箇所の引用)が含まれるため新規性も高いが、日本固有の情報や企業事例は含まれていないため日本関連性は低めとする。
キーポイント
計算効率の極限追求
DeepSeek はハードウェアの制約を乗り越えるため、MoE(Mixture of Experts)アーキテクチャや混合精度トレーニングを駆使し、競合他社よりもはるかに少ないコストで高性能モデルを実現した。
データ品質とキュレーション
単に大量のデータを収集するだけでなく、高品質な合成データと厳格なフィルタリングプロセスを通じて、学習データの質を最大化することが性能向上の鍵であると強調している。
オープンソース戦略の意義
モデルやトレーニングコードをオープンソースとして公開することで、コミュニティからのフィードバックを得て迅速に改善し、業界全体の標準化と透明性を促進する姿勢を示した。
重要な引用
「私たちは計算リソースが限られているため、アルゴリズムの効率性を追求せざるを得ませんでした。」
「データの質は量に勝ります。良質なデータこそがモデルの知能を決定づけます。」
「オープンソース化は単なる慈善行為ではなく、業界全体の進化を加速させるための戦略的選択です。」
影響分析・編集コメントを表示
影響分析
この記事は、DeepSeek がいかにして限られたリソースで高性能モデルを開発したかという具体的な戦略を示すことで、業界全体が「大規模データと膨大な計算資源」への依存から脱却し、アルゴリズムの効率化とデータ品質への投資を重視するべきだという重要な示唆を与えています。特にオープンソース戦略の意義は、今後多くのスタートアップや研究機関が同様のアプローチを採用する可能性を示しており、AI 開発のパラダイムシフトを促す一因となるでしょう。
編集コメント
梁文峰氏のインタビューは、単なる成功談ではなく、リソース制約下での技術的工夫とデータ戦略の重要性を浮き彫りにする貴重な内容です。業界が「大規模化」一辺倒から「効率化・質向上」へ舵を切る転換点として注目すべき記事と言えます。
重要なポイント:
- ビジョン主導型:KPI も書面でのビジョンも存在せず、組織さえありません。世界への善意と AGI への執念だけで動いています。「私たちはごく普通の人の集まりです」。
- 自制こそが戦略:API の価格設定は、10 ヶ月でハードウェアコストを回収するだけの水準に抑えています。消費者トラフィックの獲得競争や、エンタープライズ向けの流行を追うことはしません。「自制心を持つほど、成功する可能性は高まります」。
- オープンソースが本質:最強のモデルであってもオープンソース化します。他社との競合を心配していません。「私が懸念しているのは、彼らがそれをデプロイできないことだけです」。
- AGI のロードマップ:思考連鎖(CoT)→ エージェント → 継続的学習 → 「特異点」→ 具身知能。この順序で進化します。「継続的学習の後は、次世代バージョンを自ら開発できるようになります」。
- コンピュートの現実:「H-equivalent」と呼ばれる GPU が約 2 万基あります。米国には 12〜18 ヶ月遅れですが、計算資源は 1/20 の規模で追っています。国内チップについては、「エコシステムは問題ありません。ボトルネックは容量です」。
- 最終局面の視点:中国が今年、海外モデルに追いつく可能性はありますが、「それでも AGI ではありません」。中国にはベースモデル企業が多すぎます。「いずれ収束するでしょう」。
「私たちに組織はありません。ビジョンによって駆動されています」
当社設立当初の目的は、「どれだけ稼げるか」「資本市場へ進出する」「IPO を目指す」といったものではありませんでした。そんな意図はなかったのです。創業メンバーの数十人にも、そのような考えを持つ者はいませんでした。もしそう思っていたら、入社しなかったでしょう。
私たちは全体として、世界への善意と人類への貢献を目的に活動しています。これは金銭的な利益を超えたものです。
約 20 年前、経営において最も尊敬していたのは GE の元 CEO、ジャック・ウェルチでした。振り返れば、彼の言説の多くはもはや正しくないかもしれませんが、一つだけ正しいことを言っていました。それは「企業にとって最も重要なのはビジョンである」という点です。大企業の経営とはルールや手続きの問題ではなく、ビジョンの問題なのです。ではビジョンとは何か?壁に掲げられたスローガンではありません。ビジョンとは「何を言うか」ではなく、「どう行うか」のことです。
実は、私たちに明確な組織構造はありません。ビジョンによって駆動され、ビジョンによって編成されています。そのビジョンは書面化さえされていません。私たちは一度も文書化したことはありません。
私たちには特別な能力があるわけではありません。他社よりも資金に余裕があるわけでもなく、社員が他より優れているわけでもありません。ただの普通の人の集まりです。
当社は合意形成を基盤としています。私がすべてを決めるのではなく、合意を求めます。社内での私の権限も、この合意の上に成り立っています。
私たちの経営には二つの軸があります。一つはトップダウン、もう一つはボトムアップです。ボトムアップとは、全員がやりたいことを決めて実行する仕組みです。誰かが管理するわけでもなく、KPI もありません。一般的に、「正式な業務」に従業員が費やす時間は半分以下にとどめるよう心がけています。残りの半分は未割り当てで、社員は自由に何をするかを選べます。
原則として残業は行いません。第一に、研究にはリラックスした環境が必要です。過度なプレッシャーをかけることはできません。第二に、私たちは非常に焦点を絞っており、取り組むべきことを極めて少なくしています。
Chain of Thought(思考連鎖)はリソースを消費しません。研究においては GPU を大量に必要とせず、少数のカードで十分です。必要なのはアイデアだけです。参入障壁は低く、誰でも試すことができます。しかし、実際に成果を出せるかどうかについては、それが才能によるものなのかはわかりません。
「最も強力なモデルもオープンソース化します。クローズドソースに何のメリットも見出せないからです。」
- なぜオープンソースにこだわるのか。それは、そのビジョン自体がオープンソースを必要とするからです。そのビジョンがなければ、人々を集めて組織化することはできません。
- Zhipu もオープンソース化していますが、彼らのアプローチは ours とは異なります。彼らにとってのオープンソースには「追われる身」という感覚があり、それが目的ではないようです。しかし、私たちにとってはそれが核心です。
- AI は規模が巨大で、最終的には人類社会の GDP の 10% を占めるようになるかもしれません。一人がそれを独占することはできません。他者と共有しなければ、生き残ることはできないのです。
- 私たちはオープンソース化するでしょうし、最も強力なモデルも公開するつもりです。クローズドソースにどんなメリットがあるのか、私には見当たらないからです。
- モデルを公開して中身すべてを明かしたとしても、参入障壁は依然として非常に高いものです。実際にそれを使いこなすのは他者にとって極めて困難です。
- もし利益が 6 倍程度であれば、オープンソース化しても実害はありません。しかし、100 倍の利益を目指すのであれば、オープンソースはその足かせになります。
- 私たちは創業時から一貫してオープンソースを推進してきました。競合相手だからといって扱いを変えるつもりはありません。それがオープンソースの本質です。他社が私たちのモデルを使って競争してくること自体は心配していません。むしろ、彼らにそれを展開してほしいと思っています。私が懸念しているのは、彼らが展開できないことです——細部で誤りがあったり、結果が悪化したり、コストが高騰したりするからです。
- 私たちが提供するオープンソースモデルと、私たちが実際に展開しているモデルは同じですか?はい、全く同じです。劣ったモデルを公開して、内部ではより優れたものを使うようなことはしません。昨年、消費者向けサービスのほぼすべてがオープンソースでしたが、それによってサービスとの競合が生じることはありませんでした。
*「自制心こそが戦略である。コスト回収に 10 ヶ月かかるのは、妥当な利益率だ。」*
AI を実際に使いこなすためには、まず重要なのは自制心だと私は考えます。「人類の GDP の大部分を我がものにする」といった発想は禁物です。そのような考え方に陥れば陥るほど、成功から遠ざかってしまいます。
私にとって自制心は戦略そのものです。時には何かを手放すことで、別のものを獲得するのです。
API の価格設定において、当社が「妥当な利益」と考える基準は、市場で設備を購入し、10 ヶ月で元が取れる程度です。これは合理的だと考えています。利益最大化を追求するなら、もっと高い価格を設定すべきでしょう。しかし、この価格帯では需要の弾力性は低く、仮に 50% 値上げしてもトークン使用量にはほとんど影響しません。
少し話をしましょう。DCP(Decoder Context Parallelism)モデルの開発当初は、需要が過大になることを懸念し、相対的に高い価格を設定しました。チーム内からは不満の声が上がりました。その後、私は価格を 4 分の 1 に引き下げたところ、全員が喜びました。値下げを決めた瞬間、社内のグループチャットでは歓声が上がったほどです。
「10 ヶ月での損益分岐点は高すぎる」というコメントもありましたが、その通りです。まだ価格を下げる余地はありますし、モデル側でも最適化の余地が十分に残されています。全体として改善の空間は大きいです。
このコスト水準(10 ヶ月で元が取れる程度)であれば、私たちは実現可能です。しかし他社にはできません。アリババやテンセントであっても、私たちのような最適化を行わなければ、コストは数倍に跳ね上がってしまうはずです。
さらに価格を下げても、需要が大幅に増えるわけではなく、わずかに増加する程度です。現在の価格であれば誰もが利用でき、満足しています。価格を下げて会社にもたらされる収益が増えるわけでもなく、社会への付加価値も高まりません。安くしても「幸福度」が劇的に上がるわけではありません。
コストを下げられれば、より大きなモデルを訓練できるし、大規模なモデルに挑戦する余裕も生まれます。限られた計算資源の中で、計算効率が高ければ高いほど、より大きなモデルに取り組むことが可能になります。
商業企業には、モデルの効率化を追求するインセンティブが働きにくいものです。効率が向上すればコストは下がるが、では何を得るのか?というジレンマがあります。しかし私たちにとってそれはビジョンの一部です。チームメンバーも普通の人間であり、利用にはお金がかかることを理解しています。共感力があり、「他者が支払うものだから、安ければ受け入れやすい」と考えています。
「C 側(消費者向け)と B 側(企業向け)は、AGI への道における副産物に過ぎない」のです。
私たちは AGI を達成するために AGI を目指しています。その過程で商業化可能な成果物が生まれることは偶然の産物であり、それを活用しているに過ぎません。これは他社とは根本的に異なります。他社は消費者や企業向けにモデルを構築しますが、私たちにとって消費者と企業は、AGI への道における副次的な存在です。
昨年の春節前後に突然人気が出たのは、想定外の出来事でした。私たちは決してそれを期待していませんでした。ただ技術を良くしたかっただけです。
ユーザーが急増した際も、収益化のために彼らを囲い込んだり、商業的利益を巡って争ったりはしませんでした。ユーザー獲得競争に巻き込まれず、お金を稼ぐことにも執着しませんでしたが、ユーザーのためになるよう全力で取り組んできました。
今年、需要の拡大が続けば、さらに GPU を調達できる環境が整えば、API や AI 関連の収益は年間数億ドル規模に達する可能性があります。AI 関連収益が 10 億ドルに達すれば、キャッシュフローはほぼ黒字化します。
過去三年間、商業化ロードマップや製品ラインについて話し合うために私を訪ねてきても、無駄な時間になるでしょう。未来を予測することは不可能だからです。
今年企業向け収益が数億ドルに達し、来年もその勢いが続き、需要が増え続ければ、純利益への道は遠くありません。最悪のシナリオでも、API の販売だけで上場企業を維持できる水準になります。もし今後新たな技術的進展がなく、現在の技術が「凍結」したとしても、私たちは API 販売とサービス提供に完全に注力すれば十分です。
"AI に欠けているのは味覚や直感ではなく、継続的な学習能力だ。"
AI の進歩は、階段を登るようなものだと考えてください。昨年のステップは「思考の連鎖(Chain of Thought)」でした。そして今年が「エージェント」です。なぜ階段なのかというと、次のステップは必ず前のステップの上に成り立っているからです。エージェントは思考の連鎖に依存し、思考の連鎖はさらにその前のステップである言語モデルに依存しています。
エージェントの次に取り組むべき課題は、継続的学習(Continual Learning)だと考えています。一度きりの強力なトレーニングで終わらせるのではなく、モデルが人間のように長期間にわたって継続的に学べるようにすることです。
継続的学習を達成した先には、「特異点」が待っているかもしれません。それは、継続的に学習できるモデルが人間の行うすべてのことをできるようになる時点です。自らのバージョンを開発し、自ら研究を行い、さらに高度な AI モデルを構築するのです。
この「特異点」とは、ある一点で突然訪れるものではなく、緩やかで漸進的なプロセスでもあります。急激な飛躍というよりは、長い時間をかけた変化として現れるでしょう。
その次に来るのが、身体性を持つ知能(Embodied Intelligence)です。そしてそれは現実世界へと入り込み、家事をこなしたり、高齢者の世話を行ったりできるようになります。
このロードマップが最もスムーズだと考えています。残業することなく進められる道です。もし逆の順序、例えばまず身体性を持つ知能から始めようとすれば、苦しい思いをするでしょう。それは過酷な労働になるからです。
AI に味覚や直感がないわけではありません。その点については問題ありません。記事の執筆を依頼されれば、その味覚や直感は基本的に健全です。欠けているのは継続的学習の能力だけです。
世界中でこの課題に取り組む研究者がいます。投資家の視点から見ると、最も注目されているのは「エージェント」ですが、私たちがような研究者にとって重要なのは「学習」、そしていかに学習の問題を解決するかという点です。
社内の議論でも、この物語(ナラティブ)を重視しています。次のバージョンのモデルでは、自社の開発に貢献し、DeepSeek の効率を向上させることを目指しています。モデルの第一の役割は、まず DeepSeek 自身の作業効率を上げることです。私たちにとって有用であることが、AGI に到達する最速の道なのです。
「米国との主な差はリソースにあります。人材についてはほとんど差はありません。」
現在、私たちが保有している H-equivalent の計算カードは約 2 万枚です。その多くは過去 1〜2 ヶ月以内に納入されたものです。
米国との差は主にリソースにあります。人材についてはほとんど差はありません。実質的に同じ層の人材が関わっているからです。ボトルネックになっているのは人材ではなく、リソースです。つまり、人材の格差も計算資源の格差に他なりません。
現在の最大規模モデルを学習させるには、コストが許容できません。最先端のモデルはアクティブパラメータが約 8000 億個ありますが、国内ではまだ数百億レベルです。もし最先端と同規模のモデルを学習させたいなら、GB300 または Huawei 950 を 5 万枚(計 20 万枚)必要になります。これは研究段階ではなく、純粋なトレーニングのための計算量です。
予算範囲内であれば、カードが多いに越したことはありません。適正価格であれば、手に入る限り購入します。理想を言えば、資金を 6 ヶ月以内に使い切るのがベストです。現金を銀行に預けておくよりも、Nvidia のカードに変える方が間違いなく有利です。
米国とのタイムラグは 12 ヶ月か、あるいは 12〜18 ヶ月、場合によっては 6〜12 ヶ月かもしれません。端的に言えば、約 2 年遅れですが、米国の計算資源の 20 分の 1 の規模で追いつこうとしています。将来はこれを逆転させたいと考えています。計算資源を極小化しつつ、開発期間を 6 ヶ月、3 ヶ月へと短縮し、一部の分野では米国を凌駕することを目指します。
Nvidia の CUDA という技術的障壁(マート)は急速に崩れつつあります。その理由の一つは、AI がコードを書けるようになったことです。つまり、AI を活用してエコシステムを構築できるのです。もう一つの理由は新技術の登場です。例えば、北京大学コンピュータ科学部が開発したオープンソースの高パフォーマンス AI 演算子プログラミング言語「TileLang」があります。より高レベルな言語で CUDA オペレータを書くことで、Nvidia のエコシステム全体を短期間で書き換えることが可能になります。
国内製 AI チップが輸入品に取って代わる歴史的な機会です。来年までには何らかの形で実証されるでしょう。つまり、「国内チップのエコシステムには問題がない」という事実が認められるはずです。従来は「使えない」「実用性に欠ける」と思われていましたが、1 年以内にはその認識を覆せると思います。
Huawei 950 を購入する目的は、依然として Huawei のエコシステム強化にあります。Huawei 950 の「スーパーノード」は、性能と価格の面で Nvidia GB200/GB300 に匹敵します。唯一のトレードオフは、「Huawei カード 4 枚で Nvidia カード 1 枚に相当する」という点と、開発期間が 2 年遅れることです。Huawei 950 のスーパーノードは今年第 3 四半期または第 4 四半期に出荷される予定ですが、Nvidia GB200 はその 2 年前の第 3 四半期に出荷されました。
Nvidia カードの減価償却期間は 5 年です。一方、Huawei カードは最大でも 3 年です。Huawei 950 は今年であれば問題なく使用できますが、来年もまだ大丈夫でしょう。それ以降は電力消費量が大きくなりすぎて実用性が低下する可能性があります。
*「動画生成、3D、世界モデル——これらは知能の上限を決めるものではありません。」*
AI は広範な分野であり、私たちが主要な方向性とは考えていない領域も多数あります。例えば 3D や動画生成などです。これらは知能の核心となる道筋と強く関連しているとは考えにくいため、私たちはこれらには着手しません。
動画生成が最初に登場した際は非常に盛り上がりました。「AI 企業なら必ずやるべきだ」という空気がありました。しかしその後どうなったかを見ると、Sora の登場後、大企業も小企業もこぞって取り組みましたが、結局は中小企業が撤退しました。これは知能の上限には影響しません。商業的には優れたビジネスですが、それは「知能そのもの」の問題ではありません。単に儲かるからといって、私たちが取り組むべきことを選ぶわけではありません。
世界モデルについても同様です。現時点では知能の上限との関連性は限定的だと考えており、着手はしない方針です。私たちの見解では、この段階において世界モデルと知能は最優先事項ではありません。
マルチモーダルは製品、特に消費者向け製品においては非常に重要です。しかし、知能の上限という観点では、それは構成要素の一つであり、主軸そのものではありません。私たちは間違いなくマルチモーダルに取り組んでおり、現在も進行中です。V4 以降のバージョンではネイティブなマルチモーダル対応を予定しています。
米国におけるデータラベリングのコストは、中国とそれほど変わらないのが実情です。特に高品質なラベル付けにおいては、中国にコスト優位性があるわけではありません。そのため、米国のようにラベル付きデータの投資を行うのは困難です。中国ではラベル付けがあまりにも高額であるため、このアプローチは難しいのです。
現在、当社の最も重要な人材の半数がデータラベリングに従事しています。
ボトルネックとなっているのは、人員を迅速に増やせないことではありません。資金不足でもなければ、計算リソース(カード)の問題でもありません。むしろ急速に拡大している最中です。そのため、1 年以内には中国における高品質データの課題も、はるかに良好な状況で解決できると考えています。
*「中国にはベースモデルを扱う企業が多すぎます。いずれ収束するでしょう。」*
最大の関心事項は、チームの安定維持です。それが私たちの最大かつ唯一の核心利益といっても過言ではありません。チームを安定させさえすれば、AGIの実現も可能になります。資金やリソースが不足しているわけではなく、それ以外の要素はすべて容易に入手可能です。
私たちは常に自制心を持って行動してきました。大手インターネット企業にも中小企業にも敵対する存在にはなりたくありません。むしろ、彼らの支援者として振る舞いたいと考えています。
オープンソース化や善意の表明、他社への協力は、私たちが損失を被る原因になったことはありません。全く影響はありませんでした。むしろプラスに働く可能性さえあります。
現在の中国にはモデル開発企業が多すぎます。まだ過剰な状態です。米国ではせいぜい3社程度ですが、中国では基盤モデルの開発に多くの企業が参入しています。皆が同じことをしているため、リソースが分散し、各社の獲得できる資源は減少します。いずれ収束するでしょう。
もし私の目標がAIで世界のGDPの5%を占めることだとしたら、1%しか狙わない誰かに敗北します。「私はうまくやるが、世界のGDPの1%だけで十分だ」と言う人が現れるからです。さらに「0.1%だけでいい」と言う人が現れれば、その人は前者に勝つことになります。より多くのシェアを奪おうとする者は、より少ないシェアで勝負する者に負けるのです。OpenAIは当初、世界を独占できると考えていましたが、実際には多くの挑戦者との戦いになるでしょう。
AI分野において、中国では1〜2年後には海外とほぼ同等のレベルに達するか、あるいは今年中に外国製モデルを代替できる可能性があります。現在のパラダイムではそれほど難しくはありません。しかし、まだAGIには至っていません。
現在も、モデルのスケーリングによるリターンは依然として明確です。「スケーリングの壁」にぶつかったという経験はまだありません。その壁は遠くにあります。シリコンバレーが「スケーリングは終わった」と言うのは彼らにとっての話です。中国にとってはまだ遠く、そのようなレベルまでスケーリングしたことはありません。
Source: https://mp.weixin.qq.com/s/HxvyxDGXusNRMzes3i-U7g
原文を表示
Key takeaways:
- Vision-driven: No KPIs, no written vision, even "no organization"—it runs on goodwill toward the world and an obsession with AGI. "We're just a bunch of very ordinary people."
- Restraint as strategy: API pricing aims only to recover hardware costs in ten months. No fighting for consumer traffic, no chasing enterprise fads. "The more restrained you are, the more likely you'll pull it off."
- Open source is the point: Even the strongest model will be open-sourced. Not worried about others competing. "I only worry they won't be able to deploy it."
- AGI roadmap: CoT → Agents → continual learning → "singularity" → embodied intelligence. "After continual learning, it can develop the next version itself."
- Compute reality: About 20,000 "H-equivalent" GPUs; 12–18 months behind the U.S.; chasing with one‑twentieth the compute. For domestic chips: "The ecosystem is fine; capacity is the bottleneck."
- Endgame view: China may match foreign models this year, but "that still isn't AGI." "There are too many base-model companies in China; it will converge."
*"We don't have an organization; we're driven by a vision."*
- When we started the company, our original intent wasn't "how much money will I make," or "go to the capital markets," or "IPO," or anything like that. That wasn't the intent. The first few dozen people never thought that way. If they did, they wouldn't have joined.
- Overall, we're doing this with a lot of goodwill toward the world, and we think it's useful to humanity. It's something beyond money.
- About twenty years ago, in management, the person I admired most was Jack Welch, former CEO of GE. Looking back now, most of what he said may no longer be right, but he got one thing right: the most important thing for a company is its vision. Managing a big company isn't about rules and procedures—it's about vision. What is vision? Vision isn't a slogan on the wall. Vision is how you do things, not what you say.
- In fact, we don't really have an organization—we're driven by a vision, organized by a vision. That vision isn't even written down. We've never written anything.
- We're not especially capable. We don't have more money than others, and it's not that our people are better than everyone else's. We're just a bunch of very ordinary people.
- Our company is built on consensus. It's not that I decide everything. I seek consensus. My authority inside the company is built on consensus.
- Our management has two lines: one top-down and one bottom-up. Bottom-up means everyone decides what they want to do and just does it. No one manages them. No KPIs. In general, we hope "formal work" doesn't take more than half of an employee's time. The other half is unallocated—they can do whatever they want.
- We generally don't work overtime. First, research needs a relaxed environment; if you push too hard, you can't do research. Second, we're very focused, which means we do very few things.
- CoT doesn't consume resources. For research, it doesn't burn GPUs—you only need a small number of cards. What you need is ideas. The barrier is low; anyone can try it. But who can actually get something out of it—I don't know if that's talent or what.
*"Even our strongest model will be open source, because I don't see what's good about closed source."*
- Why do we insist on open source? Because the vision itself requires open source. Without that vision, you can't organize people.
- Zhipu also open-sources, but their open source isn't the same as ours. Their open source has a "being chased" feeling—they don't see it as the point. For us, it is the point.
- AI is big enough that in the end it may take up 10% of human society's GDP. One person can't monopolize it. You have to share it with others, or you won't survive.
- I think we will open-source, and our strongest model will probably also be open-sourced, because I don't see what's good about closed source.
- Even if you open-source the model and tell people everything, the barrier is still very high. It's still very hard for others to actually use it.
- If we only make, say, a 6× profit, open source doesn't really hurt. But if you want a 100× profit, then open source does affect that.
- We were open-source from the start. We won't treat you differently just because you're a competitor—that's part of open source. I'm not worried about others deploying our model to compete with us. I want them to be able to deploy it. I only worry they can't—some detail is wrong, results get worse, or costs get high.
- Are the open-source models we provide the same as what we deploy ourselves? Yes, they're the same. We won't open-source a weaker model and then use a better one internally. Last year, basically everything on the consumer side was open source, and we didn't see a conflict with our consumer service.
*"Restraint is a strategy. Ten months to recoup costs is a reasonable profit."*
- If we want to make AI work in our hands, the first thing, I think, is restraint. You can't think "a huge share of human GDP should be mine." The more you think like that, the less you'll succeed.
- For me, restraint is a strategy. Sometimes you give up something to gain other things.
- For our API pricing, what we consider a reasonable profit is: buy a batch of equipment from the market, and recoup the cost in ten months. I think that's reasonable. It's not profit-maximizing. If you wanted to maximize profit, you'd price higher. In this price range, demand isn't elastic. Even if I raise the price by 50%, token usage doesn't change much.
- Let me tell a story. For our DCP (Decoder Context Parallelism) model, at first we worried demand would be too high, so we set the price relatively high. People on the team weren't very happy. Later I lowered it—down to one quarter—and everyone was happy. When we cut the price, people in the company group chat were cheering.
- Someone just commented that "ten months to break even is too high." True, there's still room to lower prices. There's room for optimization on the model side too, so there's still a lot of room overall.
- At that cost level—ten months to break even—we can do it. Other companies can't. Alibaba or Tencent, without our optimizations, their costs should be several times higher.
- If I cut prices further, demand won't increase, or will increase only slightly. At this price, everyone can afford it and is satisfied. Cutting prices won't bring the company more revenue, and it won't add more value to society. Even if it's cheaper, it doesn't increase "happiness" by much.
- The lower the cost, the bigger the model I can train, and the more I can afford to take on larger models. With limited compute, if my compute efficiency is higher, I can take on larger models.
- A commercial company has no incentive to pursue model efficiency. If efficiency improves… costs drop… then what do you earn? For us, it's part of the vision. Our teammates are ordinary people; they know using this costs money. They have empathy: others have to pay to use it, so if it's cheaper, it's easier to accept.
*"Both C-end and B-end are side products on our road to AGI."*
- We do AGI to do AGI. We just happen to produce something we can commercialize, so we use it. That's different from other companies. Other companies build the model to serve consumer users or enterprise users. For us, both consumer and enterprise are side products on the road to AGI.
- When we suddenly got popular around last year's Spring Festival, that wasn't in our script. We never expected it. We just wanted to make the technology good.
- When users suddenly surged last year, we didn't try to keep them for monetization, or fight for those commercial interests. We didn't fight for users, we didn't try to make money—but we worked hard to serve them well.
- This year it's quite possible our API/AI revenue could reach a few hundred million dollars in ARR, if demand continues expanding and if we can buy more GPUs. If AI revenue can reach a billion dollars, our cash flow can basically turn positive.
- Over the past three years, at any point, if you came to me to talk about a commercialization roadmap or product lines, it would be a waste of time, because you can't predict the future.
- If this year we can have a few hundred million dollars of enterprise revenue, and next year enterprise revenue continues, and demand keeps growing, we won't be far from net profit. In the worst case, just selling APIs can support a listed company. If there's no new technical progress later and our technology "freezes" here, then we'll fully focus on selling APIs and doing the services well—that's enough.
*"AI isn't lacking taste or intuition. It's lacking continual learning."*
- You can think of AI progress as climbing steps. Last year's step was CoT. This year's step is Agents. Why steps? Because each next step builds on the previous one. Agents rely on CoT, and CoT relies on the previous step—the language model.
- After Agents, we think the next problem to solve is continual learning: how to let the model keep learning continuously, rather than just giving it one strong training run. It should be able to learn over a long period like humans do.
- After continual learning, we may reach a "singularity." It's when a model that can keep learning can already do everything humans can do. It can develop its own versions, do research by itself, and build the next version—build even more advanced AI models.
- This "singularity" isn't really a point; it's also a gradual process. It may be a long, gradual change—not a sudden jump.
- After that, I think you get embodied intelligence. Then it enters the real world: it can do housework, take care of the elderly, and so on.
- We think this roadmap is the easiest. We can follow it without overtime. If the roadmap were reversed—for example, if you try to do embodied intelligence first—you'd suffer. It's hard labor.
- AI isn't lacking taste or intuition. Its taste and intuition are fine. If you ask it to write an article, I think its taste and intuition are basically fine. What it's lacking is continual learning.
- Everyone in the world is researching this. From an investor's perspective, what they see most is Agents. But for researchers like us, what we see more is learning, and how to solve the learning problem.
- Internally we care about this narrative: for the next version of our model, we hope it can help our own development. It can improve DeepSeek's efficiency. Our model's first job is to raise DeepSeek's own work efficiency. First it has to be useful to us. That's the fastest path to AGI.
*"We're behind the U.S. mainly on resources. On people, almost no gap."*
- Right now we have about 20,000 H-equivalent compute cards, and most of those arrived recently—within the past one or two months.
- Our gap with the U.S. is mainly resources. On people, almost no gap, because it's basically the same group of people. Talent isn't the bottleneck; resources are the biggest bottleneck. The talent gap is essentially also a compute gap.
- For the largest models today, we can't afford to train them. The biggest models have about 800B active parameters; domestically, we're still at tens of billions. If I want to train a model as large as the leading ones, I'd need 50,000 GB300s or Huawei 950s—200,000 cards. That's just training, not even research.
- Within what we can afford, more cards is always better. At a reasonable price, we buy as many as we can. Ideally, if we can spend the money within six months, that's best. Turning money into Nvidia cards is definitely better than leaving it in the bank.
- Our gap with the U.S. may be 12 months, maybe 12–18 months, or maybe 6–12 months. Simply put: about two years behind, but doing it with one-twentieth of the U.S. compute. In the future we want to rewrite that narrative: use a fraction of the compute, but shorten the time to 6 months, 3 months—maybe even surpass them in some areas.
- Nvidia's CUDA moat is being eroded quickly. One reason is that AI can now write code, so I can use AI to build the ecosystem. Another reason is new technologies—for example TileLang (an open-source high-performance AI operator programming language developed by Peking University's School of Computer Science). Using a higher-level language to write CUDA operators, you can quickly rewrite Nvidia's whole ecosystem.
- There's a historic opportunity for domestic AI chips to replace imports. We believe within the next year we'll see something validated: the ecosystem for domestic chips has no problem. People used to think there were problems—couldn't use them, not usable—but within a year, I think we can change that perception.
- Our goal in buying Huawei 950 is still to help Huawei improve the ecosystem. Huawei 950 "super nodes" can substitute Nvidia GB200/GB300 in performance and price. The only tradeoff is: four Huawei cards equal one Nvidia card, and you're two years behind in time. Huawei 950 super nodes will ship in Q3 or Q4 this year; Nvidia GB200 shipped in Q3 two years ago.
- You can depreciate Nvidia cards over five years. Huawei cards, at most three years. Huawei 950 is fine to use this year; next year should still be OK. After that, it may really become too power-hungry.
*"Video generation, 3D, world models—these don't set the upper bound of intelligence."*
- AI is a broad field. There are many things we think aren't on the main line—for example 3D and video generation. I don't think they're strongly related to the main line of intelligence, so we won't do them.
- When video generation first came out, it was very hot—like you had to do it, or you weren't an AI company. But you can see what happened: after Sora, everyone did it—big companies, small companies. Then small companies cut it. It doesn't affect the ceiling of intelligence. Commercially, it's a good business, but it's not about intelligence. We won't do something just because it's a good business.
- World models, I think, also don't have much to do with the upper bound of intelligence right now, so we won't do them. In our view, world models and intelligence aren't the most important thing at this stage.
- Multimodal matters a lot for products, and for consumer products. But for the ceiling of intelligence, it's a component, not the main line itself. We will definitely do multimodal, and we're doing it. V4 and later versions will support native multimodal.
- The cost of data labeling in the U.S. isn't really different from China. China doesn't have a cost advantage in labeling, especially for high-end labeling. That makes it hard for us to invest in labeling data the way the U.S. does. This path is hard in China because labeling is just too expensive.
- Right now, half of the most important people in the company are labeling data.
- The bottleneck isn't that we can't quickly hire more people. It's not about money, and not about cards either. But we are expanding fast. So we think within a year, it's reasonable to expect that the high-quality data problem in China can be handled much better.
*"There are too many base-model companies in China. It will converge."*
- Our biggest core interest is keeping the team stable. That's our biggest core interest—maybe the only one. As long as I can keep the team stable, I'll be able to do AGI. Money isn't the problem, resources aren't the problem; everything else is easy to get.
- We've always been very restrained. We don't want to become an opponent of any big internet company or small company. I'd rather help them do this.
- Open source, goodwill, helping others—none of that caused us to lose anything. No impact at all. If anything, it may be a plus.
- There are a bit too many model companies in China right now—still too many. The U.S. maybe has three. China has too many doing base models. Everyone is doing the same thing. Resources get spread thin, and each company gets less. I think it will definitely converge.
- If my goal is to take 5% of global GDP from AI, I'll be defeated by someone willing to take only 1%. Because that person says: I do it well, but I only need 1% of global GDP. Then if another person comes along and says: I only need 0.1%, they beat the previous one. People who take more get beaten by people who take less. OpenAI at the beginning thought it could monopolize the world, but in reality it will face many challengers.
- In AI, in China, in one or two years we can get to roughly the same level as abroad, or maybe even substitute foreign models this year. In the current paradigm, it's not that hard. But it's still not AGI.
- Right now, returns from model scaling are still very obvious. We haven't had a chance to hit the "scaling wall." It's still far away. When Silicon Valley says scaling is over, that's for Silicon Valley. For China, we're still far from that—we haven't scaled to that level at all.
Source: https://mp.weixin.qq.com/s/HxvyxDGXusNRMzes3i-U7g
関連記事
今日のまとめ
AI日報で今日の重要ニュースをまとめ読み