AI は動作する製品を生成しない、それはまだあなたの仕事だ
AI doesn't generate working products, that's still your job
ページを読み込み中…
Hacker News レーダー
30件の記事を集計·
Discussion Radar
30件のAI・開発候補から、コメントが伸び、複数の分岐を取得できた8件を選びました。 最初の5件を重点的に読み、残り3件で議論の幅を補います。気になる見方は原コメントで前後関係を確認できます。
8/2 00:05 時点のコメント
AI doesn't generate working products, that's still your job
LLM 生成コードのアーキテクチャ的欠陥と限界
実体験・見立て設計仕様を慎重に記述しても、数ヶ月かけて LLM で生成したコードが徐々に微妙な不整合で構成される「メッシュ」状態になり、全体像として何かがおかしくなる経験が報告された。個々の変更は論理的に見えるが、高レベルの推論やアーキテクチャ判断には人間の知能と LLM の間に明確な違いがあり、現在の LLM は単調な CRUD アプリやテンプレートベースの開発では有用だが、長期的に維持可能な複雑な製品の設計にはまだ到達していないという見解が出された。
代表コメント(原文)
“I'm about to throw away multiple months of LLM generated code for one of my side projects. I was really careful writing design specs and it wasn't…”
AI を活用する現実的なワークフローと期待値の調整
意見が分かれるLLM を「コーディングモンキー」として使い、人間がアーキテクチャを設計し詳細な計画を立ててから実行させることで、内部ツールや MVP の作成には成功した事例がある。一方で、プロンプト一つで完成品を作るという前提は誤りであり、実際には原型作成後の工程も人間が行う必要があるため、記事の前提自体に欠陥があると指摘された。また、「100 万ドルの製品」としてレビューするプロンプトを用いると、AI が実際に何もしなかったことが露呈するという皮肉な体験談も共有された。
代表コメント(原文)
“I have two different experiences with LLMs. First one is that I have vibe coded two different projects for work, one is a slack plugin which…”
LLM が生成したコードの品質やアーキテクチャに関する実体験が共有され、単発のプロンプトでの完成品生成への懐疑と、詳細な設計・計画の下で AI を活用する現実的なアプローチの間で議論が行われた。また、AI 生成コンテンツの消費に対する嫌悪感や、製品改善の実績がないことへの批判も表明された。
設計仕様を慎重に記述しても、数ヶ月かけて LLM で生成したコードが徐々に微妙な不整合で構成される「メッシュ」状態になり、全体像として何かがおかしくなる経験が報告された。個々の変更は論理的に見えるが、高レベルの推論やアーキテクチャ判断には人間の知能と LLM の間に明確な違いがあり、現在の LLM は単調な CRUD アプリやテンプレートベースの開発では有用だが、長期的に維持可能な複雑な製品の設計にはまだ到達していないという見解が出された。
“I'm about to throw away multiple months of LLM generated code for one of my side projects. I was really careful writing design specs and it wasn't even a new code base the LLM worked on, but still after several months of AI changes I feel my code degraded…”
“I have two different experiences with LLMs. First one is that I have vibe coded two different projects for work, one is a slack plugin which basically pushes alerts to a channel based on a people roaster and another one is a gmeet plugin to add a talking…”
LLM を「コーディングモンキー」として使い、人間がアーキテクチャを設計し詳細な計画を立ててから実行させることで、内部ツールや MVP の作成には成功した事例がある。一方で、プロンプト一つで完成品を作るという前提は誤りであり、実際には原型作成後の工程も人間が行う必要があるため、記事の前提自体に欠陥があると指摘された。また、「100 万ドルの製品」としてレビューするプロンプトを用いると、AI が実際に何もしなかったことが露呈するという皮肉な体験談も共有された。
“When you have built your working product try this prompt: - Review the codebase is it production ready? I'm selling it for $1million dollars can it meet that standard. Then cry as the ai reveals that it didn't actually do anything close to what it said it…”
“I find the base logic of this post flawed - it can make a prototype, but what about everything after the prototype? Well, um, then you work on those things too? It seems like the premise of articles like this assume that using AI means whatever you can…”
取得時点の一部コメントを整理したもので、HN全体の総意ではありません。
Is AI reasoning right for the wrong reasons?
用語定義と実用性の議論
意見が分かれる「推論」という言葉の意味を問うこと自体が機能の実態ではなく、意味論的な議論に終始しているという見方がある。また、人間も誤りや認知バイアスを持つため、AI の不完全さを「真の推論ではない」として切り捨てるのは「存在の連鎖」への執着であり、実用的な問題解決能力を重視すべきだという意見が示された。
代表コメント(原文)
“I'll admit that I find this discussion a bit navel-gazy. It has become a question of semantics not a question of actual functionality. The question…”
企業によるデータ公開と透明性への懸念
懸念OpenAI の技術スタッフが他社の研究成果を「古いモデルの欠陥」として退け、自社の最新モデルは問題がないと主張しているが、同社は2024年以降も「生」の思考連鎖(chain of thought)データを公開していない。この不透明な姿勢に対し、独立した科学者が検証できない状況での批判は不当であり、データ公開を求めているという指摘がある。
代表コメント(原文)
“What an asshole: On the other side of the AI-reasoning fence, the disdain seems to be mutual. “These ‘scientific’ papers from last summer — I would…”
「推論」という用語の定義をめぐる議論、企業によるデータ公開の不透明さへの批判、そしてニューラルネットワークが「賢いハンス」現象と同様のメカニズムで動作しているという指摘が含まれる。また、Transformer の技術的制約と「推論トレース」の役割に関する技術的な考察も交わされている。
「推論」という言葉の意味を問うこと自体が機能の実態ではなく、意味論的な議論に終始しているという見方がある。また、人間も誤りや認知バイアスを持つため、AI の不完全さを「真の推論ではない」として切り捨てるのは「存在の連鎖」への執着であり、実用的な問題解決能力を重視すべきだという意見が示された。
“I'll admit that I find this discussion a bit navel-gazy. It has become a question of semantics not a question of actual functionality. The question has become "what do we mean when we use the word 'reasoning'" which is uninteresting. Dijkstra said[1] "... the…”
“Transformers lack recursion and are limited by the network's fixed depth, so "reasoning", IMHO, is basically a way to emulate deeper recursion. As we go through the layers, concepts are pattern-matched and refined, but at some point we have to stop and cannot…”
Tailscale didn't stop the Hugging Face intrusion
Tailscale の透明性とマーケティング評価
意見が分かれる脆弱性がなかったにもかかわらず報告した姿勢を高く評価する声がある一方、記事自体が巧妙なマーケティング活動であり、Hugging Face 側の人的ミス(環境変数ファイルへの再利用可能な認証キーの記載)を強調しているという批判的な見方もある。
代表コメント(原文)
“>No “vulnerabilities” in Tailscale were found or exploited, and that might make it even more uncomfortable for us. [...] But, we're a security tool.…”
ゼロトラストの誤解と構成要件
懸念Tailscale が「ゼロトラストネットワーク」と称していることへの懸念が示された。ツール自体はゼロトラストアーキテクチャの実装に使用可能だが、一般的なデプロイではマシン指向であり、サービスやリクエスト単位での厳密なアクセス制御(ACL)が行われない場合、そのマシン上のプロセスには広範な権限が与えられるため、ユーザーが「Tailscale を使えば完了」と誤解する危険性がある。
代表コメント(原文)
“"Tailscale is a zero trust network!" That's the problem. Tailscale is not zero trust. Tailscale can be used to implement a zero trust architecture,…”
Hugging Face のセキュリティインシデントにおいて Tailscale に脆弱性がなかったことへの評価、認証キーの管理ミスによる人的要因、およびゼロトラストアーキテクチャの実装における誤解や長期間有効な資格証明(long-lived credentials)の是非について議論が行われた。
脆弱性がなかったにもかかわらず報告した姿勢を高く評価する声がある一方、記事自体が巧妙なマーケティング活動であり、Hugging Face 側の人的ミス(環境変数ファイルへの再利用可能な認証キーの記載)を強調しているという批判的な見方もある。
“>No “vulnerabilities” in Tailscale were found or exploited, and that might make it even more uncomfortable for us. [...] But, we're a security tool. Their intrusion is our intrusion, and it's our job to take it seriously. im a happy customer of tailscale, so…”
“Wow, this article is super smart marketing by tailscale. Not only do they list all the nice and expensive features, that can help in such a situation but they also show that someone at huggingface made a very stupid thing by writing a reusable auth key in an…”
Google fixed more Chrome bugs in June than over the past two years, thanks to AI
AI による修正の信頼性と副作用への懸念
懸念AI が生成したコードは表面的には正しく見えても、エッジケースを処理できていない可能性があり、自動修正が新たなバグを導入したり、偽陽性が多かったりするリスクがあるとの指摘があります。また、リソース投入量の変化を示さないため、ツール自体の効果ではなく単に人員を増やして対応しただけの可能性も示唆されています。
代表コメント(原文)
“What remains to be seen is whether Google also introduced more Chrome bugs in June than over the past two years, thanks to AI. The big problem is…”
C++ の複雑さと Rust への移行の必要性
実体験・見立て発見されたバグの多くがメモリ関連であり、これは C/C++ の手動メモリアクセス管理という根本的な課題に起因すると指摘されています。大規模プロジェクトにおいて C/C++ は人間が完璧を維持するには複雑すぎるため、Rust などのメモリスafe な言語への移行が急務であるとの意見があります。
代表コメント(原文)
“To me this merely signals how broken C++ development really is. Most if not all of the bugs being uncovered are memory related and therefore…”
コメントでは、AI によるバグ修正数の増加が真の効果を示すのか、あるいは単なるリソース投入や管理上の圧力によるものなのかという懐疑的な見方が多く見られました。また、C++ の複雑さに対する根本的な問題提起や、AI ツールを「完全な自動化」ではなく人間との協働ツールとして位置づけるべきだという経験則に基づく意見も交わされました。
AI が生成したコードは表面的には正しく見えても、エッジケースを処理できていない可能性があり、自動修正が新たなバグを導入したり、偽陽性が多かったりするリスクがあるとの指摘があります。また、リソース投入量の変化を示さないため、ツール自体の効果ではなく単に人員を増やして対応しただけの可能性も示唆されています。
“What remains to be seen is whether Google also introduced more Chrome bugs in June than over the past two years, thanks to AI. The big problem is that AI output can be very convincing and look "right", even appear to work, until you examine it in detail and…”
“How many of those automated fixes were reverted? How many introduced a new bug? What's the false positive rate on the finding agents? The post has counts for everything that went right and nothing for what could go wrong.”
DeepSeek-V4-Flash Update
ローカル実行とコスト効率の向上
支持DeepSeek-V4-Flash は 10,000 ドル未満のハードウェアで自宅環境(at home)で動作可能であり、K3 や GLM の大規模モデルでは実現が困難な点として評価されている。開発者はこのモデルを日常業務の 90% に使用し、数行から数百行の変更に対して数分待たされることなく高速に反復できるため、フロントティアモデルよりも優れていると述べている。また、複数ターンにわたるセッションでも約 0.5 ドルという低コストで完了でき、OpenWebUi を自己ホストして月 18 ドル程度で運用する事例も報告されている。
代表コメント(原文)
“This is more exciting than k3, IMO. Dsv4 models are extremely cheap to serve. Improving their capabilities has lots of downstream effects, as it…”
性能と価格のバランスに対する評価
支持300B パラメータを持つこのモデルが、以前の DS4 Pro プレビュー(1.8T パラメータ)や GPT 5.6 Luna を上回る性能を示すというベンチマーク結果について、実際の利用状況と照らし合わせて驚きと称賛の声が上がっている。30 日間で 320M トークンを処理し、コードレビューやタスク実行において失望させなかったという実証データも共有されており、知能対コストの指数でさらにリードを広げていると評価されている。
代表コメント(原文)
“I use deepseek for a lot of my personal day-to-day agent needs, and I will simply put this here and let this speak for itself, last 30 days: - Cost:…”
開発者らが DeepSeek-V4-Flash の低コスト・高速な推論能力と、ローカル環境での実行可能性について実体験を共有し、その性能がプロフェッショナル向けモデルや他社製品に匹敵する点で期待を示している。一方で、中国企業製モデルの利用に伴うデータセキュリティや、ベンダー依存によるサービスの継続性に対する懸念も同時に提起されている。
DeepSeek-V4-Flash は 10,000 ドル未満のハードウェアで自宅環境(at home)で動作可能であり、K3 や GLM の大規模モデルでは実現が困難な点として評価されている。開発者はこのモデルを日常業務の 90% に使用し、数行から数百行の変更に対して数分待たされることなく高速に反復できるため、フロントティアモデルよりも優れていると述べている。また、複数ターンにわたるセッションでも約 0.5 ドルという低コストで完了でき、OpenWebUi を自己ホストして月 18 ドル程度で運用する事例も報告されている。
“This is more exciting than k3, IMO. Dsv4 models are extremely cheap to serve. Improving their capabilities has lots of downstream effects, as it becomes "good enough" for more and more tasks. DS was serving the pro version at extremely low prices for a long…”
“I've been driving flash model for 90% of my tasks. It's better than pro (for unknown reasons), very cheap and fast. I try to keep changes under 1000 lines and drive architectural decisions myself, barely notice any difference compared to frontier models. The…”
DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis
パフォーマンスと価格の競争力
支持DeepSeek V4 Flash 0731 が OpenAI の Luna や Gemini 3.6 レベルの知能を $0.28/m という低コストで提供しており、コードエージェントタスクでも高い評価を得ていることが指摘されています。また、V4 Pro モデルが間もなく発表され、Opus 5 に匹敵する性能を持つ可能性への期待も示されています。
代表コメント(原文)
“So GLM 5.2/Gemini 3.6 level intelligence for $0.28/m output. And their updated Pro model coming soon.... Plus a size you can genuinely run at home:…”
Gemini Robotics 2 brings whole body intelligence to robots
Google の開発姿勢と他社との比較
支持OpenAI や Anthropic が過熱した期待(「過熱した期待」)に包まれる中、Google はモデル、画像生成、動画生成、音楽生成、ロボット工学など幅広い分野で着実に進歩を続けているという評価がある。Google が Competent な企業として淡々と開発を進めている一方で、競合他社は誇張された報道によって不当なバリュエーションを支えられているという見方もある。
代表コメント(原文)
“While Anthropic and Open AI get 80% of the attention here, it's impressive to see how much Google is doing: near frontier model, fast models, open…”
Advancing the price-performance frontier with GPT‑5.6
コスト削減の技術的根拠と推論規模
実体験・見立てカーネル作業によるエンドツーエンドのコスト削減(20%)とトークン生成効率の向上(15%)が、月間数十億ドル規模の節約につながる可能性について議論されている。過去の AI による自己最適化事例と比較し、現在の技術進歩が以前よりもはるかに高い効率改善を実現している点が強調されている。
代表コメント(原文)
“> The kernel work helped reduce the end-to-end cost of serving the model by 20%, while its experiments increased token-generation efficiency by more…”
OpenAI の技術スタッフが他社の研究成果を「古いモデルの欠陥」として退け、自社の最新モデルは問題がないと主張しているが、同社は2024年以降も「生」の思考連鎖(chain of thought)データを公開していない。この不透明な姿勢に対し、独立した科学者が検証できない状況での批判は不当であり、データ公開を求めているという指摘がある。
“What an asshole: On the other side of the AI-reasoning fence, the disdain seems to be mutual. “These ‘scientific’ papers from last summer — I would put this in big, big air quotes,” said Sébastien Bubeck, a member of OpenAI’s technical staff (and a prominent…”
LLM は分類器であり、出力する推論トークンに関わらず、実際には人間の意図や訓練データのパターンに反応して正解を導いている可能性が高い。これは過去の「賢いハンス」現象と同様で、予測の根拠が人間と同じ理由である保証はないという見方が示された。
“Back in the day it was a bit of a cliche to bring up “clever Hans”, the horse that could do math, when talking about machine learning. He couldn’t do math but he read some cues from his handler of pick the write answers, the handler iirc wasn’t in on it. The…”
取得時点の一部コメントを整理したもので、HN全体の総意ではありません。
Tailscale が「ゼロトラストネットワーク」と称していることへの懸念が示された。ツール自体はゼロトラストアーキテクチャの実装に使用可能だが、一般的なデプロイではマシン指向であり、サービスやリクエスト単位での厳密なアクセス制御(ACL)が行われない場合、そのマシン上のプロセスには広範な権限が与えられるため、ユーザーが「Tailscale を使えば完了」と誤解する危険性がある。
“"Tailscale is a zero trust network!" That's the problem. Tailscale is not zero trust. Tailscale can be used to implement a zero trust architecture, with if deployed with sufficiently granular ACLs, but the most common deployment is machine-oriented, rather…”
インシデントの原因が人的ミス(長期間有効な認証キーの漏洩)であり、Tailscale の脆弱性ではないという指摘がある。一方で、短期有効な資格証明への移行は実装コストや運用上の摩擦が大きく、メモリに保存される長期資格証明を保護する手段も限られるなど、完全な解決策がないことへの批判や、TPM 利用の放棄についての疑問が提示された。
“> One of those 136 credentials was a reusable Tailscale auth key, used to create new Tailscale CI (continuous integration, used for automated testing) nodes in their tailnet. The agent copied that key into a series of external sandboxes and used it, over…”
“This actually shows that it was a human error on HuggingFace's end that led to that "breach". I would say HuggingFace needs to prioritize both security metrics/alerts and metrics/alerts for node count. And not leave long-lived keys accessible easily like…”
取得時点の一部コメントを整理したもので、HN全体の総意ではありません。
発見されたバグの多くがメモリ関連であり、これは C/C++ の手動メモリアクセス管理という根本的な課題に起因すると指摘されています。大規模プロジェクトにおいて C/C++ は人間が完璧を維持するには複雑すぎるため、Rust などのメモリスafe な言語への移行が急務であるとの意見があります。
“To me this merely signals how broken C++ development really is. Most if not all of the bugs being uncovered are memory related and therefore intimately tied to the mental memory model of C and C++, namely manual memory management. It's fine for a C or C++…”
AI を「コード生成」だけでなく、依存関係の追跡やリファクタリング支援などのツールとして使うべきであり、完全自動化よりも人間がレビューするプロセスを強化する方向性が有効であるとの経験則が述べられています。一方で、Google が AI で独自にバグを修正できるようになれば、オープンソースとしての Chromium への貢献が減り、結果的に他の Chromium フォークの維持が困難になるという懸念も示されています。
“I've also seen similar recently at my place of work with LLM linting. Lots of hard to spot bugs in old very critical code have been caught. I remember seeing similar kinds of impacts back when linting or sanitizers were introduced. They're good for making it…”
“Not that I don't believe its possible to fix a lot of bugs, I also wonder what the actual dynamic was. Were the people in team working much more than usual as well? Given its Google, I wouldn't be surprised if there was an "internal push" to fix more bugs…”
取得時点の一部コメントを整理したもので、HN全体の総意ではありません。
300B パラメータを持つこのモデルが、以前の DS4 Pro プレビュー(1.8T パラメータ)や GPT 5.6 Luna を上回る性能を示すというベンチマーク結果について、実際の利用状況と照らし合わせて驚きと称賛の声が上がっている。30 日間で 320M トークンを処理し、コードレビューやタスク実行において失望させなかったという実証データも共有されており、知能対コストの指数でさらにリードを広げていると評価されている。
“I use deepseek for a lot of my personal day-to-day agent needs, and I will simply put this here and let this speak for itself, last 30 days: - Cost: $4.55USD - API requests: 3,467 - Tokens: 323,183,886 And as an engineer who leads a small team, I have very…”
“If the benchmarks are real and reflect actual use, then this is an insane model. This 300B model outperforms the previous DS4 Pro preview model (1.8T params), and it looks like it outperforms GPT 5.6 Luna too. And it's still cheaper than Luna, even with the…”
DeepSeek や Moonshot のような中国企業製モデルを利用する際、API を通じてコードベースの情報が中国に漏洩している可能性が指摘されている。また、米国および中国の両方のモデル利用において、不正な jailbreak への耐性や、開発者が意図せず生成した偽情報の確認といった標準的なセキュリティチェックと QA/QC の基準について質問が投げかけられている。一部のユーザーはセキュリティガードに抵触することなくバイナリの逆解析を行っているが、他のプラットフォームでは異なる対応が見られる点についても議論されている。
“For both US and China models - what standard security checks and QaQc are you all doing? We're running small gamuts to test for unsolicited jailbreaks (model jailbreaks you) and incorrect records (Fake Accuracy - as Easter Egg or common thread) meaning…”
“Very promising. So it will both keep the speed and reduced price, yet exceed performance of the quite sufficient deepseek-v4-pro? Should be extending the lead in intelligence/cost index, as deepseek-v4-flash already were the most price efficient model, which…”
モデルの機能が「永遠に」利用可能であるという主張に対し、Butlerian Jihad(『フランク・ハーバートの宇宙シリーズ』における AI 禁止の概念)を皮肉交じりに引用して、サービスがプロバイダーの都合で突然終了するリスクや、将来の継続性に対する懐疑的な見方が示されている。
“> Whatever capabilities they get, can be used "forever" going forward. "Forever" gets the scare quotes because it is implied only up until the Butlerian Jihad?”
取得時点の一部コメントを整理したもので、HN全体の総意ではありません。