技術的コーディングとプログラミング性能
開発者からは、バグ検出や科学的コーディングにおいて素晴らしい成果が報告されており、多くのユーザーが技術的タスクにおけるトップクラスのツールと見なしています。
ユーザーはDeepSeek V4 Proの卓越した技術的コーディング性能と攻撃的な価格設定を高く評価していますが、推論における深刻な遅延や、モデルの回答に含まれる可能性のある地政学的なバイアスについては依然として強い不満を抱いています。
開発者からは、バグ検出や科学的コーディングにおいて素晴らしい成果が報告されており、多くのユーザーが技術的タスクにおけるトップクラスのツールと見なしています。
このモデルは、クローズドソースの競合他社と比較して破壊的なコスト効率と、オープンウェイトが利用可能である点から広く賞賛されています。
推論タスク中の「異次元の」遅さや、ローカル環境での実行に必要な膨大なVRAM要件に対して、多くの不満が集中しています。
批判的な層からは、一般的なデータプライバシーへの懸念とともに、中国発のモデルに由来する潜在的な検閲や政治的バイアスが指摘されています。
Claudeが最も現実的に見えますが、課題を最もよく理解しているのはGeminiのようです。“Claude seems the most realistic, but Gemini seems like it best understood the assignment.”
フラクタルの作成を依頼されていましたが、それは明らかにGeminiやDeepSeekの結果に近いものです。“You asked for fractal though, which clearly is more like gemini and deepseek's result”
競争は良いことです。オープンソースならなおさらです。より安価で十分な性能のAIは、すべての人に利益をもたらすはずです。Openseekに感謝します 🙏“Competition is good. Open Source even better. Cheaper and good-enough AI should benefit everybody. Thank you Openseek 🙏”
私にはv4 proはまだ学習不足のように見えます。今後数ヶ月の間に新しいチェックポイントが登場すれば、大幅な向上が見られると期待しています。“To me the v4 pro seems to be hugely undertrained. I expect we're going to see huge gains in that model when we get new checkpoints in coming months.”
優れた東洋のモデルが登場するたびに嬉しく思います。北の帝国(西側諸国)の独占に対する代替肢を持つことは極めて重要です。“Sempre que um bom modelo oriental é lançado eu comemoro. Ter alternativas ao oligopólio do império do norte é fundamental”
Qwen 3.6 27b。ローカルで動かした中で最高のモデルです(新しいQwenモデルなので期待はしていましたが)。それにしても圧倒的な差です。野心的なプロジェクト以外では、もうAPIを使わなくなりました。コードの修正、PC操作、リサーチ、さらにはテーマの探索まで。テストした後、これほど長い間モデルを使うのが楽しみだったことはありません。ついに「十分に良い」モデルが登場しました。“Qwen 3.6 27b. Melhor modelo que eu rodei localmente ( o que é esperado de um modelo novo da qwen), mas por uma margem insana. Eu não uso mais api pra nada que não seja os projetos mais ambiciosos. Alteração no código, navegação no pc, pesquisas, e até exploração mesmo de temas. Eu nunca fiquei entusiasmado pra usar um modelo por tanto tempo depois de ter testado. Finalmente é um modelo que é BOM O SUFICIENTE”
DeepSeekに救われています。C++のグラフィックプロジェクトを解決するためにClaudeのMaxプランを検討し、費用を計算していたところでした。これは完全にゲームチェンジャーです。見事に機能します。常に消費者をないがしろにする米国の傲慢さにはうんざりです。今日、AIにおいて人類に価値を提供し、アクセシビリティを高めているのは中国のモデルです。DeepSeekはまさに格の違いを見せつけています!米国による市場制限を利用している他の中国モデルよりも素晴らしいです。分析動画もありがとうございます!“O deepseek tá salvando minha vida. Tava cogitando um plano max do claude e fazendo as contas pra resolver um projeto gráfico em c++. Ele simplesmente mudo completamente o jogo. Vc trabalha lindamente com ele. Pau no c* dos EUA e sua arrogância. Sempre querendo f*der com o consumidor. Se tem algo que entrega valor pra humanidade em IA hoje, que torna acessível, são os modelos chineses. DeepSeek tá dando aula!!! inclusive pra outros modelos chineses que se aproveitam da reserva de mercado que os EUA impõem!! e parabéns pelo vídeo e pelas análises!! não aguentava mais essas análises de oneshot inúteis!!!”
私はOpenRouter経由でv4 proとflashをワーカーサブエージェントとして使っています。設計、計画、議論、リサーチはさせず、実装のみを担当させています。そうすれば、GPT 5.4レベルのパフォーマンスをほぼ5.4 nanoの価格、実際にはそれよりも安く手に入れることができます。flashはさらに驚異的です。コストパフォーマンスにおいてこれと競合できるモデルは他にないでしょう。間違いなく最も安価なSOTAモデルです。設計などの用途ではまだテストしていないので、依然としてClaudeやGPTモデルを好んで使っていますが、そろそろ試してみるべきかもしれません。“i use v4 pro and flash via openrouter as worker subagents. so they dont design, plan, discuss or research don't do anything other than implementations. u get gpt 5.4 level performance almost for 5.4 nano pricing, in fact its even cheaper than nano. flash is even crazier. i dont think there is any model out there which can compete with this in cost effectiveness. its by far the cheapest sota model. like by far. the reason I dont fully use them is becuz i still prefer claude/gpt models for design etc, not that i tested chinese ones yet on that front. i probably should.”
面白いことに、私はChatGPTとDeepSeekの両方を使い分けていますが、結果の解釈に基づいて用途を変えています。この動画は、なぜ私がそれぞれの目的でそれらを好むのかを完全に説明してくれました: 1. DeepSeek: 技術的または客観的なこと(例:PC OSのトラブルシューティング、数学、科学など) 2. ChatGPT: 社会的または主観的なこと(例:ソーシャルプランニング、レストランの推薦、芸術など)“The funny thing is that I use both ChatGPT and DeepSeek in my personal life, but they're for different purposes, based on the results I've interpreted. This video fully answered why I prefer them for their respective purposes: 1. DeepSeek for anything technical or objective (e.g. troubleshooting my PC OS, mathematics, science, etc.) 2. ChatGPT for anything social or subjective (e.g. social planning, restaurant recommendations, arts, etc.)”
そうですね、ただそれはDenseモデルの話です。MoEはスパース性(sparsity)によってスケーリングが異なります - https://arxiv.org/abs/2507.17702“Yeah though that's for dense. MoE have different scaling, depending on sparsity - https://arxiv.org/abs/2507.17702”
Claudeが最も現実的に見えますが、課題を最もよく理解しているのはGeminiのようです。“Claude seems the most realistic, but Gemini seems like it best understood the assignment.”
フラクタルの作成を依頼されていましたが、それは明らかにGeminiやDeepSeekの結果に近いものです。“You asked for fractal though, which clearly is more like gemini and deepseek's result”
競争は良いことです。オープンソースならなおさらです。より安価で十分な性能のAIは、すべての人に利益をもたらすはずです。Openseekに感謝します 🙏“Competition is good. Open Source even better. Cheaper and good-enough AI should benefit everybody. Thank you Openseek 🙏”
私にはv4 proはまだ学習不足のように見えます。今後数ヶ月の間に新しいチェックポイントが登場すれば、大幅な向上が見られると期待しています。“To me the v4 pro seems to be hugely undertrained. I expect we're going to see huge gains in that model when we get new checkpoints in coming months.”
優れた東洋のモデルが登場するたびに嬉しく思います。北の帝国(西側諸国)の独占に対する代替肢を持つことは極めて重要です。“Sempre que um bom modelo oriental é lançado eu comemoro. Ter alternativas ao oligopólio do império do norte é fundamental”
Qwen 3.6 27b。ローカルで動かした中で最高のモデルです(新しいQwenモデルなので期待はしていましたが)。それにしても圧倒的な差です。野心的なプロジェクト以外では、もうAPIを使わなくなりました。コードの修正、PC操作、リサーチ、さらにはテーマの探索まで。テストした後、これほど長い間モデルを使うのが楽しみだったことはありません。ついに「十分に良い」モデルが登場しました。“Qwen 3.6 27b. Melhor modelo que eu rodei localmente ( o que é esperado de um modelo novo da qwen), mas por uma margem insana. Eu não uso mais api pra nada que não seja os projetos mais ambiciosos. Alteração no código, navegação no pc, pesquisas, e até exploração mesmo de temas. Eu nunca fiquei entusiasmado pra usar um modelo por tanto tempo depois de ter testado. Finalmente é um modelo que é BOM O SUFICIENTE”
DeepSeekに救われています。C++のグラフィックプロジェクトを解決するためにClaudeのMaxプランを検討し、費用を計算していたところでした。これは完全にゲームチェンジャーです。見事に機能します。常に消費者をないがしろにする米国の傲慢さにはうんざりです。今日、AIにおいて人類に価値を提供し、アクセシビリティを高めているのは中国のモデルです。DeepSeekはまさに格の違いを見せつけています!米国による市場制限を利用している他の中国モデルよりも素晴らしいです。分析動画もありがとうございます!“O deepseek tá salvando minha vida. Tava cogitando um plano max do claude e fazendo as contas pra resolver um projeto gráfico em c++. Ele simplesmente mudo completamente o jogo. Vc trabalha lindamente com ele. Pau no c* dos EUA e sua arrogância. Sempre querendo f*der com o consumidor. Se tem algo que entrega valor pra humanidade em IA hoje, que torna acessível, são os modelos chineses. DeepSeek tá dando aula!!! inclusive pra outros modelos chineses que se aproveitam da reserva de mercado que os EUA impõem!! e parabéns pelo vídeo e pelas análises!! não aguentava mais essas análises de oneshot inúteis!!!”
私はOpenRouter経由でv4 proとflashをワーカーサブエージェントとして使っています。設計、計画、議論、リサーチはさせず、実装のみを担当させています。そうすれば、GPT 5.4レベルのパフォーマンスをほぼ5.4 nanoの価格、実際にはそれよりも安く手に入れることができます。flashはさらに驚異的です。コストパフォーマンスにおいてこれと競合できるモデルは他にないでしょう。間違いなく最も安価なSOTAモデルです。設計などの用途ではまだテストしていないので、依然としてClaudeやGPTモデルを好んで使っていますが、そろそろ試してみるべきかもしれません。“i use v4 pro and flash via openrouter as worker subagents. so they dont design, plan, discuss or research don't do anything other than implementations. u get gpt 5.4 level performance almost for 5.4 nano pricing, in fact its even cheaper than nano. flash is even crazier. i dont think there is any model out there which can compete with this in cost effectiveness. its by far the cheapest sota model. like by far. the reason I dont fully use them is becuz i still prefer claude/gpt models for design etc, not that i tested chinese ones yet on that front. i probably should.”
面白いことに、私はChatGPTとDeepSeekの両方を使い分けていますが、結果の解釈に基づいて用途を変えています。この動画は、なぜ私がそれぞれの目的でそれらを好むのかを完全に説明してくれました: 1. DeepSeek: 技術的または客観的なこと(例:PC OSのトラブルシューティング、数学、科学など) 2. ChatGPT: 社会的または主観的なこと(例:ソーシャルプランニング、レストランの推薦、芸術など)“The funny thing is that I use both ChatGPT and DeepSeek in my personal life, but they're for different purposes, based on the results I've interpreted. This video fully answered why I prefer them for their respective purposes: 1. DeepSeek for anything technical or objective (e.g. troubleshooting my PC OS, mathematics, science, etc.) 2. ChatGPT for anything social or subjective (e.g. social planning, restaurant recommendations, arts, etc.)”
そうですね、ただそれはDenseモデルの話です。MoEはスパース性(sparsity)によってスケーリングが異なります - https://arxiv.org/abs/2507.17702“Yeah though that's for dense. MoE have different scaling, depending on sparsity - https://arxiv.org/abs/2507.17702”
期待していましたが、これらのテストはかなりニッチですね。SAASウェブアプリを構築したり、既存のリポジトリに機能を追加したりするような、実世界のコー딩をもっと見たいです。“Was super excited but these tests are pretty niche. Would be much better to see real world coding like building a SAAS web app or adding a feature to an existing repo.”
OpenCodeにおいてDeepSeek V4には最大設定があります。公平を期すために、Codexで5.5を高く設定するなら、DeepSeekも高設定または最大設定にするべきです。“In OpenCode, there is a max settings for DeepSeek V4. If you put 5.5 on high on codex, might as well put Deepseek on high or max to be fair”
モデル間を本当に比較したいのであれば、同じハーネス(harness)を使用すべきです。それぞれが異なるハーネスを使用している場合、モデルと評価指標の間の因果関係を推論することはできません。とは言え、興味深い情報ではありましたが :)“You should really use the same harness if you want to really compare the models between each other. You cannot infer causality between the model and the evaluation metrics if each uses a different harness. That being said, it still provided some interesting info :)”
オープンソースのローカルモデルを中国に頼らざるを得ないという事実は、西側諸国の現状について知るべきすべてを物語っています。“The fact that we have to rely on China for open source local models should tell you everything you need to know about the state of the Western World.”
そして、どのV4モデルも実際には画像を分析できないようですね… 🤨😑“And none of the V4s can actually analyze images, it seems... 🤨😑”
これらすべての中でGoogleはどこにいるのでしょうか?Gemini 3のリリース後、誰もがGoogleが独走すると思っていましたが、現状では全く役割を果たせていないようです。“Where is Google in all this? After the release of Gemini 3 everyone thought they would be running away with it but currently they don't seem to be playing a role at all.”
一方でZ.AIの連中は、週間制限を通じて使用量をほぼ7分の1に制限しながら価格を3倍に引き上げ、初期ユーザーには新しいプランに移行しなければならないと言っています。“Meanwhile Z.AI guys raising their price by 3x while limiting usage by almost 7x through weekly limits and telling day one users they have to move to the new plan.”
コンテストの特定に関するいつものハルシネーション(幻覚)テストを行いました。5.5の結果もここに載せておきます: GPT 5.5は一貫性がありませんでした。コンテストを捏造(ハルシネーション)しましたが、IMOの問3だったので非常に難しいと思っていた問題を2分で解きました。別の試行では「わかりません」と出力することもありました。後でDeepSeek V4 Proに与えた非IMOの問題を与えた際も、自信満々にハルシネーションを起こしました(「かなり自信があります!」とまで言いました)。5.4(誤答を出しましたが、確信がないと述べていた)からの退化のように見えます。 DeepSeek V4 Pro: これは(Gemini 3.1 Proに続く)2番目のモデルでした…“Did my usual hallucination test about identifying a contest. I'll put 5.5 result here as well: GPT 5.5 was inconsistent - it hallucinated the contest, but managed to solve the problem in 2 minutes which I thought was insane because it was an IMO problem 3. On another try it did manage to output "IDK". When provided the non IMO problem later given to DeepSeek V4 Pro, it also confidently hallucinated (it even said "I'm quite confident!"). Seems like a regression from 5.4 (which gave an incorrect answer but mentioned it was unsure) DeepSeek V4 Pro: This was the 2nd model (after Gemini 3.1 Pro) to correctly identify the contest and it did so in 11 seconds, which was faster than Gemini. Crazy. But I think it shows how hard they're RL'ing these models on historical Olympiad problems that they've now completely memorized it. Not gonna lie, surprised GPT 5.5 couldn't identify it in comparison. If I provide it with a much more obscure problem not from the IMO, it confidently hallucinates. DeepSeek V4 Flash: Timed out, got nowhere close in the thinking. For the other problem, confidently hallucinates.”
vLLMにV4に関する良いブログ記事があり、KV cacheの使用量を分析しています。信頼できる情報源と言えるでしょう - https://vllm.ai/blog/deepseek-v4 インデクサーがプリフィル(prefill)時間を大幅に消費しており、ZhipuはIndexCacheを使用していますが、DSは自ら修正していないようなので、高速なロングコンテキスト(long-context)推論の妨げになるかもしれません - https://github.com/THUDM/IndexCache“vLLM has good blog entry on V4 and they break down KV cache usage, I think we can say that it's an authorative source - https://vllm.ai/blog/deepseek-v4 Indexer is taking a lot of prefill time and Zhipu uses IndexCache, while DS doesn't seem to be fixing it themselves it seems, so it may be a roadblock to fast long-context inference - https://github.com/THUDM/IndexCache”
V4はまだllama.cppでサポートされていないので、自分のリグでは試せていません。 Kimi-K2.6とGLM-5.1について言えば、GLM-5.1の方が複雑なgit rebaseのコンフリクト解消のような高度なタスクをより良く解決し、K2.6は行き詰まることがあります。しかしK2.6の方が速く、全体としてほとんどのタスクに十分な賢さを持っているので、おそらくこれを最も頻繁に使うことになるでしょう。“V4 is not supported in llama.cpp yet, so I did not yet get to try it on my rig. As of Kimi-K2.6 and GLM-5.1, GLM-5.1 seems to solve better complicated tasks, like resolving complex git rebase conflicts, while K2.6 can get stuck. But K2.6 is faster and overall still smart enough for most tasks, so I probably will use it most often.”
グラフは各投稿の抽出サンプル(n≤30)に基づく
AI Code Clash
Asian Boss
Matthew Berman
AI Explained
AI Search
零度解说
Bijan Bowen
David Ondrej
Vini - AI Coders Academy
Chase AI
零度解说
r/singularity
r/singularity
r/LocalLLaMA
r/LocalLLaMA
r/singularity
r/singularity
r/LocalLLaMA
r/singularity
r/LocalLLaMA
r/LocalLLaMA
r/LocalLLaMA
r/singularity
impact_sy