プライバシーの利点と競争上の位置付け
多くのユーザーがフロンティアAPIに代わるプライベートなローカルの選択肢としてこのモデルを評価しており、以前のGemmaバージョンと比較して優れたtool-calling機能を指摘しています。
ユーザーはGemma 4のローカルでのプライバシー保護や優れたvision機能を高く評価していますが、一方で過酷なハードウェア要件や、時折発生する論理的失敗、繰り返されるhallucination(幻覚)に不満を感じています。
多くのユーザーがフロンティアAPIに代わるプライベートなローカルの選択肢としてこのモデルを評価しており、以前のGemmaバージョンと比較して優れたtool-calling機能を指摘しています。
VRAMの制約や低いtokens-per-secondのレートについて活発な議論が行われており、特に31Bモデルや各種quantizationsが消費者向けGPUでどのように動作するかを疑問視する声が上がっています。
コーディングタスクや複雑な論理的推論テストにおける無限ループやhallucinationなど、モデルの一貫性に関する問題が浮き彫りになっています。
vision性能に対してコミュニティから高い期待が寄せられており、特にOCRタスクの精度やbounding boxesの処理能力が賞賛されています。
素晴らしい!高い知能を維持しつつサイズを縮小し、Apache 2ライセンスでスマートフォンでも動作するなんて。これこそ私たちが待ち望んでいた2026年のニュースです。Google、この調子で頑張ってください。“Amazing! shrinking size while high intelligence, Apache 2, works on Phones. This is the 2026 news we need. Keep going Google.”
ローカルで動作するGemma 4 (gemma4:e2b 7.2GB)をインストールして、取り組んでいるClifford algebraプロジェクトのためにConstraint-Dynamical Hamiltonianを導出しました。その結果は本当に素晴らしかったと断言できます…。RAGを使ってファイルを読み込み、高度な数学をこなせる思考モデルがローカルで動くなんて…最高すぎます。 :)“I installed Gemma 4 (gemma4:e2b 7.2GB) running locally derived a Constraint-Dynamical Hamiltonian for a Clifford algebra project I am working on. And you have to trust me what it provided was amazing... So I have a thinking model running locally using RAG to read files and can do quite advanced math... that is the absolute bomb.. :)”
Gemma 4へようこそ! 🎉 そしてGemma 4の開発に携わったすべての方々に感謝します。皆さんの素晴らしい仕事に、私たちは皆感謝しています。“Gemma 4 welcome! 🎉 And thanks to everyone behind Gemma 4's development. We all appreciate the incredible work you all do.”
驚きはありませんね。GemmaはミニGeminiのようなものですから、そういった分野は得意です。GLM 5.1が本領を発揮するのはコーディングです。“Not surprised. Gemma is just a mini Gemini, it's good with that stuff. Where GLM 5.1 shines is coding.”
どのように実行されたか分かりませんが、llama.cppを使用してローカルで実行している場合は、b8660 llama.cppビルド(最新バージョンにはregression、別のtokenizationの問題があります)を使用し、--temp 0.3 --top-p 0.9 --min-p 0.1 --top-k 20 を試してみてください。26Bならもっとうまくいくはずです。また、Claudeはフォーマットの良さなどを好む傾向があるため、Booleanテストは適切ではありません。judge役には以下のプロンプトを試してみてください: 私は多くのタスクで多くのAIをベンチマークしています。あなたはjudgeです。LLMごとではなく、質問ごとに確認してください。各質問を精査し、すべてのAIに対して10点満点でスコアを付け、公平に…“I don't know how you ran it, if you're running it locally using llama.cpp, use the b8660 llama.cpp build (more recent versions have a regression, another tokenization issue) and use --temp 0.3 --top-p 0.9 --min-p 0.1 --top-k 20 I am sure the 26B will do much better. Also, Claude might favor better formatting etc., a boolean test is not good. Try the below prompt for the judge: I am benchmarking many AIs in many tasks. You are a judge. Go through them question by question, not LLM by LLM. Go through each question and, for every question, give all AIs a score out of 10, and be sure to be fair with them. Later, rank them all by their total score. MAKE SURE to evaluate them correctly, not based on vibe alone (check for misinformation, hallucinations, if they are useful or not, and not on formatting). PROMPT= AI 1: ... AI 2: ....”
LLMをjudgeにするのは勘弁してください。また、テストのためにGemma 4をどのように実行しているかにもよります。llama.cpp b8665のGemma 4用新しいカスタムパーサーで解決しました。以前は、下の画像を与えただけでテストに失敗していましたが、今は解けます。“LLM as judge = no thanks. It also depends how you're running Gemma 4 for the test. The new custom parser for gemma 4 in llama.cpp b8665 has fixed it for me. Before, it failed the test of just being given the image below. Now it solves it.”
今後の方向にとてもワクワクしています。次世代は日常的な用途のほとんどでfrontier級の品質になり、Intel B70のようなシングルGPUにも収まるようになるでしょう。turbo quantのような進歩があと数回あれば、フラッグシップのスマホでもSOTAレベルが可能になります。おそらくあと2世代くらいですね。もしAIのtakeoffが完全にedgeデバイスで動作するエージェントによって行われ、大手ラボの何兆ドルもの資本が陳腐化してしまったら経済がどうなるか本気で心配ですが、AIが一部の人間に支配されない良い道に傾いているのは非常に嬉しいことです。“Super excited about the direction things are going. Next generation will be frontier quality for most daily uses and fit on a single solid GPU like the Intel B70. A couple more turbo quant type advances and we're there on SOTA phones, prob two generations. Genuinely concerned about the economy if the AI takeoff is entirely agents running on edge devices and the major labs' trillions in capital goes stale, but very glad we're leaning towards the good path where AI won't be controlled by the few.”
Gemma 4は、AIが久しぶりに成し遂げた本物の飛躍です。サイズを小さくしながら、計算リソースの消費も抑えています。私のPCで動かしていますが、20gbの占有で400gbモデルに匹敵します…正気とは思えないレベルです。しかもApache 2.0なので、これを使って作った製品は何でも販売できます。“Gemma 4 is the first actual leap AI did in a "long" time. It makes it smaller but also use less computing power. I am running it on my PC and while it takes up 20gb its equivalent to a 400gb model… insane and on Apache 2.0 so you can make and sell any product you make with it.”
Gemma 2の頃から、単なるyes manではなく、対話の質が高いので重宝しています。同調性が高すぎるのは欠陥であり、Qwenのその点は好きになれません。(私が絶対に正しいです)“Even since Gemma 2 it's been useful for being good at interacting instead of being a 'yes man' (girl). Agreeableness is a flaw and I don't like it in Qwen. (I'm absolutely right)”
qwen3 coder nextが実際のゲームロジックで4bに負けるというのは、今週見たベンチマーク結果の中で最もやる気を削がれるものでした。playwright mcpが大きな役割を果たしていることが、このバラツキの主な原因でしょう。“qwen3 coder next losing to the 4b at actual game logic is the most demoralizing benchmark result i've seen this week, playwright mcp doing the heavy lifting probably explains a lot of the variance here.”
素晴らしい!高い知能を維持しつつサイズを縮小し、Apache 2ライセンスでスマートフォンでも動作するなんて。これこそ私たちが待ち望んでいた2026年のニュースです。Google、この調子で頑張ってください。“Amazing! shrinking size while high intelligence, Apache 2, works on Phones. This is the 2026 news we need. Keep going Google.”
ローカルで動作するGemma 4 (gemma4:e2b 7.2GB)をインストールして、取り組んでいるClifford algebraプロジェクトのためにConstraint-Dynamical Hamiltonianを導出しました。その結果は本当に素晴らしかったと断言できます…。RAGを使ってファイルを読み込み、高度な数学をこなせる思考モデルがローカルで動くなんて…最高すぎます。 :)“I installed Gemma 4 (gemma4:e2b 7.2GB) running locally derived a Constraint-Dynamical Hamiltonian for a Clifford algebra project I am working on. And you have to trust me what it provided was amazing... So I have a thinking model running locally using RAG to read files and can do quite advanced math... that is the absolute bomb.. :)”
Gemma 4へようこそ! 🎉 そしてGemma 4の開発に携わったすべての方々に感謝します。皆さんの素晴らしい仕事に、私たちは皆感謝しています。“Gemma 4 welcome! 🎉 And thanks to everyone behind Gemma 4's development. We all appreciate the incredible work you all do.”
驚きはありませんね。GemmaはミニGeminiのようなものですから、そういった分野は得意です。GLM 5.1が本領を発揮するのはコーディングです。“Not surprised. Gemma is just a mini Gemini, it's good with that stuff. Where GLM 5.1 shines is coding.”
どのように実行されたか分かりませんが、llama.cppを使用してローカルで実行している場合は、b8660 llama.cppビルド(最新バージョンにはregression、別のtokenizationの問題があります)を使用し、--temp 0.3 --top-p 0.9 --min-p 0.1 --top-k 20 を試してみてください。26Bならもっとうまくいくはずです。また、Claudeはフォーマットの良さなどを好む傾向があるため、Booleanテストは適切ではありません。judge役には以下のプロンプトを試してみてください: 私は多くのタスクで多くのAIをベンチマークしています。あなたはjudgeです。LLMごとではなく、質問ごとに確認してください。各質問を精査し、すべてのAIに対して10点満点でスコアを付け、公平に…“I don't know how you ran it, if you're running it locally using llama.cpp, use the b8660 llama.cpp build (more recent versions have a regression, another tokenization issue) and use --temp 0.3 --top-p 0.9 --min-p 0.1 --top-k 20 I am sure the 26B will do much better. Also, Claude might favor better formatting etc., a boolean test is not good. Try the below prompt for the judge: I am benchmarking many AIs in many tasks. You are a judge. Go through them question by question, not LLM by LLM. Go through each question and, for every question, give all AIs a score out of 10, and be sure to be fair with them. Later, rank them all by their total score. MAKE SURE to evaluate them correctly, not based on vibe alone (check for misinformation, hallucinations, if they are useful or not, and not on formatting). PROMPT= AI 1: ... AI 2: ....”
LLMをjudgeにするのは勘弁してください。また、テストのためにGemma 4をどのように実行しているかにもよります。llama.cpp b8665のGemma 4用新しいカスタムパーサーで解決しました。以前は、下の画像を与えただけでテストに失敗していましたが、今は解けます。“LLM as judge = no thanks. It also depends how you're running Gemma 4 for the test. The new custom parser for gemma 4 in llama.cpp b8665 has fixed it for me. Before, it failed the test of just being given the image below. Now it solves it.”
今後の方向にとてもワクワクしています。次世代は日常的な用途のほとんどでfrontier級の品質になり、Intel B70のようなシングルGPUにも収まるようになるでしょう。turbo quantのような進歩があと数回あれば、フラッグシップのスマホでもSOTAレベルが可能になります。おそらくあと2世代くらいですね。もしAIのtakeoffが完全にedgeデバイスで動作するエージェントによって行われ、大手ラボの何兆ドルもの資本が陳腐化してしまったら経済がどうなるか本気で心配ですが、AIが一部の人間に支配されない良い道に傾いているのは非常に嬉しいことです。“Super excited about the direction things are going. Next generation will be frontier quality for most daily uses and fit on a single solid GPU like the Intel B70. A couple more turbo quant type advances and we're there on SOTA phones, prob two generations. Genuinely concerned about the economy if the AI takeoff is entirely agents running on edge devices and the major labs' trillions in capital goes stale, but very glad we're leaning towards the good path where AI won't be controlled by the few.”
Gemma 4は、AIが久しぶりに成し遂げた本物の飛躍です。サイズを小さくしながら、計算リソースの消費も抑えています。私のPCで動かしていますが、20gbの占有で400gbモデルに匹敵します…正気とは思えないレベルです。しかもApache 2.0なので、これを使って作った製品は何でも販売できます。“Gemma 4 is the first actual leap AI did in a "long" time. It makes it smaller but also use less computing power. I am running it on my PC and while it takes up 20gb its equivalent to a 400gb model… insane and on Apache 2.0 so you can make and sell any product you make with it.”
Gemma 2の頃から、単なるyes manではなく、対話の質が高いので重宝しています。同調性が高すぎるのは欠陥であり、Qwenのその点は好きになれません。(私が絶対に正しいです)“Even since Gemma 2 it's been useful for being good at interacting instead of being a 'yes man' (girl). Agreeableness is a flaw and I don't like it in Qwen. (I'm absolutely right)”
qwen3 coder nextが実際のゲームロジックで4bに負けるというのは、今週見たベンチマーク結果の中で最もやる気を削がれるものでした。playwright mcpが大きな役割を果たしていることが、このバラツキの主な原因でしょう。“qwen3 coder next losing to the 4b at actual game logic is the most demoralizing benchmark result i've seen this week, playwright mcp doing the heavy lifting probably explains a lot of the variance here.”
>Q8_0でさえ、長いドキュメントで0.45、非ラテン文字で0.24のKL散逸を示しています。Q8_0からQ5_K_Sにかけて全カテゴリーでほぼ倍増していますが、科学とツール使用は一貫して最も低いままです(Q8_0で0.07と0.08)。これは重要な発見のように思えます。ほとんどの人はQ8_0がBF16と実質的に同じだと思い込んでいますから。“>Even Q8_0 shows a KL of 0.45 on long documents and 0.24 on non-Latin scripts. All categories roughly double from Q8_0 to Q5_K_S, but science and tool use remain the lowest throughout (0.07 and 0.08 at Q8_0). This looks like it's a significant finding. Most people assume Q8_0 to be virtually the same as BF16.”
何とも比較していないなら、この投稿の意図は何ですか? Q2-Q3のquant(これでも一部の用途には使えます!)とq4キャッシュを使えば、おそらく収まるでしょう。^ 収まるかどうかなど誰も気にしません。誰もが品質が保たれているかを気にしています。そして、収まるということは、速いということです。驚きもありません。“So what's the point of the post if you're not comparing it with anything? you can take a Q2-Q3 quant (still usable for some use!) and q4 cache and it would probably fit as well. ^ Nobody cares if it fits, everybody cares if the quality is there And if it fits = it's fast. Surprise.”
真のテストはtokens-per-secondではありません。256k以降もモデルが自身の出力を確実に読み取れるかどうかです。そこでこれらのquantsは破綻します。“the real test isn't tokens-per-second. it's whether the model still reads back its own output reliably after 256k. that's where these quants break.”
Q8がQ4よりもパフォーマンスが低いというのは奇妙に思えます。テストした正確なquantsのリンクを貼ってもらえますか?“Seems odd that the Q8 would perform worse than the Q4. Can you link the exact quants you tested?”
速度は申し分ないですが、KV quantによってどれほど悪影響が出ていますか?“Speeds all well and good but how badly does it suffer from the KV quant?”
私のRTX PRO 4500 32GB GPUでは、qwen3.5-27bの方がコンテキスト長(115K)を長く確保できます。ウェイトのVRAM占有が少ないからです。サーバーにlobsterが住んでいるような状況では、これはかなり重要です。“i know that on my RTX PRO 4500 32GB GPU, I get more context length (115K) with qwen3.5-27b, because its weight occupy less VRAM. Which is sorta important when you have a lobster living in your server.”
「car wash test」はあまり良くありません。なぜなら、それはLLMが犯す可能性のある、ほぼ無限に存在するembodied reasoningや常識の失敗例の中で、最もよく知られたものに過ぎないからです。モデル製作者は学習でそのような例を一つ修正することはできても、すべてを修正することは不可能です。“The 'car wash test' is not very good, because it's the best known example of a nearly infinite number of embodied reasoning/common sense fails an LLM can make. Model makers can patch one such example in training, they cannot patch them all.”
26Bと35B、両方のMOEを試しました。どちらもdenseモデルと比較して大幅に劣化しており、SLOPまみれです。26Bが最悪の戦犯です。“I've tried both MOEs - 26B and 35B. Both are significant downgrades compared to the dense models and loaded with SLOP. 26B is the worst culprit.”
他に投稿者のプロンプトを実行して確認した人はいますか?その回答が得られなかったと言いたいわけではありませんが、あくまで一つのデータポイントに過ぎません。私の経験では、突然めちゃくちゃ遅くなりました。賢さは健在ですが、なんていうか、レーシングチームだったのに今は歩いているような気分です。“Has anyone else run the OPs prompt to check on their own? Not saying they DIDN'T get that answer, but it's just one data point. My experience has been that it is suddenly effing SLOW. Still smart, but damn, feels like walking when we used to be a racecar team.”
興味がある方のために、私のDGXでGemma 4 31Bを実行したところ、Q8は6tps、Q4は10tpsでした。コンテキストウィンドウをフルでロードしましたが、数通のメッセージで新しいチャットを始めたので、使ったのはごく一部です。また、フルコンテキストでQ8を完全にロードするには、101gb必要でした O_o“For those interested, i ran Gemma 4 31B on my DGX, Q8 was doing 6tps and the Q4 was doing 10tps. It was loaded with full context window, but only using a tiny bit since i did a new chat with a few messages. Also to fully load the Q8 with full context, it was 101gb O_o”
グラフは各投稿の抽出サンプル(n≤30)に基づく
NetworkChuck
Google for Developers
DIY Smart Code
Teacher's Tech
零度解说
Zero to MVP
ByteMonk
Prasadtechintelugu
零度解说
Zero to MVP
Bart Slodyczka
Ishan Sharma
r/LocalLLaMA
r/LocalLLaMA
r/LocalLLaMA
r/LocalLLaMA
r/LocalLLaMA
r/LocalLLaMA
r/LocalLLaMA
r/LocalLLaMA
r/LocalLLaMA
r/LocalLLaMA
r/LocalLLaMA
themanmaran