Technical Coding and Programming Performance
Developers report superb results in bug detection and scientific coding, with many considering it a top-tier tool for technical tasks.
Users celebrate the Technical Coding performance and aggressive Pricing of DeepSeek V4 Pro, but significant frustration persists regarding extreme Inference Latency and potential Geopolitical Bias within the model's outputs.
Developers report superb results in bug detection and scientific coding, with many considering it a top-tier tool for technical tasks.
The model is widely praised for its disruptive cost efficiency and the availability of open weights compared to closed-source rivals.
Significant complaints focus on 'next level' slowness during reasoning tasks and the massive VRAM requirements for local hardware execution.
Skeptics highlight potential censorship and political biases tied to the model's Chinese origin, alongside general data privacy anxieties.
Claude seems the most realistic, but Gemini seems like it best understood the assignment.
You asked for fractal though, which clearly is more like gemini and deepseek's result
Competition is good. Open Source even better. Cheaper and good-enough AI should benefit everybody. Thank you Openseek 🙏
To me the v4 pro seems to be hugely undertrained. I expect we're going to see huge gains in that model when we get new checkpoints in coming months.
Sempre que um bom modelo oriental é lançado eu comemoro. Ter alternativas ao oligopólio do império do norte é fundamental
Qwen 3.6 27b. Melhor modelo que eu rodei localmente ( o que é esperado de um modelo novo da qwen), mas por uma margem insana. Eu não uso mais api pra nada que não seja os projetos mais ambiciosos. Alteração no código, navegação no pc, pesquisas, e até exploração mesmo de temas. Eu nunca fiquei entusiasmado pra usar um modelo por tanto tempo depois de ter testado. Finalmente é um modelo que é BOM O SUFICIENTE
O deepseek tá salvando minha vida. Tava cogitando um plano max do claude e fazendo as contas pra resolver um projeto gráfico em c++. Ele simplesmente mudo completamente o jogo. Vc trabalha lindamente com ele. Pau no c* dos EUA e sua arrogância. Sempre querendo f*der com o consumidor. Se tem algo que entrega valor pra humanidade em IA hoje, que torna acessível, são os modelos chineses. DeepSeek tá dando aula!!! inclusive pra outros modelos chineses que se aproveitam da reserva de mercado que os EUA impõem!! e parabéns pelo vídeo e pelas análises!! não aguentava mais essas análises de oneshot inúteis!!!
i use v4 pro and flash via openrouter as worker subagents. so they dont design, plan, discuss or research don't do anything other than implementations. u get gpt 5.4 level performance almost for 5.4 nano pricing, in fact its even cheaper than nano. flash is even crazier. i dont think there is any model out there which can compete with this in cost effectiveness. its by far the cheapest sota model. like by far. the reason I dont fully use them is becuz i still prefer claude/gpt models for design etc, not that i tested chinese ones yet on that front. i probably should.
The funny thing is that I use both ChatGPT and DeepSeek in my personal life, but they're for different purposes, based on the results I've interpreted. This video fully answered why I prefer them for their respective purposes: 1. DeepSeek for anything technical or objective (e.g. troubleshooting my PC OS, mathematics, science, etc.) 2. ChatGPT for anything social or subjective (e.g. social planning, restaurant recommendations, arts, etc.)
Yeah though that's for dense. MoE have different scaling, depending on sparsity - https://arxiv.org/abs/2507.17702
Claude seems the most realistic, but Gemini seems like it best understood the assignment.
You asked for fractal though, which clearly is more like gemini and deepseek's result
Competition is good. Open Source even better. Cheaper and good-enough AI should benefit everybody. Thank you Openseek 🙏
To me the v4 pro seems to be hugely undertrained. I expect we're going to see huge gains in that model when we get new checkpoints in coming months.
Sempre que um bom modelo oriental é lançado eu comemoro. Ter alternativas ao oligopólio do império do norte é fundamental
Qwen 3.6 27b. Melhor modelo que eu rodei localmente ( o que é esperado de um modelo novo da qwen), mas por uma margem insana. Eu não uso mais api pra nada que não seja os projetos mais ambiciosos. Alteração no código, navegação no pc, pesquisas, e até exploração mesmo de temas. Eu nunca fiquei entusiasmado pra usar um modelo por tanto tempo depois de ter testado. Finalmente é um modelo que é BOM O SUFICIENTE
O deepseek tá salvando minha vida. Tava cogitando um plano max do claude e fazendo as contas pra resolver um projeto gráfico em c++. Ele simplesmente mudo completamente o jogo. Vc trabalha lindamente com ele. Pau no c* dos EUA e sua arrogância. Sempre querendo f*der com o consumidor. Se tem algo que entrega valor pra humanidade em IA hoje, que torna acessível, são os modelos chineses. DeepSeek tá dando aula!!! inclusive pra outros modelos chineses que se aproveitam da reserva de mercado que os EUA impõem!! e parabéns pelo vídeo e pelas análises!! não aguentava mais essas análises de oneshot inúteis!!!
i use v4 pro and flash via openrouter as worker subagents. so they dont design, plan, discuss or research don't do anything other than implementations. u get gpt 5.4 level performance almost for 5.4 nano pricing, in fact its even cheaper than nano. flash is even crazier. i dont think there is any model out there which can compete with this in cost effectiveness. its by far the cheapest sota model. like by far. the reason I dont fully use them is becuz i still prefer claude/gpt models for design etc, not that i tested chinese ones yet on that front. i probably should.
The funny thing is that I use both ChatGPT and DeepSeek in my personal life, but they're for different purposes, based on the results I've interpreted. This video fully answered why I prefer them for their respective purposes: 1. DeepSeek for anything technical or objective (e.g. troubleshooting my PC OS, mathematics, science, etc.) 2. ChatGPT for anything social or subjective (e.g. social planning, restaurant recommendations, arts, etc.)
Yeah though that's for dense. MoE have different scaling, depending on sparsity - https://arxiv.org/abs/2507.17702
Was super excited but these tests are pretty niche. Would be much better to see real world coding like building a SAAS web app or adding a feature to an existing repo.
In OpenCode, there is a max settings for DeepSeek V4. If you put 5.5 on high on codex, might as well put Deepseek on high or max to be fair
You should really use the same harness if you want to really compare the models between each other. You cannot infer causality between the model and the evaluation metrics if each uses a different harness. That being said, it still provided some interesting info :)
The fact that we have to rely on China for open source local models should tell you everything you need to know about the state of the Western World.
And none of the V4s can actually analyze images, it seems... 🤨😑
Where is Google in all this? After the release of Gemini 3 everyone thought they would be running away with it but currently they don't seem to be playing a role at all.
Meanwhile Z.AI guys raising their price by 3x while limiting usage by almost 7x through weekly limits and telling day one users they have to move to the new plan.
Did my usual hallucination test about identifying a contest. I'll put 5.5 result here as well: GPT 5.5 was inconsistent - it hallucinated the contest, but managed to solve the problem in 2 minutes which I thought was insane because it was an IMO problem 3. On another try it did manage to output "IDK". When provided the non IMO problem later given to DeepSeek V4 Pro, it also confidently hallucinated (it even said "I'm quite confident!"). Seems like a regression from 5.4 (which gave an incorrect answer but mentioned it was unsure) DeepSeek V4 Pro: This was the 2nd model (after Gemini 3.1 Pro) to correctly identify the contest and it did so in 11 seconds, which was faster than Gemini. Crazy. But I think it shows how hard they're RL'ing these models on historical Olympiad problems that they've now completely memorized it. Not gonna lie, surprised GPT 5.5 couldn't identify it in comparison. If I provide it with a much more obscure problem not from the IMO, it confidently hallucinates. DeepSeek V4 Flash: Timed out, got nowhere close in the thinking. For the other problem, confidently hallucinates.
vLLM has good blog entry on V4 and they break down KV cache usage, I think we can say that it's an authorative source - https://vllm.ai/blog/deepseek-v4 Indexer is taking a lot of prefill time and Zhipu uses IndexCache, while DS doesn't seem to be fixing it themselves it seems, so it may be a roadblock to fast long-context inference - https://github.com/THUDM/IndexCache
V4 is not supported in llama.cpp yet, so I did not yet get to try it on my rig. As of Kimi-K2.6 and GLM-5.1, GLM-5.1 seems to solve better complicated tasks, like resolving complex git rebase conflicts, while K2.6 can get stuck. But K2.6 is faster and overall still smart enough for most tasks, so I probably will use it most often.
Graph based on sampled comments per item (n≤30)
AI Code Clash
Asian Boss
Matthew Berman
AI Explained
AI Search
零度解说
Bijan Bowen
David Ondrej
Vini - AI Coders Academy
Chase AI
零度解说
r/singularity
r/singularity
r/LocalLLaMA
r/LocalLLaMA
r/singularity
r/singularity
r/LocalLLaMA
r/singularity
r/LocalLLaMA
r/LocalLLaMA
r/LocalLLaMA
r/singularity
impact_sy