코딩 역량 및 성능
사용자들은 1M 컨텍스트 창과 강력한 에이전트 능력을 높이 평가하며, 복잡한 프로그래밍 및 멀티 레포 작업에 강력한 도구라고 판단하고 있습니다.
개발자들은 Sonnet 5의 코딩 능력과 쉬운 통합 성능에 열광하고 있지만, 높은 가격 책정과 사이버 보안 작업에 대한 제한적인 안전 가드레일에 대해서는 많은 불만을 토로하고 있습니다.
사용자들은 1M 컨텍스트 창과 강력한 에이전트 능력을 높이 평가하며, 복잡한 프로그래밍 및 멀티 레포 작업에 강력한 도구라고 판단하고 있습니다.
높은 토큰 비용과 초기 가격 모델에 대한 상당한 비판이 집중되고 있으며, 이로 인해 이전 Opus 모델이나 경쟁사 모델에 비해 접근성이 떨어진다는 지적이 있습니다.
모델의 제한적인 사이버 보안 기능에 대해 논란이 있으며, 일부 사용자들은 이를 정부 주도의 검열이나 게이트키핑을 위한 '신어(Newspeak)'로 해석하고 있습니다.
개발자들은 이 모델을 이전 버전의 드롭인 대체품으로 빠르게 채택하고 있으며, GitHub Copilot 및 기존 API 워크플로우에 쉽게 통합된다는 점에 주목하고 있습니다.
드디어 새로운 Sonnet 모델이 나왔네요“Finally a new sonnet model”
GLM-5.2와 Sonnet 4.6으로 기사를 다시 써봤습니다. LLM은 비결정적이라 결과가 완전히 다르더군요. 하지만 GLM-5.2는 수작업으로 수정해야 할 미세한 실수가 많았던 반면, Sonnet은 두 번째 시도에서 모든 실수를 찾아내고 수정했습니다. 기획이나 코딩에서도 비슷한 상황이었어요. GLM-5.2는 서류상으로는 좋아 보이지만 실제 사용 결과는 달랐습니다. 제가 Claude나 GLM-5.2의 대변인은 아니지만... :) 2022년 11월부터 매일 LLM 모델을 사용해 온 결과, 모든 일반적인 테스트 결과는 각자의 프로젝트에서 직접 확인해 봐야 한다는 것을 깨달았습니다.“I have tried to rewrite an article with GLM-5.2 and with Sonnet 4.6. Completely different results as LLM is non-deterministic. But GLM-5.2 made a lot of subtle mistakes that needed to be corrected by hand. On the opposite, Sonnet found and corrected all mistakes in the second round. Similar situation was with planning and coding. GLM-5.2 seems to be good “on paper” but the real usage results was different. And I am not an attorney for Claude or GLM-5.2… :) But as I’ve been using LLM models daily since Nov 2022 I have realized that all common tests have to be confirmed in your project - there is no “one model rules them all” - you need to dig out a specific model from that LLM haystack with thousands of models. Benchmarks help but they start to be similar to fuel consumption specs in car ads - real consumption is different for everybody :)”
지금까지 짧게 테스트해 본 결과, 4.6보다 결과물도 좋고 속도도 약간 더 빨라서 확실히 업그레이드되었다는 느낌을 받았습니다.“The limited time I've had to test it so far, it's given me better results than 4.6 and a little quicker, so I think it's a noticeable step up.”
> 왜 이런 걸 자랑하는 걸까요? 마치 사람들이 사이버 보안 작업에 모델을 쓰고 싶어 한다는 걸 알면서도 의도적으로 그 능력을 거부하는 것 같습니다. Anthropic이 여기서 정확히 뭐라고 말하길 원하시나요? "전 세계에 저렴하게 제공하려는 이 모델이 해킹을 정말 잘합니다"라고요? Sonnet이 사이버 보안에 형편없다고 말하는 것이 여러 좋지 않은 선택지 중에서 그들이 할 수 있는 가장 합리적인 말입니다.“> Why would they brag about something like this? It's like they know people want to use models to perform cybersecurity tasks yet knowingly deny them the ability. What exactly do you want Anthropic to say here? "This model, the one we are about to give to the entire world for cheap, is really good at hacking"? Saying Sonnet is terrible at cybersecurity is the most reasonable thing they can say, out of a lot of bad options.”
이 시점에서는 무능한 미국 정부에 본때를 보여주기 위해서라도 Anthropic이 오픈 소스로 전환했으면 좋겠네요.“At this point I want Anthropic to go open source just to give the middle to the braindead US government.”
Fable을 다시 돌려줄 방법이나 찾아내는 게 좋을 거예요 😂“They better figure out how to get us Fable back😂”
fix: Claude Sonnet 5 적응형 사고(adaptive thinking) 요청 지원 — #98386 해결 관련: #98254 ## 해결된 문제 Claude Sonnet 5를 선택한 사용자가 기본, 적응형, 낮음 또는 꺼짐 사고 설정을 사용할 때 Anthropic API 요청 실패가 발생하거나 요청된 사고 모드가 손실되는 문제를 수정했습니다. Claude Sonnet 5는 Anthropic 적응형 사고를 사용합니다.“fix: support Claude Sonnet 5 adaptive thinking requests — Closes #98386 Related: #98254 ## What Problem This Solves Fixes an issue where users selecting Claude Sonnet 5 could hit Anthropic API request failures or lose the requested thinking mode when using default, adaptive, low, or off thinking settings. Claude Sonnet 5 uses Anthropic adaptive thinking se”
fix(model-core): claude-sonnet-5를 GA 1M-컨텍스트 모델로 인식 (fixes #5788) — ## 요약 1M 토큰 컨텍스트 창을 지원하지만 OMO 게이트가 이를 인식하지 못해 Anthropic 제공자의 기본값인 200K로 되돌아갔고, 압축(compaction)이 약 5배 정도 너무 빨리 트리거되었습니다. ## 원인 [ ] 1M 창을 두 개의 정규식으로 제어하는데, 어느 쪽도 일치하지 않았습니다. 첫 번째 브랜치는 다음에 즉시 오는 것을 요구하는데...“fix(model-core): recognize claude-sonnet-5 as GA 1M-context model (fixes #5788) — ## Summary has a 1M-token context window, but OMO's gate did not recognize it, so fell back to the 200K default on Anthropic providers and triggered compaction roughly 5x too early. ## Root Cause [ ]( gates the 1M window on two regexes: matches neither: the first branch requires immediately after ,”
> 그리고 Opus 4.8은 통과율이 더 높으면서도 여전히 더 저렴합니다. Opus만큼 텍스트를 쏟아내지 않는 한 의심스럽네요. Opus 4.8은 말 그대로 텍스트를 토하듯이 내뱉습니다. 장기적으로 볼 때, 특히 여기저기서 캐시 미스가 발생하면 비용의 대부분은 그 추가된 컨텍스트에서 나옵니다.“> And Opus 4.8 is still cheaper for a higher pass rate Unless it spams as much as Opus, I doubt it. Opus 4.8 literally spams text like puke. On a longer run especially if you get cache misses here and there the bulk of the cost is all the extra context it adds.”
서비스가 시작된 이후로 계속 써보고 있는데... 정말 제 역할을 톡톡히 하네요. Rust 플러그인과 서버 관리 패널을 코딩할 수 있고, 플러그인 설정과 패널을 연결해 줘서 .json 파일을 수정하거나 변경 사항을 적용하려고 플러그인을 리로드할 필요가 없습니다. 마음에 들어요. 제가 요청하는 모든 것을 Opus처럼 원샷으로 해결해 주는데 더 싸고 빠른 건가요? 잘 모르겠지만 사용 제한도 그렇게 빨리 줄어들지 않네요.“been using it since it is up and live and man... it is doing it's job like... can code Rust plugins and server managemenet panel and connect plugin settings with panel so you don't need to deal with editing .json files and reloading the plugin to activate changes... I liked it. one shotting everything I asked for like Opus but cheaper or faster? I dk... and usage limit is not dropping that fast”
드디어 새로운 Sonnet 모델이 나왔네요“Finally a new sonnet model”
GLM-5.2와 Sonnet 4.6으로 기사를 다시 써봤습니다. LLM은 비결정적이라 결과가 완전히 다르더군요. 하지만 GLM-5.2는 수작업으로 수정해야 할 미세한 실수가 많았던 반면, Sonnet은 두 번째 시도에서 모든 실수를 찾아내고 수정했습니다. 기획이나 코딩에서도 비슷한 상황이었어요. GLM-5.2는 서류상으로는 좋아 보이지만 실제 사용 결과는 달랐습니다. 제가 Claude나 GLM-5.2의 대변인은 아니지만... :) 2022년 11월부터 매일 LLM 모델을 사용해 온 결과, 모든 일반적인 테스트 결과는 각자의 프로젝트에서 직접 확인해 봐야 한다는 것을 깨달았습니다.“I have tried to rewrite an article with GLM-5.2 and with Sonnet 4.6. Completely different results as LLM is non-deterministic. But GLM-5.2 made a lot of subtle mistakes that needed to be corrected by hand. On the opposite, Sonnet found and corrected all mistakes in the second round. Similar situation was with planning and coding. GLM-5.2 seems to be good “on paper” but the real usage results was different. And I am not an attorney for Claude or GLM-5.2… :) But as I’ve been using LLM models daily since Nov 2022 I have realized that all common tests have to be confirmed in your project - there is no “one model rules them all” - you need to dig out a specific model from that LLM haystack with thousands of models. Benchmarks help but they start to be similar to fuel consumption specs in car ads - real consumption is different for everybody :)”
지금까지 짧게 테스트해 본 결과, 4.6보다 결과물도 좋고 속도도 약간 더 빨라서 확실히 업그레이드되었다는 느낌을 받았습니다.“The limited time I've had to test it so far, it's given me better results than 4.6 and a little quicker, so I think it's a noticeable step up.”
> 왜 이런 걸 자랑하는 걸까요? 마치 사람들이 사이버 보안 작업에 모델을 쓰고 싶어 한다는 걸 알면서도 의도적으로 그 능력을 거부하는 것 같습니다. Anthropic이 여기서 정확히 뭐라고 말하길 원하시나요? "전 세계에 저렴하게 제공하려는 이 모델이 해킹을 정말 잘합니다"라고요? Sonnet이 사이버 보안에 형편없다고 말하는 것이 여러 좋지 않은 선택지 중에서 그들이 할 수 있는 가장 합리적인 말입니다.“> Why would they brag about something like this? It's like they know people want to use models to perform cybersecurity tasks yet knowingly deny them the ability. What exactly do you want Anthropic to say here? "This model, the one we are about to give to the entire world for cheap, is really good at hacking"? Saying Sonnet is terrible at cybersecurity is the most reasonable thing they can say, out of a lot of bad options.”
이 시점에서는 무능한 미국 정부에 본때를 보여주기 위해서라도 Anthropic이 오픈 소스로 전환했으면 좋겠네요.“At this point I want Anthropic to go open source just to give the middle to the braindead US government.”
Fable을 다시 돌려줄 방법이나 찾아내는 게 좋을 거예요 😂“They better figure out how to get us Fable back😂”
fix: Claude Sonnet 5 적응형 사고(adaptive thinking) 요청 지원 — #98386 해결 관련: #98254 ## 해결된 문제 Claude Sonnet 5를 선택한 사용자가 기본, 적응형, 낮음 또는 꺼짐 사고 설정을 사용할 때 Anthropic API 요청 실패가 발생하거나 요청된 사고 모드가 손실되는 문제를 수정했습니다. Claude Sonnet 5는 Anthropic 적응형 사고를 사용합니다.“fix: support Claude Sonnet 5 adaptive thinking requests — Closes #98386 Related: #98254 ## What Problem This Solves Fixes an issue where users selecting Claude Sonnet 5 could hit Anthropic API request failures or lose the requested thinking mode when using default, adaptive, low, or off thinking settings. Claude Sonnet 5 uses Anthropic adaptive thinking se”
fix(model-core): claude-sonnet-5를 GA 1M-컨텍스트 모델로 인식 (fixes #5788) — ## 요약 1M 토큰 컨텍스트 창을 지원하지만 OMO 게이트가 이를 인식하지 못해 Anthropic 제공자의 기본값인 200K로 되돌아갔고, 압축(compaction)이 약 5배 정도 너무 빨리 트리거되었습니다. ## 원인 [ ] 1M 창을 두 개의 정규식으로 제어하는데, 어느 쪽도 일치하지 않았습니다. 첫 번째 브랜치는 다음에 즉시 오는 것을 요구하는데...“fix(model-core): recognize claude-sonnet-5 as GA 1M-context model (fixes #5788) — ## Summary has a 1M-token context window, but OMO's gate did not recognize it, so fell back to the 200K default on Anthropic providers and triggered compaction roughly 5x too early. ## Root Cause [ ]( gates the 1M window on two regexes: matches neither: the first branch requires immediately after ,”
> 그리고 Opus 4.8은 통과율이 더 높으면서도 여전히 더 저렴합니다. Opus만큼 텍스트를 쏟아내지 않는 한 의심스럽네요. Opus 4.8은 말 그대로 텍스트를 토하듯이 내뱉습니다. 장기적으로 볼 때, 특히 여기저기서 캐시 미스가 발생하면 비용의 대부분은 그 추가된 컨텍스트에서 나옵니다.“> And Opus 4.8 is still cheaper for a higher pass rate Unless it spams as much as Opus, I doubt it. Opus 4.8 literally spams text like puke. On a longer run especially if you get cache misses here and there the bulk of the cost is all the extra context it adds.”
서비스가 시작된 이후로 계속 써보고 있는데... 정말 제 역할을 톡톡히 하네요. Rust 플러그인과 서버 관리 패널을 코딩할 수 있고, 플러그인 설정과 패널을 연결해 줘서 .json 파일을 수정하거나 변경 사항을 적용하려고 플러그인을 리로드할 필요가 없습니다. 마음에 들어요. 제가 요청하는 모든 것을 Opus처럼 원샷으로 해결해 주는데 더 싸고 빠른 건가요? 잘 모르겠지만 사용 제한도 그렇게 빨리 줄어들지 않네요.“been using it since it is up and live and man... it is doing it's job like... can code Rust plugins and server managemenet panel and connect plugin settings with panel so you don't need to deal with editing .json files and reloading the plugin to activate changes... I liked it. one shotting everything I asked for like Opus but cheaper or faster? I dk... and usage limit is not dropping that fast”
더 유능해졌다"는 건 일주일에 몇 가지 작업만 해도 사용 제한에 걸린다는 뜻이겠네요 😆“"More Capable" mean a couple tasks a week and you've hit usage limits 😆”
조만간 우리 모두 비싸서 최고급 모델은 쓰지도 못하게 될 겁니다.“Soon we'll all be priced out of the best models”
Fable 5는 제가 가진 21년산 글렌피딕 같을 거예요. 절대 손도 안 대겠죠. 저는 제 Haiku, 즉 싼 맥주나 홀짝일 겁니다. ㅋㅋ.“Fable 5 will be like my bottle of 21 year Glenfiddich. Never touched. I will sip my Haiku a.k.a cheap beer. Lol.”
잘 모르겠네요. 그 프롬프트 3개에 비용이 얼마나 들었나요? Fable을 테스트하기엔 적절한 유형의 테스트가 아닌 것 같습니다.“I don't know, how much did it cost for those 3 prompts. I feel like these aren't the correct types of tests for Fable”
GPT 5.5 medium이 벤치마크도 더 좋고 가격도 더 싼데... Sonnet 5도 좀 더 저렴하게 만들었으면 좋았을 텐데요!“GPT 5.5 medium benchmarks better and cheaper... I wish they made sonnet 5 cheaper!”
Fable을 써봤는데, 웹사이트 HTML UI도 다 못 끝냈는데 몇 분 만에 유료 세션 제한에 걸렸습니다. 20달러 플랜에서는 쓸모가 없네요.“I tried fable, not even completed the html ui for my website, and hit my paid sessions within minute. useless on $20 plan.”
무슨 소리를 하는 건지 모르겠네요. 거의 비슷한 가격에 Opus 4.8 low로 더 좋은 결과를 얻을 수 있는데 왜 medium에서 Sonnet 5를 쓰나요? 그리고 Opus 4.8 med와 high가 Sonnet high와 xhigh보다 더 좋고 저렴합니다.“i dont know what your are talking about. why chosing sonnet 5 on med, when you can have better results with opus 4.8 on low for nearly the same price? And Opus 4.8 on med and high is better and cheaper than sonnet on high and xhigh.”
확실히 실망스럽습니다. 중요한 일에는 Opus 4.8을 쓰고, 중요하지 않은 일에는 Haiku를 쓰겠어요.“Definitely a let down. Opus 4.8 for everything that matters and Haiku for things that don't.”
로컬 AI를 쓰고 개인 데이터를 보호하겠다고 1,000만 원짜리 PC랑 맥 스튜디오를 잔뜩 사놓고서는, 이제 와서 Anthropic한테 자기 신분증 정보를 넘겨주고 있다니요.“You buy a 10k pc and a bunch of mac studios and say that it's for local AIs and you want personal data and now you say you are giving away your ID to Anthropic.”
Deepseek은 모든 것을 5배나 저렴하게 만들었는데, 얘네들은... 에휴...“Deepseek made everything 5 times cheaper, and these... ugh...”
그래프는 각 게시물의 추출 샘플(n≤30) 기반
Zubair Trabzada | AI Workshop
WorldofAI
AsapGuide
Alex Finn
WorldofAI
Brock Mesarich | AI for Non Techies
WorldofAI
Aura Labs
Bijan Bowen
Chase AI
IndyDevDan
Teacher's Tech
marinesebastian
rickylabs/netscript
SebaBoler/vanguard
SebaBoler/vanguard
boserh/garmin-coach
code-yeongyu/oh-my-openagent
BHarper77/brady-cli
link-assistant/hive-mind
YiftachCohen/token-tab
stevencarpenter/token-auditor
khanhphan1311/git-ralph
a5c-ai/babysitter
codemie-ai/docs
earendil-works/pi
outofrange-consulting/omp-dev-team
hank9999/kiro.rs
kirodotdev/Kiro
sepivip/SeekerClaw
earendil-works/pi
nechmads/hot-metal
timenco/dialectic-pr
cmudco/dare-backend
openclaw/openclaw
robinebers/openusage
danielnogueira8/LinkedInViralPostsSwipeFile
polyant-ai/polyant
tensorblock
echos-keeper
XueyingJia
XueyingJia
mradermacher