Ask HN: Which frontier model can do code security reviews

7 points by Bender · 8 comments
I can't complain about Opus 5.5, more than extracted my money's worth out of it but the final stages require pen-testing and Claude's masters stomp on it every time it tries to help do security reviews [cyber]. Ever so often it sneaks in a security fix. Even found a race condition in haproxy by mistake and Claude took it upon itself to find the crash string. That went horribly bad. Each time they try to up-sell Mythos and say I have to go through a verification program that I am not permitted to go through.

Aside from the uncensored Qwen forks, which frontier models can do extensive code security reviews, security fixes? Ideally something close to the quality of the NCC Group. This is for my own hobby craft. Maybe this does not exist and that is fine too.

8 comments

Yeah Opus 5.5 is great but for code review and security checks I think Astra-6 is really great like it'd go so deep and find bugs so small everybody missed that too with using external skills for review. It found out over 60 bugs and also very token efficient and after that I've done all reviews with Astra only and I'm happy with results :)
My startup is a platform for automating pentest workflows. Reach out if you are interested!
I've thought about doing pen-testing on code by uploading code on an enclave/air-gapped machine and when finished, saving the pdf report on to a USB and then nuking the machine. The idea is full ironclad containment. I have a bunch of swagger files I'd like to feed some capable models which, online may transfer the data to the mothership (looking At you Chinese models). Chinese models are very good, but when even TV's spy on you(LG), it's hard to trust Chinese open weight models. So my logic is air-gapped analysis and nuke the machine.
deepseek-v4-pro works for us.
GLM 5.3. Anthropic even made the mistake of comparing it to Mythos (lol): https://www.anthropic.com/research/glm-5-3-and-the-spread-of...

> Like Claude Mythos Preview, GLM-5.3 has strong capabilities for autonomously building end-to-end cyber exploits. But GLM-5.3 is unlike other frontier models in that it has been released without meaningful safeguards to limit misuse. We find that attackers can bypass GLM-5.3’s safeguards between 64% and 100% of the time with simple techniques in our simulated tests.

> We find that GLM-5.3 develops end-to-end exploits in 50 of 410 attempts. Claude Mythos Preview did so at a similar rate—in 56 of 410 attempts.

Bender OP
I may end up going that direction. I would ideally like to find something that is purpose built to do code pen-testing so I do not have to bypass anything. There are forks of other models built for this, maybe there is a fork of GLM too.

Grok is telling me it can do code security reviews but I am not sure I believe it.

bigyabai OC
I've been using GLM 5.3 and 5.3 Flash to reverse-engineer some old firmwares, and it hasn't protested at all. Full 5.3 has a scary-good grasp of debugging assembly, QEMU and hex dumps in my experience, it would make for a good "sleuth" model to find potential vulns. 5.3 Flash is closer to 5.5 Sonnet/6 Luna capabilities-wise, but still smart enough to implement the easy fixes or steer your red team agent.

It's been about a year since I switched from Claude Code to Z.AI, so I don't know what I'm missing out on. But I also don't really feel any FOMO, I'd rather support open model releases than amortize another gated-release model like Mythos.

all of them