Ask HN: Which frontier model can do code security reviews

I can't complain about Opus 5.5, more than extracted my money's worth out of it but the final stages require pen-testing and Claude's masters stomp on it every time it tries to help do security reviews [cyber]. Ever so often it sneaks in a security fix. Even found a race condition in haproxy by mistake and Claude took it upon itself to find the crash string. That went horribly bad. Each time they try to up-sell Mythos and say I have to go through a verification program that I am not permitted to go through.

Aside from the uncensored Qwen forks, which frontier models can do extensive code security reviews, security fixes? Ideally something close to the quality of the NCC Group. This is for my own hobby craft. Maybe this does not exist and that is fine too.

6 points | by Bender 15 hours ago

3 comments

  • samuelknight 11 hours ago
    My startup is a platform for automating pentest workflows. Reach out if you are interested!
  • bigyabai 15 hours ago
    GLM 5.3. Anthropic even made the mistake of comparing it to Mythos (lol): https://www.anthropic.com/research/glm-5-3-and-the-spread-of...

    > Like Claude Mythos Preview, GLM-5.3 has strong capabilities for autonomously building end-to-end cyber exploits. But GLM-5.3 is unlike other frontier models in that it has been released without meaningful safeguards to limit misuse. We find that attackers can bypass GLM-5.3’s safeguards between 64% and 100% of the time with simple techniques in our simulated tests.

    > We find that GLM-5.3 develops end-to-end exploits in 50 of 410 attempts. Claude Mythos Preview did so at a similar rate—in 56 of 410 attempts.

    • Bender 15 hours ago
      I may end up going that direction. I would ideally like to find something that is purpose built to do code pen-testing so I do not have to bypass anything. There are forks of other models built for this, maybe there is a fork of GLM too.

      Grok is telling me it can do code security reviews but I am not sure I believe it.

      • bigyabai 10 hours ago
        I've been using GLM 5.3 and 5.3 Flash to reverse-engineer some old firmwares, and it hasn't protested at all. Full 5.3 has a scary-good grasp of debugging assembly, QEMU and hex dumps in my experience, it would make for a good "sleuth" model to find potential vulns. 5.3 Flash is closer to 5.5 Sonnet/6 Luna capabilities-wise, but still smart enough to implement the easy fixes or steer your red team agent.

        It's been about a year since I switched from Claude Code to Z.AI, so I don't know what I'm missing out on. But I also don't really feel any FOMO, I'd rather support open model releases than amortize another gated-release model like Mythos.

  • segmondy 8 hours ago
    all of them