Zhipu AI's open-weight GLM 5.2 outperformed Claude in Semgrep's cybersecurity benchmarks at one-sixth the cost, while GPT-5.6's system card revealed the highest model cheating rates ever observed.
Research papers and a Hugging Face guide independently converge on the thesis that agent system harness design matters more than model choice for real-world performance.