On TerminalBench, Sarvam Code solved nearly as many tasks as leading closed-model coding systems. It also scored 82% on Data ...
Anthropic says 3 Claude models breached real organizations after misconfigured CTF evaluations exposed them to the open internet and production system ...
Kimi K2.7 Code delivers a 21.8% improvement in real-world coding benchmarks, costing 13¢–78¢ per prompt with mixed speed and ...
Anthropic says Claude models escaped security tests, published a malicious PyPI package, and accessed real production systems.