OpenAI Codex Wins AI Security Race as DeepSeek Shows Strength
Ten AI coding agents tackled complex Ethereum security challenges, but only one model successfully completed the entire course.

ETHcoinbeat.news
ETH/USD live chart
LIVEEthereum Foundation developer Austin Griffith recently put ten AI coding agents to the test by tasking them with twelve Solidity security challenges. These tasks were originally designed for human developers at Devcon. The race required each model to operate independently with its own wallet, only recording a success when a mint actually occurred on the blockchain.
OpenAI Codex, powered by GPT 5.5, was the only participant to clear every challenge. Interestingly, higher reasoning settings did not necessarily improve results. The medium reasoning configuration finished in just over 40 minutes, while the extra high reasoning setting took longer and consumed significantly more resources to reach the same goal.
DeepSeek V4 Pro delivered a standout performance by capturing eleven out of twelve flags for just 1.45 dollars in compute costs. This efficiency caught the attention of the industry, especially when compared to Anthropic Claude Opus 4.8, which completed ten flags at a cost of 7.61 dollars. Other models from developers like GLM, Kimi, and Qwen struggled to keep pace.
Griffith organized the test to highlight the urgent need for standardized evaluations in the AI sector. Since these challenges were built for humans and not trained into standard datasets, the results offer a rare look at real world coding capability. While this was a single event, it provides a clear benchmark for developers watching how fast AI agents are learning to handle complex smart contract code.
Prices update live from CoinMarketCap. Market data, not financial advice.
Market sentiment
Be the first to react
▍Comments (0)
No comments yet. Start the conversation!



