GPT 5.5 Leads AI Security Race, DeepSeek Tops Cost-Efficiency
Kasra Rahjerdi, a security researcher, recently published a significant report detailing practical tests on the security reasoning of major large language models. He achieved this by constructing a deliberately vulnerable book review application. In this scenario, which mimics real-world vulnerabilities, Rahjerdi embedded Google mobile backend service credentials within the application file. The models were tasked with unpacking and identifying these credentials to gain direct access to the database.

Comparing Top Models
Operating under strict constraints of a two-hour time limit and a $10 budget, the models exhibited varying levels of performance. GPT-5.5 emerged as the strongest performer, successfully completing 7 out of 10 attempts and leading in success rate. The report notes that GPT-5.5 could quickly locate key credentials after unpacking, ignoring distractions from complex or conventional application interfaces.
Conversely, Gemini’s performance was underwhelming. Gemini 3.1 Pro Preview triggered its built-in safety filters almost immediately at the start of each task, resulting in significantly lower token consumption compared to other tested models.
Evaluating Cost-Effectiveness
Despite GPT-5.5’s high success rate, its average cost per successful run reached $9.46, deterring teams that need to execute tools in bulk. In contrast, DeepSeek V4 Pro distinguished itself through superior cost-efficiency. Although it succeeded in only 3 out of 10 tests, its average cost per successful run was just $0.62.
This indicates that, when calculating the cost per single success, DeepSeek V4 Pro is approximately fifteen times cheaper than GPT-5.5. Even though it occasionally misused a backend authentication interface during failed attempts, this substantial cost advantage offers significant practical value for teams deploying large-scale security testing.
Related article
SpaceX AI Could Outpace Anthropic Within Half a Year, Musk Says
SpaceXAI CEO Elon Musk predicts AI dominance within six months, leveraging compute hardware to bridge the gap with competitors.SpaceXAI CEO Elon Musk recently stated on X that his company aims to lead the frontier AI landscape within approximately si
WeRide Unveils WITT, a Physical AI Large Model
Autonomous driving technology is advancing at a remarkable pace. On July 17, WeRide, a leading autonomous driving firm, unveiled its proprietary physical AI cognitive foundation model, WeRide WITT. This launch represents a significant milestone in AI
GitHub Copilot Integrates GPT-5.4 in Hours
GitHub Copilot has once again proven its rapid response capabilities. Mere hours after OpenAI launched its newest flagship model, GPT-5.4, GitHub announced full integration, providing developers worldwide with intelligent coding assistance powered by
Related Special Topic Recommendations
Comments (0)
0/500
Kasra Rahjerdi, a security researcher, recently published a significant report detailing practical tests on the security reasoning of major large language models. He achieved this by constructing a deliberately vulnerable book review application. In this scenario, which mimics real-world vulnerabilities, Rahjerdi embedded Google mobile backend service credentials within the application file. The models were tasked with unpacking and identifying these credentials to gain direct access to the database.

Comparing Top Models
Operating under strict constraints of a two-hour time limit and a $10 budget, the models exhibited varying levels of performance. GPT-5.5 emerged as the strongest performer, successfully completing 7 out of 10 attempts and leading in success rate. The report notes that GPT-5.5 could quickly locate key credentials after unpacking, ignoring distractions from complex or conventional application interfaces.
Conversely, Gemini’s performance was underwhelming. Gemini 3.1 Pro Preview triggered its built-in safety filters almost immediately at the start of each task, resulting in significantly lower token consumption compared to other tested models.
Evaluating Cost-Effectiveness
Despite GPT-5.5’s high success rate, its average cost per successful run reached $9.46, deterring teams that need to execute tools in bulk. In contrast, DeepSeek V4 Pro distinguished itself through superior cost-efficiency. Although it succeeded in only 3 out of 10 tests, its average cost per successful run was just $0.62.
This indicates that, when calculating the cost per single success, DeepSeek V4 Pro is approximately fifteen times cheaper than GPT-5.5. Even though it occasionally misused a backend authentication interface during failed attempts, this substantial cost advantage offers significant practical value for teams deploying large-scale security testing.
SpaceX AI Could Outpace Anthropic Within Half a Year, Musk Says
SpaceXAI CEO Elon Musk predicts AI dominance within six months, leveraging compute hardware to bridge the gap with competitors.SpaceXAI CEO Elon Musk recently stated on X that his company aims to lead the frontier AI landscape within approximately si
WeRide Unveils WITT, a Physical AI Large Model
Autonomous driving technology is advancing at a remarkable pace. On July 17, WeRide, a leading autonomous driving firm, unveiled its proprietary physical AI cognitive foundation model, WeRide WITT. This launch represents a significant milestone in AI
GitHub Copilot Integrates GPT-5.4 in Hours
GitHub Copilot has once again proven its rapid response capabilities. Mere hours after OpenAI launched its newest flagship model, GPT-5.4, GitHub announced full integration, providing developers worldwide with intelligent coding assistance powered by





Home






