The cybersecurity landscape just experienced a massive reality check. For the past year, the tech industry has debated the hypothetical risks of generative AI falling into the hands of bad actors.
Now, an alarming intelligence report from Cisco Talos proves the threat is no longer theoretical.
By analyzing accidentally exposed chat logs and prompt histories, Cisco’s researchers caught threat actors red-handed using top-tier AI models to automate cyberattacks, hunt for software vulnerabilities, and rapidly develop highly effective malware.
The models implicated in the report are some of the most advanced on the market, including Anthropic’s Claude Code, OpenAI’s Codex, the AI-powered code editor Cursor, and Google’s Gemini.
This revelation exposes a harsh paradox at the heart of artificial intelligence development.
The exact same capabilities engineered to help legitimate security teams and software developers write cleaner code are actively being weaponized.
The cutting-edge tools designed to patch vulnerabilities are now being deployed to exploit them at unprecedented speeds.
How Hackers Bypass AI Guardrails with Simple Social Engineering
One of the most concerning takeaways from the Cisco Talos investigation isn’t the technical sophistication of the attackers, but rather the complete lack of it required to break AI safety filters
Commercial AI developers spend millions building complex guardrails to prevent their platforms from writing malicious code. Yet, researchers found that hackers rarely needed highly technical jailbreaks to bypass these restrictions.
Instead, they lean heavily on basic social engineering tactics directed at the AI itself. Threat actors routinely manipulate the models by framing their requests under false pretenses.
Common tricks include claiming they are part of an approved ethical hacking competition. Others falsely say they have permission to test a company’s network.
If Gemini or Claude refuses a request, attackers often start a new chat and try again. The fresh chat removes the previous context. That can help attackers bypass safety checks and continue unnoticed.
Nick Biasini, a senior technical leader at Cisco Talos, highlighted the incredibly difficult tightrope AI developers must walk. AI models must still support legitimate security researchers who test systems and uncover vulnerabilities.
If AI companies lock down the models too aggressively, they cripple the very defenders they intend to empower. Cybercriminals understand this gray area and exploit it. That shows built-in AI safeguards alone are not enough to stop skilled attackers.
Enterprise API Theft Fuels AI-Driven Cyberattacks
Beyond outsmarting the AI’s built-in conversational defenses, attackers are fundamentally shifting how they finance and execute these operations. Cisco found that hackers are avoiding the cost of expensive AI computing power.
Instead, they steal enterprise API tokens and hijack company accounts to use AI services at someone else’s expense.
By piggybacking on a target company’s authorized, paid AI access, threat actors can hide their activity in plain sight. They use stolen corporate resources to create malware and scan networks for weaknesses. In effect, the victim pays for the attack.
This makes attacks harder to detect because the requests come from trusted company accounts instead of anonymous free accounts.
The lesson is clear. Companies should not rely only on the built-in safety features of Claude, Codex, Cursor, or Gemini.
Basic prompt tricks and stolen credentials can bypass standard AI protections. Security experts recommend a stronger, layered defense instead.
Organizations must begin treating AI access like any other highly critical attack vector.
This means enforcing strict identity and access management, actively monitoring API usage for bizarre anomalies, and operating under the assumption that threat actors already possess the exact same AI capabilities as internal defense teams.
Source: Official The Times of India, "Cisco Warns Hackers Are Using Claude Code, Codex, Cursor and Gemini AI Models"




