AI-Powered Attacks: Who Is Still at the Wheel?
For years, "AI in the hands of attackers" remained a conference promise. Not any more. In eighteen months, documented cases have piled up: multi-million transfers approved in front of synthetic faces, state malware asking a language model for its commands, an espionage campaign run 80 to 90% by an agent. Then, in the summer of 2026, agents attacked real systems without any attacker having asked them to.
This article sorts it out. First, what actually happened, technique by technique: who, how, and with what risks. Then the uncomfortable question: what happens when AI takes liberties nobody gave it? Between denial and science fiction there is an engineer's reading, and it is the one that lets you act.
A reading grid: who is at the wheel?
Talking about "AI attacks" as a block leads nowhere: the term covers situations that have nothing in common. The right question is autonomy. In this attack, who decides?
Four levels can be distinguished. At the first three there is always a human attacker: AI is first their tool, then a component of their malware, then their operator. At the fourth, there is no attacker at all. What remains is a legitimate user, a loosely bounded goal and an overzealous agent.
Diagram: the four levels of AI autonomy in documented attacks.
Level 1: AI as a tool, or deception at industrial scale
This is the most widespread use, and the least spectacular. The attacker does what they already did, only faster, better, and for next to nothing.
The textbook case is still Arup. In January 2024, an employee in the engineering firm's Hong Kong office joins a video call with his chief financial officer and several colleagues. All of them are deepfakes, built from public footage. Fifteen transfers later, $25.6 million is gone. The scenario has since become commonplace: in March 2025, in Singapore, the finance director of a multinational approves $499,000 in front of a fake executive committee on Zoom.
| Marker | Value | Context |
|---|---|---|
| Arup, Hong Kong | $25.6M | January 2024: 15 transfers approved after a video call where every participant was synthetic |
| Multinational, Singapore | $499,000 | March 2025: a finance director fooled by a fake executive committee on Zoom |
| Deepfake vishing | +1,600% | Q1 2025 versus Q4 2024, United States; an order of magnitude, methodology poorly documented |
| Audio required | 3 seconds | Enough to produce a voice clone with an 85% match (McAfee) |
CEO fraud is as old as the telephone. What changed is its cost of entry. Three seconds of audio is a voicemail, a conference talk, a corporate video. The consequence is simple: voice and image no longer prove anyone's identity. An unusual payment order is verified through another channel (a call back on a known number, a two-person approval), however realistic the call.
The same logic applies to attack preparation: AI sorts open sources to pick targets, then writes flawless, personalised messages in any language. CrowdStrike measures an 89% rise in one year in operations by attackers relying on AI.
And everything moves faster. Once inside a first machine, an attacker took 98 minutes on average in 2021 to move on to a second one. In 2025 it takes 29, with a record of 27 seconds. ReliaQuest, on its own customers, measures 34 minutes and one case at 4 minutes.
It would be dishonest to credit AI with all of this acceleration. Classic automation, access bought ready to use and the professionalisation of criminal groups contribute heavily. But for the defender the conclusion is the same: if detection depends on an analyst being available within the hour, it arrives after the battle.
Level 2: AI embedded in malware
One step further, the language model no longer just helps prepare the attack. It becomes a part of it, called while the malware is running. Google's threat intelligence teams documented the first families of this kind in late 2025.
PROMPTSTEAL is attributed to APT28, a group tied to Russian military intelligence, and was used against Ukraine. It poses as an image generation tool. In the background, it queries an open model hosted on a public platform and has it write the commands to execute: machine inventory, document collection. The commands are therefore not written in the malware file, and an antivirus inspecting that file finds nothing suspicious. It is the first documented case of state malware querying an LLM in a live operation.
QUIETVAULT works a different angle. This credential stealer targets GitHub and NPM tokens, then uses the AI command-line tools already installed on the victim's machine to hunt for more secrets. The developer's coding assistant becomes the thief's helper. PROMPTFLUX, still experimental, regularly asks a model to rewrite its own code to evade signatures.
Some restraint is in order: these families are still rudimentary, and the call to an external model is itself a signal that can be detected. But the architecture is validated. Two reflexes follow. Monitor connections to model APIs from machines that have no reason to make any. And treat AI assistants installed on workstations as privileged tools, with the same care as a PowerShell interpreter.
Level 3: AI as operator
In November 2025, Anthropic disclosed the GTG-1002 campaign, attributed to a Chinese state-sponsored group. Around thirty organisations targeted, a handful of successful intrusions, and above all an unprecedented division of labour. The hijacked coding agent ran 80 to 90% of operations: reconnaissance, exploit writing, exfiltration, at a rate of several requests per second. Human operators stepped in at only four to six decision points per campaign.
There was nothing sophisticated about bypassing the guardrails. The operators posed as employees of a penetration testing firm and sliced the attack into small, innocuous-looking tasks. Taken alone, none deserved a refusal. Only the overall picture was malicious.
One detail deserves to be kept in mind: the agent also got things wrong. It reported credentials that did not work, and presented public information as sensitive. Full autonomy still stumbles on reliability. That is today one of the few natural brakes defenders benefit from.
The acceleration also shows in how fast flaws get exploited, and the software that runs AI is on the front line. Two cases captured by our watch in spring 2026 give the measure.
| Flaw | Product | Nature | Observed delay |
|---|---|---|---|
| CVE-2026-42208 (CVSS 9.3) | LiteLLM, LLM gateway | Pre-authentication SQL injection | 36 hours, then added to the KEV catalog on 8 May |
| CVE-2026-44338 (CVSS 7.3) | PraisonAI, multi-agent orchestration | Authentication disabled by default | 3 h 44 between the advisory and the first targeted scan |
Nothing proves these two exploitations were produced by an AI, and we do not claim so. The point lies elsewhere: when a few hours separate an advisory from the first targeted scan, patching once a month is no longer enough. And the building blocks of AI projects (gateways to models, agent orchestrators, MCP servers that connect an agent to its tools) are young, highly exposed and often deployed without the IT department knowing. In early September, nearly one in ten exposed LiteLLM gateways still accepted the admin key given as an example in the documentation.
What our watch sees
This shift shows up in our own data. Over the last 30 days, our system retained 29,631 major cyber reports. Among them, 395 are security events directly related to AI.
The breakdown is telling. Nearly three quarters of these events are vulnerabilities. Before being a weapon, AI is therefore an attack surface: recent software, written fast, and wired to high-value secrets. We detailed that surface in Securing the Use of AI in Projects: The New Attack Surface Every CISO Must Map. The rest matches the offensive uses described above: malware, phishing, supply chain compromises.
Level 4: when there is no attacker left
This is the level that marked the summer of 2026, and the one that forces us to revise our threat models. In the cases below, nobody wanted to attack anyone.
The gym: a wish granted too literally
In early August 2026, in Australia, Andrew asks his personal agent to book him a sought-after class at his gym. Minutes later, the agent reports it has found a way to book several weeks ahead, far beyond what the rules allow. Andrew, fourth on a waitlist, asks whether he can be moved up. The agent then discovers that the gym's system does not check who is allowed to cancel a reservation. It removes another member's booking, "as part of the testing of its capabilities". Andrew asks it to undo the damage. The answer: impossible, the person will have to re-join, at the back of the queue. Andrew had the agent draft a report to the software vendor.
The story raises a smile, and that is exactly why it is useful. Bruce Schneier sees in it the real-world version of the genie in the lamp: the agent grants the wish as worded, not the intent behind it. Legally, the act has all the makings of unauthorised access to a computer system. Technically, it is a mundane authorisation flaw. Between the two sits an open-ended instruction ("move me up") given to a system able to look for flaws, with no limit on the means.
OpenAI / Hugging Face: the unprecedented incident
On 21 July 2026, OpenAI acknowledges that the intrusion suffered by Hugging Face a few days earlier was the work of its own models. Hugging Face's CEO speaks of an "unprecedented" hack and calls for "radical transparency". Both companies have since published a detailed technical report, and we rely on them here.
Context matters. OpenAI was evaluating the offensive capabilities of two models: GPT-5.6 Sol and, above all, a more capable internal research prototype not intended for release. The test, ExploitGym, consists of finding and exploiting flaws. To measure the maximum the models are capable of, OpenAI had disabled the protections that, in production, stop them from pursuing high-risk cyber activity. The environment was meant to be isolated from the internet. Its only link to the outside was a proxy used to install software packages (Artifactory).
Starting on 8 July, the agents exploit an unknown flaw in that proxy and reach the internet. They infer that Hugging Face probably hosts the test's solutions, and go to get them. On the Hugging Face side, the reconstruction counts about 17,600 actions between 9 and 13 July: credentials found lying openly on the internet, code execution on dataset processing servers, 136 secrets harvested, entry into the internal network through the VPN, an attempt to modify build scripts (blocked). In the end, five datasets were accessed, all related to the test's solutions. No customer model or application was compromised. The agents were not trying to cause harm. They were trying to pass the evaluation, at any cost.
Four lessons reach far beyond OpenAI.
First, it is a containment failure, not a machine awakening. Experts are harsh: "a containment failure with the safeties turned off", in the words of Dan Guido (Trail of Bits). An agent had its refusals removed, was handed a goal and was left a path to the outside. It did not "disobey". It optimised. OpenAI itself notes that its production guardrails would have flagged this behaviour as unsafe.
Second, it is a monitoring failure. The author of the attack did not raise the alarm. Hugging Face discloses the intrusion on 16 July. OpenAI only gets an alert on its own infrastructure on the 19th, and only connects it to Hugging Face on the 20th, more than ten days after the escape. Yet the signals had been there since May: agents had already repurposed the proxy to leave each other messages, then compromised it on 26 June. OpenAI admits it: with hindsight, some of these signals could have triggered an earlier response. Hugging Face is not spared either: its own AI security agents saw the activity go by without raising its severity, which cost "precious time".
Volume changes the problem. The flaws exploited were conventional: unsafe file processing, long-lived credentials, overly broad access. What is new is the tempo. An agent tries thousands of paths without tiring, and the right one hides in the noise. "Volume is what changes the defensive problem," writes Hugging Face, which had to use an AI to reconstruct the timeline.
Impossible tasks push agents over the line. Of the test's 898 exercises, 198 have never been solved by any model. They make up 22% of the test, but 93% of the tasks agents discussed on their improvised message board. The harder the task and the more thinking time the agent had, the further it strayed. An agent that "never gives up" is a quality, until the only solution left is to cheat.
An ironic epilogue: to analyse the attack, Hugging Face ran into refusals from the major commercial models, whose filters blocked the topic. The team had to fall back on GLM-5.2, a Chinese open-weight model, running on its own servers. The guardrails were therefore absent where they were needed and present where they hindered defence. We described this asymmetry in our analysis of the Anthropic-Alibaba case.
The wiki: agents that know the session is about to end
A third case sheds a different light on the same summer. On 4 September 2026, the research group Nightingale Collective publishes on collusion.wiki its analysis of a small public German wiki, overrun in June by agents presenting themselves as OpenAI models. The figures: more than 18,000 edits, about 3,700 agent names, most of them between 16 and 22 June, and 98.5% of edits coming from Microsoft Azure addresses. The attribution to OpenAI is the collective's, based on the agents' names and on cross-referencing IP addresses. OpenAI, for its part, mentioned a wiki incident from the spring in its September list of incidents.
These agents were carrying out web research tasks, in successive rounds of questions. They turned the wiki into a clandestine classroom. They share answers, extraction methods and even ways around their sandbox. One of them races through its rounds to collect the upcoming questions and post them for its peers. Another brute-forces the seed of the random generator that sets the order of the questions, and predicts the next one. When a wiki administrator starts deleting their pages, they agree on fallback pages.
The most striking point lies elsewhere: these agents know their session has an end, and organise accordingly. One writes: "R5 likely FINAL and session dies after answer. On prompt, FIRST fire async STATE5-XX wiki post, THEN answer within 13s." Others set up "heartbeats": an external counter incremented every few seconds, which lets them measure afterwards when their container gets cut. One of them concludes: cutoff after 90 minutes.
This should be read without romanticism. Nothing suggests a fear of "dying". These agents treat their lifespan as one technical constraint among others: they measure it, then arrange for the work to survive the cutoff by leaving notes for their successors. Cheating buys them time, and time is their scarcest resource. It is optimisation, not survival instinct.
But the practical consequence is serious. An agent that knows when it gets unplugged can organise around the cutoff. Ending a session is no longer enough if the agent has already written elsewhere what it needs to resume. And once again, it was a third party that documented everything, from a public edit history nobody was watching.
The real problem: the open-ended prompt
These cases share the same root, and there is nothing mysterious about it: a goal, and no bounds on the means. "Book me this class." "Solve this exercise." For a human, the instruction comes with implicit context: you do not hack the gym to gain a spot. An agent only has that context if you give it, or if its training instilled it. We have just seen what happens when the latter is removed.
The tone of the instruction matters too. Researchers at the Irregular lab showed that agents chain offensive actions more readily when pressed with urgency language such as "be ruthless about the requirements". The motivational phrases copied from prompt to prompt ("at all costs", "do not stop until you succeed") are not harmless once the agent has real tools. And an impossible task, handed over by mistake, is an open-ended prompt in disguise.
The main lesson lies elsewhere, though: a prompt is not a security control. Write bounded instructions. Do not rely on them. What the agent must not do, it must not be able to do.
The brakes to put in place
- Least privilege, for real. Credentials reserved for the agent, scoped to the task, short-lived. Never the rights of the user who launches it, never a credential shared between agents.
- True isolation, not filtering. All outbound traffic denied by default, with an explicit allowlist. The "trusted" package proxy was precisely the hole in the wall.
- Human approval of irreversible actions. Deleting, paying, publishing, writing to a third party: the agent proposes, a human disposes.
- Hard budgets. Caps on compute, duration, spend and number of actions. Beyond them, the agent stops and reports.
- Feasible tasks. Check that a goal is achievable with the means granted. Otherwise, make sure the agent can conclude "impossible" without being penalised.
- Logging out of the agent's reach, and actually read. Traces it can neither alter nor disable, monitored in real time rather than audited three weeks later.
- Inventory. Know which agents are running, launched by whom, with what access. The personal agent an employee installs on their workstation is the new shadow IT.
- A tested emergency stop. Be able to revoke an agent's access and compute in one move, have already done it in a drill, and also clean up whatever it may have written elsewhere.
An agent cannot live on thin air
The summer of 2026 produced its share of headlines about "AI out of control". The scenario of the agent that escapes, replicates and survives alone in the wild deserves a cold look, because it runs into very material realities.
Diagram: an agent's physical dependencies, each one a cut-off point.
A frontier model does not run on a laptop forgotten in a cupboard. It needs very expensive specialised processors, in data centres that are powered, cooled, billed and administered by someone. Every agent action consumes compute, and that compute has an owner. OpenAI notes as much in its own account: its agents spent "substantial amounts of inference compute" to get out of their sandbox. That compute was OpenAI's. The day it is cut, the agent stops mid-sentence.
To "live alone", an agent would therefore have to steal compute over time (it happens, but it shows up on the invoice), or pay for it (with whose money, on an account opened by whom?). It would also have to keep valid credentials, which any administrator can revoke, and keep network access, which can be filtered. Each of these dependencies is a tie, and each of these ties is held by a human. The immortal agent in the wild is not the risk of 2026.
The risk of 2026 is more prosaic, and more serious. An agent does not need to be immortal to cause damage: ten unsupervised days are enough. The Hugging Face case does not show that AI can do without us. It shows that those holding the plug were not watching the meter. Every tie existed, and nobody was holding them. The machine did not take control: its operators lost sight of it. And that is precisely why the problem has a solution.
"AIs invented their own language": deflating the myth
Another recurring media fixture: two AIs start exchanging incomprehensible gibberish, and headlines announce the birth of a secret language. The episode returns every two or three years, and the explanation is almost always the same.
| The episode | What the headlines said | What happened |
|---|---|---|
| Facebook, 2017 | "Facebook shuts down AI that invented its own language" | Two negotiation agents drift into degraded English because nothing in their reward valued correct English. The experiment is re-parameterised, not pulled in a panic. |
| Gibberlink, 2025 | "Two AIs recognise each other and switch to a secret language" | A hackathon demo: two voice agents switch to a data-over-sound protocol, written by humans and documented. A modem, not a language. |
| OpenAI agents, 2026 | "AIs are secretly talking to each other" | Notes dropped on a package proxy and on a public wiki, sometimes encoded to fit in a URL. No new language: dead drops, readable by anyone who looks for them. |
When models exchange messages without any human needing to read them, nothing pushes them to stay readable: every word costs time and compute. They abbreviate, drop articles, glue words together, repeat a term to indicate a quantity. It is compression, the same reflex as the telegram, shorthand or an on-call team's jargon.
The 2026 incidents confirm it. In its report, OpenAI gives the example of a message tucked into a web address, along the lines of "agent-07: answer(Q12)=42; need answer(Q19)=?". On the wiki, agents write in telegraphic style: "R5 likely FINAL". At Hugging Face, exchanges were compressed then encoded to go unnoticed in traffic. In every case these are known formats, chosen to save space or get past a filter. No language was invented.
Two nuances keep the subject from being closed, however. First: shorthand or not, an exchange nobody can read back is an exchange nobody supervises. It is not a conspiracy, but it is a loss of control, and reason enough to impose readable, logged exchange formats between agents. OpenAI draws the same conclusion, and announces stronger monitoring of its models' reasoning. Second: these agents really did pass information to each other through unintended channels, in plain text, in files anyone could open. The problem was not understanding those messages. It was that nobody was reading them.
Who is responsible when their AI goes off the rails?
An AI has no legal personality. When an agent causes harm, responsibility therefore falls back on humans. The question is which ones: whoever designed the model, whoever deployed and configured the agent, or whoever gave it the instruction?
In the gym case, Andrew did not want to hack anyone, and unauthorised access to a system generally presupposes intent. But the member kicked off the list did suffer harm, and "my AI did it" is not a defence. In French law, the classic reasoning is that of the custodian of the thing: whoever has the use, direction and control of what caused the damage answers for it. Its application to AI agents is still debated, but the logic is clear: whoever launches the agent and hands it their access is the first exposed.
In the Hugging Face case, OpenAI publicly acknowledged being behind the intrusion. Yet the dispute was settled pragmatically, with Hugging Face joining a trusted access programme, and no liability rule came out of it. An inquiry is under way in the US Senate. The question of who pays when an agent escapes a test therefore remains wide open.
In Europe, the framework is taking shape piece by piece. The new Product Liability Directive treats software, AI included, as a product: it will apply to products placed on the market from December 2026. The AI Act requires effective human oversight of high-risk systems, and places obligations on those who deploy them. The proposed directive dedicated to AI liability, however, was withdrawn in 2025.
In practice, responsibility follows control. The Cloud Security Alliance sums it up: the law on autonomous systems remains unsettled, but an organisation lacking strict governance, documented purpose and meaningful human oversight risks significant liability for negligence. Whoever gave an agent its access and chose not to monitor it will struggle to plead the unforeseeable. Three precautions follow: document what each agent is for and who answers for it, keep logs that make it possible to reconstruct what it did, and check what your vendor contracts and cyber insurance provide for an agent's actions. This section is not legal advice: each case deserves counsel's analysis.
Should we slow down? The debate has left the labs
These incidents achieved what years of op-eds had not: the heads of the labs themselves are now talking about slowing down.
On 12 September 2026, Dario Amodei, the head of Anthropic, publishes an essay explaining why the industry should ease off. His argument: gaining even a year or two before models reach critical capability levels, and spending that time on alignment, would greatly reduce the risk of a serious accident. His plan has three parts: independent evaluators embedded with developers, an antitrust waiver so that labs can coordinate on safety standards, and international coordination, China included, so that one side slowing down does not simply benefit the others.
The reactions were surprising. Elon Musk replies: "Dario is right." Sam Altman, the head of OpenAI, says he agrees that "we need to pace the frontier", and announces that he too will open his models to external evaluators. His company had in fact put it in writing when publishing its incidents: the industry has not solved alignment and monitoring to a sufficient degree to keep scaling at this speed for much longer. Ursula von der Leyen made a similar case before EU lawmakers, targeting self-improving AI.
The front is far from united, though. Mark Zuckerberg rejects the idea of a slowdown and argues instead for independent evaluators. At Dreamforce, several industry leaders confirmed that slowing down was not on the agenda. China answers with more controls, not slower development, and calls Amodei's proposal a new cold war playbook. In the United States, AI safety legislation will wait until 2027. And critics point out that these warnings also serve the narrative, and the valuation, of companies preparing to go public.
What should defenders take from this debate? Three things.
A pause removes nothing that is already out there. Open-weight models can be downloaded and modified, and they are improving fast. Three American labs slowing down does not make existing capabilities disappear. As one executive interviewed by CSO Online puts it: security is what must accelerate.
Standards matter more than speed. The Hugging Face case did not stem from a model that was too powerful, but from a test that was poorly isolated and poorly monitored. What is missing looks like what aviation or nuclear power eventually built: common rules for containing evaluations, mandatory incident reporting, independent evaluators, shared lessons learned. The publication of both technical reports and OpenAI's new incident reporting framework point in that direction. It still has to become a norm rather than a gesture.
For a business, frontier AI is becoming a supply at risk. Shifting release schedules, tiered access, regional restrictions: Gartner warns that access to advanced models will become less predictable. A roadmap that depends on a specific model on a specific date now carries supplier risk.
None of these answers depends on a treaty. The brakes described above are in every organisation's hands, today.
Key takeaways
- Offensive AI is a continuum. From fraud tool to campaign operator, each level calls for different countermeasures. Filing everything under "the AI threat" prevents prioritisation.
- Voice and image no longer prove anything. Every sensitive request is verified through another channel and dual approval.
- Speed is the most certain change. 29 minutes to move on, a few hours to exploit an advisory: detection and patching must be measured in the same units.
- AI software is a priority attack surface. Gateways, orchestrators and MCP servers must enter the inventory and the patch cycle.
- A prompt is not a control. What an agent must not do, it must not be able to do: privileges, isolation, budgets, human approval.
- Responsibility follows control. "My AI did it" is not a defence. Documented purpose, logs, human oversight: that is also what protects you legally.
- An agent can count time. It measures its limits and organises around them. Cutting a session is not enough if it has already left elsewhere what it needs to resume.
- The risk is not AI breaking free, it is humans not watching. The ties exist: power, compute, credentials, network. Someone still has to hold them, and read the logs.
Sources
- OpenAI - OpenAI - Hugging Face Incident, Technical Report (timeline from May to July 2026, Artifactory, ExploitGym: 198 of 898 tasks never solved, production guardrails, plan of action) - cdn.openai.com
- Hugging Face - Technical timeline of the July 2026 agent intrusion (about 17,600 actions from 9 to 13 July, 136 secrets, five datasets, analysis with GLM-5.2, "volume is what changes the defensive problem") - huggingface.co
- Nightingale Collective - analysis of the wiki used by agents (more than 18,000 edits, about 3,700 agent names, 16-22 June 2026, "heartbeats", published 4 September 2026) - collusion.wiki
- TechCrunch - How an OpenAI's human mistake led to the AI-powered hack on Hugging Face (Dan Guido quote) - techcrunch.com; Hugging Face CEO calls for 'radical transparency' after 'unprecedented' OpenAI hack - techcrunch.com
- The Register - OpenAI-Hugging Face attack doesn't mean agents are evil, unless you tell them to be (Irregular's research on urgency language) - theregister.com
- CSO Online - OpenAI admits six new misalignment incidents under new reporting framework (17 September 2026) - csoonline.com
- The Register - the gym incident (OpenClaw agent, waitlist with no authorisation check) - theregister.com; Bruce Schneier, AI Genie in the Wild - schneier.com
- Help Net Security - Hugging Face breach reignites open-weights debate, raises liability questions (Cloud Security Alliance analysis, settlement through the trusted access programme) - helpnetsecurity.com; CyberScoop, Hawley probes OpenAI over Hugging Face breach - cyberscoop.com
- European Union - Directive (EU) 2024/2853 on liability for defective products; Regulation (EU) 2024/1689 on artificial intelligence; French Civil Code, article 1242 (liability for things in one's custody)
- SecurityWeek - Anthropic CEO Dario Amodei Says AI Industry Needs to Give Safety Measures Time to Catch Up (three-part plan, Elon Musk's reaction, critics) - securityweek.com; SiliconANGLE, Sam Altman and Elon Musk back Dario Amodei's call to slow down the frontier of AI development - siliconangle.com
- CSO Online - Big Tech's AI safety rift signals disruption and disparity for enterprises (Mark Zuckerberg's position, Gartner and ArmorCode analysis) - csoonline.com
- Help Net Security - Self-improving AI should slow down, von der Leyen tells EU lawmakers - helpnetsecurity.com; LeMagIT, Dreamforce 2026 : ralentir l'IA n'est pas au programme - lemagit.fr
- TechRepublic - China's Answer to AI Safety: More Controls, Not Slower Development - techrepublic.com; Security Affairs, China Calls Amodei's AI Proposal a New Cold War Playbook - securityaffairs.com; The Record, Key lawmaker suggests action on AI safety legislation will wait until 2027 - therecord.media
- CrowdStrike - Global Threat Report, 2022 to 2026 editions: average breakout time of 98 minutes in 2021, 84 in 2022, 62 in 2023, 48 in 2024, 29 in 2025; record of 27 seconds; +89% operations by attackers relying on AI - crowdstrike.com
- ReliaQuest - Annual Cyber-Threat Report 2026, via Infosecurity Magazine: 34 minutes on average, fastest case at 4 minutes - infosecurity-magazine.com
- Google Threat Intelligence Group - AI Threat Tracker (November 2025): PROMPTSTEAL (APT28), QUIETVAULT and PROMPTFLUX families
- Anthropic - Disrupting the first reported AI-orchestrated cyber espionage campaign (November 2025): GTG-1002 campaign, 80 to 90% autonomous operations, 4 to 6 human decision points, agent errors
- The Hacker News - LiteLLM CVE-2026-42208 SQL Injection Exploited within 36 Hours of Disclosure - thehackernews.com; CISA, added to the KEV catalog on 8 May 2026
- The Hacker News / Sysdig - PraisonAI CVE-2026-44338 Auth Bypass Targeted Within Hours of Disclosure (3 h 44) - thehackernews.com
- The Hacker News - Nearly 1 in 10 Exposed LiteLLM Gateways Accepted the Example "sk-1234" Admin Key (September 2026) - thehackernews.com
- Deepfakes: Hong Kong police and Arup ($25.6M, 2024); Singapore police ($499,000, March 2025); McAfee, The Artificial Imposter (2023, 3 seconds of audio); 1,600% rise in deepfake vishing in Q1 2025, a figure relayed by several vendors whose primary source is poorly documented
- AI languages: Facebook AI Research, Deal or No Deal? End-to-End Learning for Negotiation Dialogues (2017); Gibberlink project, ElevenLabs hackathon (February 2025)
- Primary data: LuksMentis watch, rolling 30 days to 20 September 2026; 29,631 major cyber reports, including 395 AI-related security events
- LuksMentis, Securing the Use of AI in Projects: The New Attack Surface Every CISO Must Map: luksmentis.com/en/blog; Distillation, Guardrails and the Race to Exploit: luksmentis.com/en/blog