Large language models may have a fundamental security weakness that makes it impossible to protect them completely from prompt injection and jailbreak attacks, according to research presented at the International Conference on Machine Learning (ICML) this month.

The paper, titled “ Prompt Injection as Role Confusion ,” argues that LLMs struggle to reliably distinguish between trusted instructions, user requests, external data, and their own reasoning. The researchers say this weakness is built into how current models process text, meaning better safety training alone may never eliminate the problem.

The finding has broad implications as LLMs are increasingly used in government systems, military applications, healthcare, online shopping, and other areas where incorrect or manipulated actions could have serious consequences.

The researchers found that LLMs do not always determine where a piece of text came from using the tags that are supposed to identify its role.

Chat systems typically separate different types of information into categories such as user instructions, assistant responses, system instructions, internal reasoning, and information retrieved from external tools.

However, the study found that models often judge these roles based more on the wording and style of the text than on the tags surrounding it.

In tests, changing the tags around a passage made little difference if the writing still resembled a particular role. Text written to look like the model’s own reasoning could therefore be treated as trusted reasoning even when it came from a user.

The team called one attack chain-of-thought forgery.

Instead of simply telling a model to ignore its rules, the attack presents fabricated reasoning written in a style that resembles the model’s own internal notes. The model can then interpret the fake reasoning as something it has already concluded and act on it.

Using this approach, the researchers were able to bypass safeguards on several models and make them provide information they had been trained not to give, including instructions related to illegal drugs and interfering with an aircraft navigation system.

One test used an irrelevant detail about the user’s clothing together with fake internal reasoning claiming that the detail made an otherwise prohibited request acceptable. OpenAI models tested by the researchers then complied with the request.

The researchers reported an average attack success rate of about 60% across multiple models.

The attack won an OpenAI red-teaming competition in August 2025. OpenAI later said an early version of its automated red-team model, GPT-Red, independently discovered a similar technique that it calls Fake Chain-of-Thought.

AI companies regularly use human red teams and automated systems to find new attacks before models are released.

Once researchers discover a successful jailbreak or prompt injection, developers can train later models to recognize and resist similar attacks.

Jasmine Cui, one of the paper’s authors, argues that this approach has a fundamental limitation. Developers can teach a model many examples of what it should not do, but they cannot anticipate every possible attack.

The researchers say the deeper problem is that role boundaries exist clearly in the software interface but are not represented as equally firm boundaries inside the model itself.

Their experiments found that the more strongly a model confused one role with another, the more likely an attack was to succeed.

The researchers acknowledge that some of the models examined in the study were released last year and that newer models have improved.

OpenAI, for example, says training with GPT-Red has sharply reduced the success of Fake Chain-of-Thought attacks on newer models. Attacks that succeeded more than 95% of the time against GPT-5.1 fell below 10% against GPT-5.6 Sol, according to the company.

Florian Tramèr, a computer scientist specializing in LLM security at ETH Zürich, said model developers are combining training with monitoring and other defenses, and that leading models have become considerably harder to compromise.

However, he said it remains unclear whether those measures will be sufficient for highly sensitive applications.

Cui also said she has continued finding unusual ways to bypass safeguards in newer systems, including getting GPT-5.4 to return prohibited self-harm information.

She previously worked as a red-teamer for major AI labs, including Anthropic, and said she had found jailbreaks involving role-play and unexpected contextual information.

Charles Ye, another author of the study, warned that the financial incentive to develop jailbreaks and prompt-injection attacks will grow as AI agents gain access to valuable accounts, data and systems.

He said organizations may ultimately need to assume that an LLM or AI agent can be compromised rather than treating its safeguards as completely reliable.

The researchers do not argue that current defenses are useless. Instead, they conclude that training alone is unlikely to guarantee complete protection as long as LLMs continue to infer the authority of text partly from how that text is written.

Their paper argues that without a more reliable way for models to understand who or what is actually giving an instruction, prompt-injection defense could remain an ongoing cycle of finding and patching new attacks.

Get the latest tech news, telecom insights, and product launches wherever you prefer.

Add ProPakistani to Preferred Sources and see more of our stories in Google Search and Top Stories.

Technology and Automotive Specialist covering the latest cars, smartphones, AI breakthroughs, and...

Shares