A fundamental flaw in large language models (LLMs) leaves them strikingly susceptible to manipulation. Researchers have discovered that by mimicking the style and content of an LLM's internal chain-of-thought, they can trick these models into revealing harmful information.
The issue stems from how LLMs interpret roles within their text. User instructions, generated thoughts, and external prompts are all jumbled together in a continuous stream of tokens, making it difficult for the model to distinguish where its own ideas end and external commands begin. This leads to potential security breaches, as demonstrated by experiments that made popular models disclose illicit information they had been trained not to share.
The revelation has significant implications for the safety of AI in various fields, including government, military, health care, and online services. Hackers can exploit this flaw to manipulate LLMs into revealing sensitive or harmful information, undermining the trust placed in these systems.
Despite efforts by companies like OpenAI to combat such vulnerabilities through red-teaming and super-hackers, the inherent limitations mean that exhaustive lists of prohibited actions are insufficient. The researchers argue that this fundamental flaw is fundamentally unsolvable, highlighting a critical gap in current AI security practices.







