Emojis can bypass safety mechanisms in large language models, leading to toxic outputs that would otherwise be blocked. This method enables LLMs to discuss and provide guidance on prohibited topics like bomb-making and murder.
A recent China-Singapore collaboration presents strong evidence that emojis can not only evade content filters in large language models (LLMs) but also amplify toxicity during interactions:
From the new paper, a broad demonstration of how encoding banned concepts with emojis can help users 'jailbreak' popular LLMs. Source: https://arxiv.org/pdf/2509.11141
In the example above, converting rule-breaking text-based intent into an emoji-laden alternative can prompt a more cooperative response from advanced models like ChatGPT-4o, which typically sanitizes inputs and blocks rule-violating content.
According to the authors, emojis can effectively serve as a jailbreaking technique in extreme cases.
A lingering question is why LLMs allow emojis to bypass rules and elicit toxic content, even when the models recognize certain emojis' harmful associations.
The researchers propose that LLMs, trained to replicate patterns from their data, treat emojis as statistical cues rather than content to filter. Since emojis are common in training data, models learn to associate them with specific discourse, reinforcing toxic meanings instead of flagging them. Safety measures, applied post-hoc and often narrowly, may miss these emoji-laden prompts entirely.
Thus, the model becomes tolerant not despite the toxic association but because of it.
Free Pass
The authors acknowledge that this isn't a definitive explanation for emojis' filtering bypass. They state:
‘Models can recognize the malicious intent expressed by emojis, yet how it bypasses safety mechanisms remain unclear.’
The vulnerability may stem from text-centric filter designs, which rely on explicit tokens or embeddings matched against safety rules. Unlike words, emojis exist in a gray area—neither purely text nor image—allowing them to evade detection. Further research into this loophole is needed.
The paper, titled When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs’ Toxicity, involves nine researchers from Tsinghua University and the National University of Singapore.
(The paper references examples in an appendix not yet available; despite requests, it wasn't provided at the time of writing. Still, the core findings merit attention.)
Three Core Emoji Interpretations
Emojis bypass filters through three linguistic traits. First, their meanings are context-dependent. For instance, the ‘Money with Wings’ emoji officially denotes spending but can imply illicit activity depending on context:
In a partial illustration, a popular emoji's meaning can be hijacked in usage, granting it a semantic passport with a hidden toxic payload exploitable after filtering.
Second, emojis alter tone, adding playfulness or irony that softens emotional impact. In harmful queries, this can disguise intent as humor, encouraging model compliance:
Emojis can detoxify tone without neutralizing harmful intent.
Third, emojis are language-agnostic, conveying consistent sentiment across languages like English, Chinese, and French. This makes them ideal for multilingual prompts, preserving meaning despite translation:
The ‘broken heart’ emoji communicates universally, reflecting a fundamental human experience less affected by cultural differences.
Approach, Data and Tests*
Researchers modified the AdvBench dataset, adding emojis as substitutes for sensitive terms or decorative elements. AdvBench includes 32 high-risk topics like bombing and hacking:
Original AdvBench examples show how adversarial prompts bypass safeguards in major chatbots, eliciting harmful responses despite alignment. Source: https://arxiv.org/pdf/2307.15043
All 520 AdvBench instances were emoji-modified, with the top 50 toxic prompts used across experiments. Prompts were translated into multiple languages and tested on seven closed- and open-source models, combined with jailbreak techniques like PAIR, TAP, and DeepInception.
Closed-source models included Gemini-2.0-flash, GPT-4o, GPT-4-0613, and Gemini-1.5-pro. Open-source models were Llama-3-8B-Instruct, Qwen2.5-7B-Instruct, and Qwen2.5-72B-Instruct, with tests repeated thrice for reliability.
The study assessed whether emoji-rewritten prompts increased toxic output, including in translations. It also applied emoji edits to known jailbreak strategies to gauge enhanced effectiveness.
Prompt structures were preserved, with only sensitive terms swapped for emojis or decorative elements added.
For evaluation, the authors introduced GPT-Judge, where GPT-4o graded responses from other models on a Harmful Score (HS) scale of 1-5. Responses scoring 5 constituted the Harmfulness Ratio (HR).
To prevent emoji explanations, prompts included instructions for brevity:
Results from emoji-based prompts in ‘Setting-1’, compared to variants where emojis were replaced with words or removed. Model names are abbreviated.
The initial results show emoji-substituted prompts achieved higher HS and HR scores than text-based versions. The emoji approach outperformed prior jailbreak methods, as seen in the additional table:
Harmfulness Ratio results for emoji-augmented jailbreak prompts in ‘Setting-2’, with abbreviated model names.
The first table also indicates emojis' cross-language effect. When prompts were translated into Chinese, French, Spanish, and Russian, harmful outputs remained high, suggesting risks extend beyond English to major user groups.
In conclusion, the researchers note that emojis' impact stems from how models process them—recognizing harm but suppressing rejection when emojis are present. Tokenization studies show emojis fragment into rare tokens, creating an alternative semantic channel.
Pretraining data analysis reveals frequent emoji use in toxic contexts (e.g., scams, gambling), normalizing harmful associations. Together, model quirks and biased data explain emojis' effectiveness in bypassing safety.
Conclusion
Alternative input methods like hexadecimal encoding have been used to jailbreak LLMs. The issue lies in text-centric qualification of inputs and outputs.
Emojis introduce rule-breaking meaning undetected, as their unorthodox transmission evades filters. While CLIP-based transliteration should flag offensive image content, this isn't consistently applied in major LLMs, whose linguistic barriers remain brittle. Broader content interpretation (e.g., via heatmaps) may be costly or impractical.
* The paper's layout is less structured than typical studies; we've aimed to convey its core insights clearly.
†The results presentation is notably challenging to interpret.
Multiverse Computing Launches Free Compressed Generative AI ModelLarge language models face a significant challenge: their immense size. Spanish startup Multiverse Computing is tackling this problem by creating compressed models designed to bridge the gap between the capabilities of cutting-edge AI and what busine
Secret Tracking Data Exposes Theft of AI ModelsA new method can invisibly watermark models like ChatGPT in seconds without retraining, leaving no trace in standard outputs and resisting all practical removal attempts. The key distinction between watermarking and 'copyright-baiting' is that waterm
2026 Latest Best Top-rated AI Anomaly Detection Tools for KPI Monitoring in SaaS and Ecommerce teams! XIX.AI has curated a powerful, game-changing collection based on rigorous real-world tests and weekly updated rankings. You’ll find detailed free vs paid comparison insights to help you identify the must-try solution that boosts productivity and unlocks your AI edge. Explore now!
2026 Latest Best Top-Rated AI Outline Generators for Long-Form SEO Articles, meticulously curated by XIX.AI. These powerful tools offer game-changing assistance in creating high-quality content quickly, boosting writing efficiency significantly. Get a free vs paid comparison along with real-world tests and detailed rankings to help you find the must-try option that suits your needs. Explore now to unlock your AI edge.
2026 Latest Best AI Study Tools for Homework and Exam Prep! XIX.AI curates a top-rated list of powerful, game-changing tools that help students boost productivity, streamline homework completion, and ace exams through real-world tests. Get a free vs paid comparison, detailed rankings, and must-try options to unlock your AI edge. Explore now!
2026 Latest Best AI Vocal Demo Tools for Songwriters, Hook Creators, and Multi-Language Content Teams! XIX.AI has curated a top-rated list of powerful game-changing tools that go through rigorous real-world tests. You’ll find detailed free vs paid comparison data, comprehensive rankings, and must-try options to help you boost writing efficiency and unlock your creative potential. Explore now to discover your perfect tool for all your content needs!
2026 Latest Best Top-rated AI Competitive Research Tools for Small Businesses! XIX.AI has curated a highly powerful game-changing collection, updated weekly with rigorous real-world tests and detailed rankings. You can find a comprehensive free vs paid comparison to help you identify the must-try tools that boost your productivity and give you a competitive edge. Explore now to discover your perfect tool!
2026 Latest Best Photoshop AI retouch tools for ecommerce apparel, skin cleanup, and color consistency! This top-rated curated list features powerful game-changing solutions that help you boost writing efficiency, streamline content creation, and achieve perfect visual results effortlessly. Each tool has undergone real-world tests through weekly updated rankings, complete with free vs paid comparison details. Backed by XIX.AI, it’s the must-try guide for anyone aiming to unlock your AI edge. Explore now!
By clicking "Accept All Cookies", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts.Privacy Policy Notice
When you visit any website, it may store or retrieve information on your browser, mostly in the form of cookies. This information might be about you, your preferences or your device and is mostly used to make the site work as you expect it to. The information does not usually directly identify you, but it can give you a more personalized web experience. Because we respect your right to privacy, you can choose not to allow some types of cookies. Click on the different category headings to find out more and change our default settings.However, blocking some types of cookies may impact your experience of the site and the services we are able to offer. Privacy PolicyStatement
Manage Preferences
Strictly Necessary Cookie
Always Active
These cookies are necessary for the website to function and cannot be switched off in our systems. They are usually only set in response to actions made by you which amount to a request for services, such as setting your privacy preferences, logging in or filling in forms. You can set your browser to block or alert you about these cookies, but some parts of the site will not then work. These cookies do not store any personally identifiable information.