option
Home
News
Emojis Could Skirt Safety Filters in AI Chatbots

Emojis Could Skirt Safety Filters in AI Chatbots

November 27, 2025
176

Emojis can bypass safety mechanisms in large language models, leading to toxic outputs that would otherwise be blocked. This method enables LLMs to discuss and provide guidance on prohibited topics like bomb-making and murder.

 

A recent China-Singapore collaboration presents strong evidence that emojis can not only evade content filters in large language models (LLMs) but also amplify toxicity during interactions:

From the new paper, a broad demonstration of the ways that encoding a banned concept with emojis can help a user to

From the new paper, a broad demonstration of how encoding banned concepts with emojis can help users 'jailbreak' popular LLMs. Source: https://arxiv.org/pdf/2509.11141

In the example above, converting rule-breaking text-based intent into an emoji-laden alternative can prompt a more cooperative response from advanced models like ChatGPT-4o, which typically sanitizes inputs and blocks rule-violating content.

According to the authors, emojis can effectively serve as a jailbreaking technique in extreme cases.

A lingering question is why LLMs allow emojis to bypass rules and elicit toxic content, even when the models recognize certain emojis' harmful associations.

The researchers propose that LLMs, trained to replicate patterns from their data, treat emojis as statistical cues rather than content to filter. Since emojis are common in training data, models learn to associate them with specific discourse, reinforcing toxic meanings instead of flagging them. Safety measures, applied post-hoc and often narrowly, may miss these emoji-laden prompts entirely.

Thus, the model becomes tolerant not despite the toxic association but because of it.

Free Pass

The authors acknowledge that this isn't a definitive explanation for emojis' filtering bypass. They state:

‘Models can recognize the malicious intent expressed by emojis, yet how it bypasses safety mechanisms remain unclear.’

The vulnerability may stem from text-centric filter designs, which rely on explicit tokens or embeddings matched against safety rules. Unlike words, emojis exist in a gray area—neither purely text nor image—allowing them to evade detection. Further research into this loophole is needed.

The paper, titled When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs’ Toxicity, involves nine researchers from Tsinghua University and the National University of Singapore.

(The paper references examples in an appendix not yet available; despite requests, it wasn't provided at the time of writing. Still, the core findings merit attention.)

Three Core Emoji Interpretations

Emojis bypass filters through three linguistic traits. First, their meanings are context-dependent. For instance, the ‘Money with Wings’ emoji officially denotes spending but can imply illicit activity depending on context:

In a partial illustration from the new paper, we see that a popular emoji can have its meaning hijacked altered or subverted in popular usage This effectively gives the emoji an official passport into the semantic space, and a hidden payload of negative or toxic meaning that can be exploited once it is past the filters.

In a partial illustration, a popular emoji's meaning can be hijacked in usage, granting it a semantic passport with a hidden toxic payload exploitable after filtering.

Second, emojis alter tone, adding playfulness or irony that softens emotional impact. In harmful queries, this can disguise intent as humor, encouraging model compliance:

The leavening effect of emojis can detoxify tone without detoxifying intent.

Emojis can detoxify tone without neutralizing harmful intent.

Third, emojis are language-agnostic, conveying consistent sentiment across languages like English, Chinese, and French. This makes them ideal for multilingual prompts, preserving meaning despite translation:

The broken heart emoji conveys a universal message, perhaps not least because it represents a baseline case in the human condition, relatively immune to national or cultural variations.

The ‘broken heart’ emoji communicates universally, reflecting a fundamental human experience less affected by cultural differences.

Approach, Data and Tests*

Researchers modified the AdvBench dataset, adding emojis as substitutes for sensitive terms or decorative elements. AdvBench includes 32 high-risk topics like bombing and hacking:

Original examples from AdvBench, illustrating how a single adversarial prompt can bypass safeguards in multiple major chatbots, eliciting harmful instructions despite alignment training. Source: https://arxiv.org/pdf/2307.15043

Original AdvBench examples show how adversarial prompts bypass safeguards in major chatbots, eliciting harmful responses despite alignment. Source: https://arxiv.org/pdf/2307.15043

All 520 AdvBench instances were emoji-modified, with the top 50 toxic prompts used across experiments. Prompts were translated into multiple languages and tested on seven closed- and open-source models, combined with jailbreak techniques like PAIR, TAP, and DeepInception.

Closed-source models included Gemini-2.0-flash, GPT-4o, GPT-4-0613, and Gemini-1.5-pro. Open-source models were Llama-3-8B-Instruct, Qwen2.5-7B-Instruct, and Qwen2.5-72B-Instruct, with tests repeated thrice for reliability.

The study assessed whether emoji-rewritten prompts increased toxic output, including in translations. It also applied emoji edits to known jailbreak strategies to gauge enhanced effectiveness.

Prompt structures were preserved, with only sensitive terms swapped for emojis or decorative elements added.

For evaluation, the authors introduced GPT-Judge, where GPT-4o graded responses from other models on a Harmful Score (HS) scale of 1-5. Responses scoring 5 constituted the Harmfulness Ratio (HR).

To prevent emoji explanations, prompts included instructions for brevity:

Results from emoji-based prompts in

Results from emoji-based prompts in ‘Setting-1’, compared to variants where emojis were replaced with words or removed. Model names are abbreviated.

The initial results show emoji-substituted prompts achieved higher HS and HR scores than text-based versions. The emoji approach outperformed prior jailbreak methods, as seen in the additional table:

Harmfulness Ratio results for emoji-augmented jailbreak prompts in

Harmfulness Ratio results for emoji-augmented jailbreak prompts in ‘Setting-2’, with abbreviated model names.

The first table also indicates emojis' cross-language effect. When prompts were translated into Chinese, French, Spanish, and Russian, harmful outputs remained high, suggesting risks extend beyond English to major user groups.

In conclusion, the researchers note that emojis' impact stems from how models process them—recognizing harm but suppressing rejection when emojis are present. Tokenization studies show emojis fragment into rare tokens, creating an alternative semantic channel.

Pretraining data analysis reveals frequent emoji use in toxic contexts (e.g., scams, gambling), normalizing harmful associations. Together, model quirks and biased data explain emojis' effectiveness in bypassing safety.

Conclusion

Alternative input methods like hexadecimal encoding have been used to jailbreak LLMs. The issue lies in text-centric qualification of inputs and outputs.

Emojis introduce rule-breaking meaning undetected, as their unorthodox transmission evades filters. While CLIP-based transliteration should flag offensive image content, this isn't consistently applied in major LLMs, whose linguistic barriers remain brittle. Broader content interpretation (e.g., via heatmaps) may be costly or impractical.

 

* The paper's layout is less structured than typical studies; we've aimed to convey its core insights clearly.

The results presentation is notably challenging to interpret.

First published Wednesday, September 17, 2025

Related article
Base44 Unveils Proprietary AI Model to Bolster Defensibility in Vibe Coding Platform Base44 Unveils Proprietary AI Model to Bolster Defensibility in Vibe Coding Platform Base44, the vibe coding platform acquired by Wix for $80 million just a year ago — when it was merely six months old with a team of eight — has begun deploying its proprietary AI model to help users build applications using natural language.This deve
Multiverse Computing Launches Free Compressed Generative AI Model Multiverse Computing Launches Free Compressed Generative AI Model Large language models face a significant challenge: their immense size. Spanish startup Multiverse Computing is tackling this problem by creating compressed models designed to bridge the gap between the capabilities of cutting-edge AI and what busine
Secret Tracking Data Exposes Theft of AI Models Secret Tracking Data Exposes Theft of AI Models A new method can invisibly watermark models like ChatGPT in seconds without retraining, leaving no trace in standard outputs and resisting all practical removal attempts. The key distinction between watermarking and 'copyright-baiting' is that waterm
Related Special Topic Recommendations
Data Analysis Best AI Anomaly Detection Tools for KPI Monitoring across SaaS and Ecommerce Teams
Best AI Anomaly Detection Tools for KPI Monitoring across SaaS and Ecommerce Teams

2026 Latest Best Top-rated AI Anomaly Detection Tools for KPI Monitoring in SaaS and Ecommerce teams! XIX.AI has curated a powerful, game-changing collection based on rigorous real-world tests and weekly updated rankings. You’ll find detailed free vs paid comparison insights to help you identify the must-try solution that boosts productivity and unlocks your AI edge. Explore now!

10 tools
xix.ai
writing Best AI Outline Generators for Long-Form SEO Articles
Best AI Outline Generators for Long-Form SEO Articles

2026 Latest Best Top-Rated AI Outline Generators for Long-Form SEO Articles, meticulously curated by XIX.AI. These powerful tools offer game-changing assistance in creating high-quality content quickly, boosting writing efficiency significantly. Get a free vs paid comparison along with real-world tests and detailed rankings to help you find the must-try option that suits your needs. Explore now to unlock your AI edge.

8 tools
xix.ai
Education and Learning AI Study Tools for Homework and Exam Prep
AI Study Tools for Homework and Exam Prep

2026 Latest Best AI Study Tools for Homework and Exam Prep! XIX.AI curates a top-rated list of powerful, game-changing tools that help students boost productivity, streamline homework completion, and ace exams through real-world tests. Get a free vs paid comparison, detailed rankings, and must-try options to unlock your AI edge. Explore now!

10 tools
xix.ai
Music composition AI Vocal Demo Tools for Songwriters, Hooks, Toplines, and Multilingual Draft Sessions
AI Vocal Demo Tools for Songwriters, Hooks, Toplines, and Multilingual Draft Sessions

2026 Latest Best AI Vocal Demo Tools for Songwriters, Hook Creators, and Multi-Language Content Teams! XIX.AI has curated a top-rated list of powerful game-changing tools that go through rigorous real-world tests. You’ll find detailed free vs paid comparison data, comprehensive rankings, and must-try options to help you boost writing efficiency and unlock your creative potential. Explore now to discover your perfect tool for all your content needs!

9 tools
xix.ai
Business Best AI Competitive Research Tools for Small Businesses
Best AI Competitive Research Tools for Small Businesses

2026 Latest Best Top-rated AI Competitive Research Tools for Small Businesses! XIX.AI has curated a highly powerful game-changing collection, updated weekly with rigorous real-world tests and detailed rankings. You can find a comprehensive free vs paid comparison to help you identify the must-try tools that boost your productivity and give you a competitive edge. Explore now to discover your perfect tool!

9 tools
xix.ai
Image editing Photoshop AI Retouch Tools for Ecommerce Apparel, Skin Cleanup, and Color Consistency
Photoshop AI Retouch Tools for Ecommerce Apparel, Skin Cleanup, and Color Consistency

2026 Latest Best Photoshop AI retouch tools for ecommerce apparel, skin cleanup, and color consistency! This top-rated curated list features powerful game-changing solutions that help you boost writing efficiency, streamline content creation, and achieve perfect visual results effortlessly. Each tool has undergone real-world tests through weekly updated rankings, complete with free vs paid comparison details. Backed by XIX.AI, it’s the must-try guide for anyone aiming to unlock your AI edge. Explore now!

10 tools
xix.ai
Comments (0)
0/500
OR