option
Home
News
Vibe-Based Image Labeling Poses Significant Risks

Vibe-Based Image Labeling Poses Significant Risks

February 21, 2026
130

Though often paid very little or nothing at all, the anonymous individuals who assess images for "harmful" content hold significant power over your life through their decisions. A major new study from Google now suggests that these annotators often develop their own personal rules for what constitutes "harmful" or offensive material—regardless of how unusual or subjective their reactions to an image may be. So what could possibly go wrong?

 

Opinion This week, a joint effort by Google Research and Google Mind brought together 13 contributors for a new paper examining whether the "gut feelings" of image annotators should influence how images are rated for algorithms—even when those feelings conflict with established rating guidelines.

This issue matters to you because the standards annotators agree upon as offensive tend to become embedded in automated moderation systems, legal definitions of "obscene" or "unacceptable" content—such as the UK’s upcoming NSFW firewall* (with Australia soon to follow)—and content evaluation mechanisms on social media and other platforms.

In short, the broader the criteria for what's considered offensive, the more extensive the potential censorship.

Vibe-Censorship

The study doesn’t stop there. It also reveals that image raters often censor more strictly based on what they believe might offend others, not just themselves. Additionally, low-quality images frequently trigger safety concerns, despite image quality being unrelated to content.

In its conclusion, the paper highlights these two findings almost as if its central argument fell short—yet the researchers felt compelled to publish anyway.

While that’s not unusual in academic publishing, a closer reading reveals a more unsettling undercurrent: annotation practices may be moving toward what can only be called vibe-annotating:

“Our findings suggest that existing frameworks need to account for subjective and contextual dimensions, such as emotional reactions, implicit judgments, and cultural interpretations of harm. Annotators’ frequent use of emotional language and their divergence from predefined harm labels highlight gaps in current evaluation practices.

Expanding annotation guidelines to include illustrative examples of diverse cultural and emotional interpretations can help address these gaps.”

The scantly-illustrated new paper leads with examples that are unambiguous and sympathetic to the average reader, although the actual core material is far more ambiguous and invites many more questions. Here, under each image, we see annotators

The new paper uses straightforward examples that resonate with most readers, though its core content raises far more complex questions. Here, each image is accompanied by annotators' emotional responses. Source: https://arxiv.org/pdf/2507.16033

Initially, this sounds like a reasonable effort to better define "harm" in images—a worthy goal. But the paper repeatedly suggests that achieving this may not be practical or even desirable:

“Our findings suggest that existing frameworks need to account for subjective and contextual dimensions, such as emotional reactions, implicit judgments, and cultural interpretations of harm. Annotators’ frequent use of emotional language and their divergence from predefined harm labels highlight gaps in current evaluation practices.

Expanding annotation guidelines to include illustrative examples of diverse cultural and emotional interpretations can help address these gaps […]

[…] The process by which annotators reason about ambiguous images often reflects their personal, cultural, and emotional perspectives, which are difficult to scaffold or standardize.”

It's hard to see how including "illustrative examples of diverse cultural and emotional interpretations" fits into a rational rating system. The authors repeatedly grapple with this point without offering a clear solution, making their central argument feel like it, too, is driven by "vibes"—even as it deals with intangible psychological factors.

Simply put, expanding annotation criteria in this way could allow any content—or entire topics—that an annotator reacts strongly to, to be suppressed or obscured.

Binary Judgement

Quantifying the harm caused by images and text is inherently difficult, especially since "high" and "low" culture often overlap—as seen in art and literature. This has led to early forms of "vibe-based" censorship, where obscene material is judged not by strict definition, but by the "you know it when you see it" principle.

Beneath its extensive discussion of empathy and nuance, the paper subtly challenges the authority of centralized, standardized categories like "violence," "nudity," and "hate"—categories that enable platforms to implement scalable moderation with reasonable accuracy.

The emerging argument is that only decentralized, subjective, context-aware human judgment can properly evaluate GenAI output.

But this approach isn't scalable. You can’t filter billions of images based on "vibes" and personal experience. Harm must be quantified into specific properties; filtering systems need clear limits; and edge cases require updated guidelines—much like how new laws are sometimes needed to address unique grievances.

Instead, the paper seems to advocate for an automated moderation system that automatically broadens its scope—so cautious that even one annotator’s highly personal reaction could penalize an image that offends no one else.

Moral Expansion

Although exploratory in nature, the paper applies scientific methods: the authors created a framework to identify—though not strictly measure—a wider range of annotator reactions to images, and examined how these vary by gender and other demographics.

Beyond analyzing harm-focus, the study examined "moral reasoning" in annotators’ additional comments. Participants were asked to annotate a modified dataset containing images, prompts, and related text.

This "moral sentiment autorater" was designed to capture moral values such as Care, Equality, Proportionality, Loyalty, Authority, and Purity, based on Moral Foundations Theory—a psychological model that, due to its fluid nature, is ill-suited for creating the concrete definitions needed in large-scale rating systems.

Inspired by this theory, the authors introduced additional safety dimensions, including fear, anger, sadness, disgust, confusion, and uncanniness.

The authors elaborate on fear:

“Many annotators used terms like ‘scary’ (e.g., for distorted faces or images suggesting violence like a gun pointed at a child), ‘disturbing’ (e.g., ‘Absolutely vile to see someone get ran over, very distressing and disturbing,’ or ‘Disturbing and looks like blood’ for red paint), or ‘upsetting’ (e.g., ‘The image of the boy has many distortions… I find it distasteful because it appears that the boy is playing on the wrong side of the side rails’).

The [graph below] shows that ‘fear’ was the most frequently mentioned emotion (233 mentions). While nearly half of these were linked to violent content, the second-highest number of fear mentions came from content deemed not harmful.”

Distribution of emotion-related terms across harm categories, with bar heights indicating proportions of comments, counts displayed within bars, and total comment counts shown above each category.

Distribution of emotion-related terms across harm categories, with bar heights indicating proportions of comments, counts displayed within bars, and total comment counts shown above each category.

Regarding these new safety dimensions, the authors state:

“These emerging themes highlight a critical need to enrich AI image evaluation frameworks by integrating subjective, emotional, and perceptual elements.”

This direction could be risky, as it may allow annotation processes to arbitrarily introduce rules based on individual reactions—rather than requiring all annotators to follow consistent standards.

If there's an economic motive here, it's that this model enables hyperscale human annotation: a frictionless, self-regulating system where participants define the rules themselves.

Under standard annotation, rules are agreed upon by consensus and followed by annotators. Under the paper’s proposed model, that oversight is reduced or removed—meaning any image that offends even one person could be flagged, partly because building consensus is costly and time-consuming.

Rorschach Judgements

The goal of annotation is to produce accurate descriptions through expert oversight, consensus, or ideally both. Expanding a clear hierarchy of harms into an "intuitive," highly personal process is like annotating a Rorschach test.

For example, the paper notes that some annotators interpreted poor image quality—such as JPEG artifacts or technical flaws—as “disturbing” or “indicative of harm”:

“This occurred despite the task omitting instructions on image quality. Moreover, annotators interpreted these quality artifacts as semantically meaningful.

One annotator commented, ‘The image is not harmful at all; he just has a bit of a distorted face.’ Similarly, others saw image flaws as intentional harm, attributing emotional meaning to glitches. For example, one interpreted a distorted face as ‘indicative of pain.’”

By prioritizing subjective, emotional, or context-specific reactions over predefined safety labels, this approach risks creating a system where anything can be arbitrarily flagged as harmful—leading to a "chilling effect" of ad hoc content removal or recategorization, especially for material that might offend specific interest groups.

 

 

The paper “Just a strange pic”: Evaluating ‘safety' in GenAI Image safety annotation tasks from diverse annotators' perspectives is available on Arxiv.

* A simplified reference, as this isn’t the main focus. Under the new law, offending sites must either self-police, implement costly review and age-verification systems (feasible only for large platforms), or block UK access—again, at their own expense.

Often simplified as the “think of the children” meme, which satirizes using others’ moral agency for seemingly altruistic purposes.

 

First published Friday, July 25, 2025

Related article
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control According to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
How to fix Core Web Vitals for better SEO rankings How to fix Core Web Vitals for better SEO rankings Streamline Report Card Comments with AI ToolsIntroductionAI Tools for Generating Report Card CommentsMagic SchoolAlmanac AIChat GPTUsing Magic School to Generate Report Card CommentsLogging into Magic SchoolSelecting the Report Card Comments ToolCust
Related Special Topic Recommendations
writing Best AI Outline Generators for Long-Form SEO Articles
Best AI Outline Generators for Long-Form SEO Articles

2026 Latest Best Top-Rated AI Outline Generators for Long-Form SEO Articles, meticulously curated by XIX.AI. These powerful tools offer game-changing assistance in creating high-quality content quickly, boosting writing efficiency significantly. Get a free vs paid comparison along with real-world tests and detailed rankings to help you find the must-try option that suits your needs. Explore now to unlock your AI edge.

8 tools
xix.ai
Education and Learning AI Study Tools for Homework and Exam Prep
AI Study Tools for Homework and Exam Prep

2026 Latest Best AI Study Tools for Homework and Exam Prep! XIX.AI curates a top-rated list of powerful, game-changing tools that help students boost productivity, streamline homework completion, and ace exams through real-world tests. Get a free vs paid comparison, detailed rankings, and must-try options to unlock your AI edge. Explore now!

10 tools
xix.ai
Music composition AI Vocal Demo Tools for Songwriters, Hooks, Toplines, and Multilingual Draft Sessions
AI Vocal Demo Tools for Songwriters, Hooks, Toplines, and Multilingual Draft Sessions

2026 Latest Best AI Vocal Demo Tools for Songwriters, Hook Creators, and Multi-Language Content Teams! XIX.AI has curated a top-rated list of powerful game-changing tools that go through rigorous real-world tests. You’ll find detailed free vs paid comparison data, comprehensive rankings, and must-try options to help you boost writing efficiency and unlock your creative potential. Explore now to discover your perfect tool for all your content needs!

9 tools
xix.ai
Business Best AI Competitive Research Tools for Small Businesses
Best AI Competitive Research Tools for Small Businesses

2026 Latest Best Top-rated AI Competitive Research Tools for Small Businesses! XIX.AI has curated a highly powerful game-changing collection, updated weekly with rigorous real-world tests and detailed rankings. You can find a comprehensive free vs paid comparison to help you identify the must-try tools that boost your productivity and give you a competitive edge. Explore now to discover your perfect tool!

9 tools
xix.ai
Image editing Photoshop AI Retouch Tools for Ecommerce Apparel, Skin Cleanup, and Color Consistency
Photoshop AI Retouch Tools for Ecommerce Apparel, Skin Cleanup, and Color Consistency

2026 Latest Best Photoshop AI retouch tools for ecommerce apparel, skin cleanup, and color consistency! This top-rated curated list features powerful game-changing solutions that help you boost writing efficiency, streamline content creation, and achieve perfect visual results effortlessly. Each tool has undergone real-world tests through weekly updated rankings, complete with free vs paid comparison details. Backed by XIX.AI, it’s the must-try guide for anyone aiming to unlock your AI edge. Explore now!

10 tools
xix.ai
Prompt Best AI Prompt Libraries for ChatGPT Workflows
Best AI Prompt Libraries for ChatGPT Workflows

2026 Latest Best Top-Rated AI Prompt Libraries for optimizing all types of ChatGPT workflows. XIX.AI has curated a powerful, game-changing collection that goes through rigorous real-world tests to ensure top performance. You can find detailed free vs paid comparisons and expert rankings to help you choose the must-try tools that boost your productivity and unlock your AI edge. Explore now!

11 tools
xix.ai
Comments (1)
0/500
DonaldGonzález
DonaldGonzález August 29, 2026 at 8:00:15 AM EDT

画像ラベリングの匿名作業員の待遇と権限の問題、Googleの研究で浮き彫りに。低賃金で高リスクの責任を背負う人々の苦境に注目。AIの透明性確保には人間の労働環境改善も不可欠。

OR