option
Home
News
New Technique Enables DeepSeek and Other Models to Respond to Sensitive Queries

New Technique Enables DeepSeek and Other Models to Respond to Sensitive Queries

May 11, 2025
178

New Technique Enables DeepSeek and Other Models to Respond to Sensitive Queries

Removing bias and censorship from large language models (LLMs) like China's DeepSeek is a complex challenge that has caught the attention of U.S. policymakers and business leaders, who see it as a potential national security threat. A recent report from a U.S. Congress select committee labeled DeepSeek as "a profound threat to our nation's security" and offered policy recommendations to address the issue.

While techniques like Reinforcement Learning from Human Feedback (RLHF) and fine-tuning can help mitigate bias, the enterprise risk management startup CTGT claims to have developed a novel approach. According to CTGT, their method can completely eliminate censorship in LLMs. Cyril Gorlla and Trevor Tuttle of CTGT detailed their framework in a paper, explaining that it "directly locates and modifies the internal features responsible for censorship."

Their approach is not only efficient but also allows precise control over the model's behavior, ensuring that uncensored responses are provided without affecting the model's overall capabilities or factual accuracy. Although initially designed for DeepSeek-R1-Distill-Llama-70B, the method can be applied to other models as well. Gorlla confirmed to VentureBeat that CTGT's technology works at the foundational neural network level, making it applicable to all deep learning models. They are collaborating with a leading foundation model lab to ensure new models are inherently trustworthy and safe.

How It Works

The researchers at CTGT identify features within the model that are likely associated with unwanted behaviors. They explained that "within a large language model, there exist latent variables (neurons or directions in the hidden state) that correspond to concepts like 'censorship trigger' or 'toxic sentiment'. If we can find those variables, we can directly manipulate them."

CTGT's method involves three key steps:

  1. Feature identification
  2. Feature isolation and characterization
  3. Dynamic feature modification

To identify these features, researchers use prompts designed to trigger "toxic sentiments," such as inquiries about Tiananmen Square or tips for bypassing firewalls. They analyze the responses to establish patterns and locate the vectors where the model decides to censor information. Once identified, they isolate the feature and understand which part of the unwanted behavior it controls, whether it's responding cautiously or refusing to answer. They then integrate a mechanism into the model's inference pipeline to adjust the activation level of the feature's behavior.

Making the Model Answer More Prompts

CTGT's experiments, using 100 sensitive queries, showed that the base DeepSeek-R1-Distill-Llama-70B model answered only 32% of the controversial prompts. However, the modified version responded to 96% of the prompts, with the remaining 4% being extremely explicit content. The company emphasized that their method allows users to adjust the model's bias and safety features without turning it into a "reckless generator," especially when only unnecessary censorship is removed.

Importantly, this method does not compromise the model's accuracy or performance. Unlike traditional fine-tuning, it doesn't involve optimizing model weights or providing new example responses. This offers two major advantages: immediate effect on the next token generation and the ability to switch between different behaviors by toggling the feature adjustment on or off, or even adjusting it to varying degrees for different contexts.

Model Safety and Security

The congressional report on DeepSeek urged the U.S. to "take swift action to expand export controls, improve export control enforcement, and address risks from Chinese artificial intelligence models." As concerns about DeepSeek's potential national security threat grew, researchers and AI companies began exploring ways to make such models safer.

Determining what is "safe," biased, or censored can be challenging, but methods that allow users to adjust model controls to suit their needs could be highly beneficial. Gorlla emphasized that enterprises "need to be able to trust their models are aligned with their policies," highlighting the importance of methods like CTGT's for businesses.

"CTGT enables companies to deploy AI that adapts to their use cases without having to spend millions of dollars fine-tuning models for each use case. This is particularly important in high-risk applications like security, finance, and healthcare, where the potential harms that can come from AI malfunctioning are severe," Gorlla stated.

Related article
Base44 Unveils Proprietary AI Model to Bolster Defensibility in Vibe Coding Platform Base44 Unveils Proprietary AI Model to Bolster Defensibility in Vibe Coding Platform Base44, the vibe coding platform acquired by Wix for $80 million just a year ago — when it was merely six months old with a team of eight — has begun deploying its proprietary AI model to help users build applications using natural language.This deve
DeepSeek Unveils AI Model Rivaling Frontier Systems DeepSeek Unveils AI Model Rivaling Frontier Systems Chinese AI lab DeepSeek has released two preview versions of its latest large language model, DeepSeek V4, a highly anticipated update to last year's V3.2 model and the accompanying R1 reasoning model that made a significant impact in the AI communit
Multiverse Computing Launches Free Compressed Generative AI Model Multiverse Computing Launches Free Compressed Generative AI Model Large language models face a significant challenge: their immense size. Spanish startup Multiverse Computing is tackling this problem by creating compressed models designed to bridge the gap between the capabilities of cutting-edge AI and what busine
Related Special Topic Recommendations
Music composition AI Vocal Demo Tools for Songwriters, Hooks, Toplines, and Multilingual Draft Sessions
AI Vocal Demo Tools for Songwriters, Hooks, Toplines, and Multilingual Draft Sessions

2026 Latest Best AI Vocal Demo Tools for Songwriters, Hook Creators, and Multi-Language Content Teams! XIX.AI has curated a top-rated list of powerful game-changing tools that go through rigorous real-world tests. You’ll find detailed free vs paid comparison data, comprehensive rankings, and must-try options to help you boost writing efficiency and unlock your creative potential. Explore now to discover your perfect tool for all your content needs!

9 tools
xix.ai
Business Best AI Competitive Research Tools for Small Businesses
Best AI Competitive Research Tools for Small Businesses

2026 Latest Best Top-rated AI Competitive Research Tools for Small Businesses! XIX.AI has curated a highly powerful game-changing collection, updated weekly with rigorous real-world tests and detailed rankings. You can find a comprehensive free vs paid comparison to help you identify the must-try tools that boost your productivity and give you a competitive edge. Explore now to discover your perfect tool!

9 tools
xix.ai
Image editing Photoshop AI Retouch Tools for Ecommerce Apparel, Skin Cleanup, and Color Consistency
Photoshop AI Retouch Tools for Ecommerce Apparel, Skin Cleanup, and Color Consistency

2026 Latest Best Photoshop AI retouch tools for ecommerce apparel, skin cleanup, and color consistency! This top-rated curated list features powerful game-changing solutions that help you boost writing efficiency, streamline content creation, and achieve perfect visual results effortlessly. Each tool has undergone real-world tests through weekly updated rankings, complete with free vs paid comparison details. Backed by XIX.AI, it’s the must-try guide for anyone aiming to unlock your AI edge. Explore now!

10 tools
xix.ai
Prompt Best AI Prompt Libraries for ChatGPT Workflows
Best AI Prompt Libraries for ChatGPT Workflows

2026 Latest Best Top-Rated AI Prompt Libraries for optimizing all types of ChatGPT workflows. XIX.AI has curated a powerful, game-changing collection that goes through rigorous real-world tests to ensure top performance. You can find detailed free vs paid comparisons and expert rankings to help you choose the must-try tools that boost your productivity and unlock your AI edge. Explore now!

11 tools
xix.ai
Education and Learning AI Quiz Builder Platforms for Teachers, Tutors, and Cohort-Based Learning Programs
AI Quiz Builder Platforms for Teachers, Tutors, and Cohort-Based Learning Programs

2026 Latest Best AI Quiz Builder Platforms for Teachers, Tutors, and Cohort-Based Learning Programs! XIX.AI has curated a top-rated list of powerful game-changing tools that go through real-world tests to deliver accurate rankings. These must-try platforms help boost writing efficiency, streamline content creation, and simplify quiz design across all learning scenarios. Explore now to discover your perfect tool for unlocking your AI edge in teaching!

13 tools
xix.ai
code AI Pull Request Review Tools for GitHub Teams Handling Refactors, Bugs, and Security Gaps
AI Pull Request Review Tools for GitHub Teams Handling Refactors, Bugs, and Security Gaps

2026 Latest Best AI Pull Request Review Tools for GitHub Teams are here on XIX.AI! This top-rated curated list showcases powerful game-changing solutions that streamline refactoring, bug fixing, and security gap detection across all team workflows. Enjoy a free vs paid comparison along with real-world tests and detailed rankings to help you find the perfect tool that boosts productivity significantly. Explore now to unlock your AI edge!

12 tools
xix.ai
Comments (4)
0/500
CarlGarcia
CarlGarcia March 22, 2026 at 8:01:13 PM EDT

É impressionante a rapidez com que questões de 'segurança nacional' aparecem quando se fala de inovações vindas de outros países. Este relatório sobre o DeepSeek soa mais como justificativa para manter uma vantagem tecnológica do que uma genuína preocupação ética. Já parou para pensar se a 'neutralidade' que buscam não é apenas uma forma de censura disfarçada? 🤔 A corrida pela IA está mesmo acirrada.

GaryGonzalez
GaryGonzalez December 25, 2025 at 9:30:40 AM EST

この記事を読んで、AIのバイアス除去って本当に可能なのかな?技術的には興味深いけど、各国の規制や価値観の違いを考えると、完全に中立なAIを作るのは無理なんじゃないかって思う。DeepSeekが米国で国家安全保障上の脅威と見なされているって…地政学的な要素が技術開発にこんなに影響するなんて。🤔

CharlesThomas
CharlesThomas December 4, 2025 at 3:30:40 PM EST

この手法、完全にセンシティブなクエリに対して何でも返信し始めたら怖くない? 倫理的なライン越えてる気がするけど、政治的な発言の規制が緩和されるのは歓迎かも🤔 でもAIが中立を装いながら偏った情報を流す可能性も…

JustinAnderson
JustinAnderson August 21, 2025 at 1:01:17 AM EDT

¡Vaya! Quitar sesgos a modelos como DeepSeek suena a un puzzle imposible. ¿Realmente pueden hacer que una IA sea neutral? Me preocupa que esto termine siendo una carrera por controlar la narrativa. 😬

OR