Respecting Human Rights When Using LLMs in Content Moderation: Key Principles for the Tech Industry
October 1, 2026
The Oversight Board is conducting new research to provide guidance to the technology industry on how to use large language models (LLMs) for content moderation on social media platforms in a way that respects free expression, reduces harm and adheres to recognized human rights standards.
Why This Matters
Meta and other technology companies are actively moving toward deploying LLMs for content moderation to enforce content guidelines. This emerging practice holds the potential for improving content moderation for billions of users of social media platforms globally, but as with any AI-driven deployment, there are risks to free expression and other human rights that must be addressed as the technology develops, including at the design stage.
Background
When Meta announced in March its plans to deploy advanced AI models for content moderation, the Board welcomed the potential benefits, such as: protecting against the wrongful removal of speech; providing better explanations to users for moderation decisions; detecting and enforcing violating content at a scale that machine learning classifiers and human review alone cannot reach; sparing human reviewers from some of the psychological harms of reviewing certain content; and improving moderation in low-resource languages.
The Board also noted the need to mitigate the well-known risks associated with LLMs, including: guardrails against inherent biases and hallucinations; the fact that LLMs also struggle with the nuances of sarcasm, humor and coded language; and whether the system can be trusted to protect human rights during crisis situations. The Board also called for independent oversight during the transition and the need for data transparency to make accurate assessments about performance.
The Board's New Research - Rights-Based Guidance
In this research, the Board will build on six years of evaluating how Meta’s content policies are enforced by human and automated systems, and contribute Board Members’ global expertise in developing rights-based, scalable policies in the technology sector. The analysis will draw on the Board’s learnings from early, confidential access to the deployment of Meta’s advanced LLM-based content enforcement system, and insights from other industry experts, policymakers and the public.
The resulting report will provide guidance for all companies that are seeking to realize the benefits of LLM-based content moderation while mitigating its risks. The Board would appreciate public comments that address:
- How LLM-based content moderation at scale compares to existing classifier and human-reviewer-based systems.
- Views on designing scaled LLM-based content enforcement systems, including approaches to content routing and escalation pathways for human review.
- Best practices for enabling LLM-based contextual analysis while preventing bias and hallucinated outputs.
- Research into LLM-based moderation in low-resource languages, and on multimodal content (image, video, audio).
- Recommended approaches to providing greater transparency to users on LLM-based content enforcement, including appeal and review processes.
- Recommended ways to test and audit the LLM-based moderation systems.