A cross-border community for researchers with openness, equality and inclusion
Context-Aware Real-Time Audio and Image Toxicity Moderation via Multi-Agent Reinforcement Learning for Cybersecurity in Networked Communication Platforms
ID:67 View protection:Participant Only Updated time:2026-07-25 18:02:18 Views:29 Online

Start Time:2026-07-30 15:25

Duration:15min

Session:[S3] Cyber Security [S3-2] Cyber Security

Abstract

Toxic content propagation across networked communication

platforms constitutes an emerging cybersecurity

challenge, as real-time audio and image channels introduce

attack surfaces that bypass conventional text-only defences. Automated

content moderation in online communication platforms

increasingly requires coverage beyond text, as toxic content is

frequently delivered through audio and image channels that textonly

systems cannot address. This paper presents a dual-modality

toxic content detection and moderation pipeline operating across

audio and image inputs in real time. For audio, the Roblox

voice-safety-classifier-v2, a WavLM transformer pretrained on

over 100,000 hours of real gaming voice chat, generates sixdimensional

toxicity probability scores mapped to the Jigsaw

taxonomy via a novel cross-taxonomy semantic alignment, with

faster-whisper providing word-level timestamps for surgical muting

of precisely the toxic speech segments rather than entire

clips. For image moderation, CLIP ViT-B/32 encodes images

against toxicity-describing natural language prompts to produce

a nine-dimensional feature vector, with flagged content reposted

as blurred spoilers. Proximal Policy Optimisation reinforcement

learning agents trained with an asymmetric reward structure

penalising false negatives over false positives achieve 97.5%

accuracy on unseen audio evaluation data with 91.0% reward efficiency,

and 87.1% reward efficiency on image data, significantly

outperforming rule-based, random, always-mute, and alwaysallow

baseline policies. A per-user cross-modal trust score system

with progressive escalation is deployed as a real-time automated

Discord moderation bot validated through live user interactions.

Keywords
audio toxicity detection, image moderation, reinforcement learning, proximal policy optimisation, WavLM, CLIP, gaming platforms, content moderation, cybersecurity
Speaker
Vemula Yashodha
Amrita University

Post comments
Verification Code Change Another
All comments
Important Dates
  • Conference date

    07-30

    2026

    -

    08-01

    2026

  • 07-28 2026

    Draft paper submission deadline

  • 07-28 2026

    Registration deadline

Sponsored By

The United Societies of Science

Organized By

Kongunadu College of Engineering and Technology

Contact info
×

USS WeChat Official Account

USSsociety

Please scan the QR code to follow
the wechat official account.