WayToClawEarn
High impactORAVYS / Hacker News / 华尔街日报

Mercor 4TB voice data leak: 40,000 AI contractors’ voiceprints and IDs were stolen at the same time, 5 class-action lawsuits have been filed

Hacker group Lapsus$ leaked approximately 4TB of contractor data from the AI ​​data annotation platform Mercor, involving more than 40,000 people. This is the first time in the industry that high-fidelity voice + ID document + facial selfie - the complete raw materials for a deepfake attack - have been leaked simultaneously.

WayToClawEarn EditorialPublished Apr 28, 2026Updated Aug 8, 2026

Editorial review of public sources · AI-assisted drafting. How we work · Original source

Core conclusion

In early April 2026, the hacker group Lapsus$ leaked approximately 4TB of AI contractor data on the dark web, involving more than 40,000 annotators on the Mercor platform. This batch of data is far more dangerous than ordinary data leaks: for the first time, it combines and stores "high-fidelity voice samples + government ID scans + face selfies", providing a complete set of raw materials for deepfake attacks. Five class-action lawsuits have been filed within 10 days of the leak.

Key Points

  • Incident time: April 4, 2026 (leak date)
  • Breach size: ~4TB, covering 40,000+ Mercor contractors -Leaked content: 2-5 minutes of studio quality voice acting per person + passport/driver's license scan + webcam selfie
  • Attacking organization: Lapsus$

Event details

Mercor is an AI data annotation platform that provides training data services to global AI companies. The contractor's registration process requires the submission of: a scan of a passport or driver's license, a webcam selfie, and a voice recording of a script being read in a quiet environment. These three types of data are stored linked in the same row in Mercor’s database—which happens to be all the input material needed for the voice cloning service.

The Wall Street Journal reported in February 2026 that high-quality voice cloning today only requires about 15 seconds of clean reference audio. And each recording Mercor leaked averaged 2-5 minutes long — well above the threshold. With the addition of a verified identity document, an attacker can both clone the voice and use the document information to increase the fraud credibility of the clone.

This is not the first AI training data leak, but ORAVYS (the voice forensics company that discovered the incident) emphasized in its report: "Voice leaks in the past decade were either call center recordings stolen but unable to be linked to identities, or ID documents leaked but without audio. Mercor merged the two." This "ID + voice + selfie" trinity leak is one of the most serious data security incidents in the industry to date.

Key Impact

DimensionsChangeWhat it means to usRecommended actions
Personal securityBiometrics (voiceprint + face) of 40,000 people have been leakedVoiceprints cannot be changed, unlike passwordsMonitor abnormal voice calls immediately and be wary of AI voice phishing
Enterprise risksBank/enterprise voice authentication systems are facing threatsThe reliability of voiceprint verification systems has been greatly reducedEvaluate upgrading to multi-factor authentication and reduce pure voice verification
Industry trustThe security standards of the AI data annotation industry are questionedUser confidence in AI companies' data protection capabilities has declinedCheck security compliance certification when choosing a data service provider
Legal5 class-action lawsuits filed"Informed consent" for data collection will face tighter scrutinyCompliance teams should update privacy terms in advance

Types of attacks possible with this leak

The following attack methods have been confirmed to be used in the wild by ORAVYS:

  1. Voice cloning fraud: Use a recording of more than 15 seconds to clone the victim’s voice, and combine it with ID information for bank phone transfer verification
  2. Deepfake Video Call: 2024 Arup Finance employee wired ~$25M following multi-person deepfake video call - Mercor leak provides better training data than publicly available footage
  3. Identity Forgery: Scanned copies of driver’s licenses can be directly used to open accounts, apply for loans and other traditional identity thefts

Adaptation suggestions

Protection suggestions for AI practitioners

  • If you are registered as a contractor with Mercor, immediately monitor your credit report and bank account for any abnormalities
  • Stay alert for "voice verification calls" from family members, even if the voices sound completely authentic
  • All sensitive operations (transfers, password changes) should enable hardware keys or multi-factor authentication and do not rely on voice verification

Implications for AI content producers

  • Data privacy compliance is not an optional configuration, but a life-or-death line for business
  • When using AI tools to process sensitive data, check their Data Processing Agreement (DPA)
  • Consider incorporating data masking steps into automated workflows

Further reading

The two-sided nature of the AI industry: the more powerful the tool, the heavier the responsibility.

This voice data leak is in sharp contrast to Taylor Swift’s trademark registration of a sound at the same time. Swift's legal team is trying to protect sound rights through trademarks - IP lawyers noted celebrity sound trademarks "have not yet been tested in court." This shows that when AI synthesis capabilities run ahead of the law, both individuals and companies need to proactively establish protection mechanisms.

Security blind spots in the data annotation industry

The Mercor incident exposed systemic risks in the AI data annotation industry: in order to provide high-quality training data, annotation platforms often require contractors to submit biometric information that far exceeds that required for actual annotation tasks. As an HN community comment noted: "Contractors submit studio-quality voice and ID scans solely for data annotation—data that has nothing to do with the annotation work."

Tool entry

The following tools are involved in the text, and the platform side will automatically match the maintained tools library to trigger the floating card:

  • OpenAI, ChatGPT, Claude, DeepSeek, Gemini, n8n, LangGraph, Hermes Agent

Internal link guidance

View source →

Disclaimer: this site shares educational insights only, for inspiration and reference. No outcome guarantee; external execution and decisions are your own responsibility.