OpenAI Is Firing the Contractors Who Used AI to Train Its AI
AI & ML

OpenAI Is Firing the Contractors Who Used AI to Train Its AI

Internal documents obtained by 404 Media show OpenAI terminating contractors for using tools like GPTZero and Grammarly while rating ChatGPT responses, a rule meant to protect the integrity of human feedback data that trains the model.

PublishedSeptember 24, 2026
Read time6 min read
Share

The rule that created the firing wave

The instruction is unambiguous in the internal documents 404 Media reviewed: do not use AI detection tools, or AI yourself. That single line covers a surprisingly wide net, explicitly naming GPTZero, Grammarly, and AI translation tools alongside the more obvious prohibition on using a chatbot to write or check a response rating outright before submitting it for review. The breadth of the ban signals that OpenAI has already seen contractors try to route around narrower rules by using tools that do not look like generative AI on their face, and closed those loopholes one by one as each new workaround surfaced during internal review of contractor output patterns over time.

More than 10,000 contractors work across the various rating projects this rule applies to, according to the internal documents 404 Media reviewed as part of its reporting. At that scale, even a small percentage of violations represents a meaningful volume of corrupted training data flowing into the model's alignment process, which is presumably why enforcement has been aggressive enough that contractors describe it internally as the single fastest way to lose the work entirely, faster even than most other quality violations on the platform.

How they actually get caught

The detection methods described are more behavioral than technical, which is itself a telling design choice given the underlying problem the company is trying to catch. Reviewers look for repetitive language patterns across a contractor's ratings, distinctive AI writing habits like overuse of em dashes, and task completion times that are implausibly fast for someone genuinely reading and evaluating a response line by line before scoring it. None of these signals require running the contractor's own output through a separate AI detector, which is notable given how unreliable those tools are known to be when used as a primary enforcement mechanism rather than one weak input weighed alongside several others in a broader review process.

Internal Slack channels reportedly feature contractors flagging suspicious examples to each other, asking whether a given rating looks AI-generated, with affirmative answers common enough among peers to suggest the underlying problem is genuinely not rare across the workforce. That peer-level awareness, combined with management-side pattern detection, describes an enforcement environment that relies on triangulating multiple weak signals rather than leaning on one definitive test, which is the realistic approach given that no reliable, generally trusted AI-detection tool actually exists on the market today.

Why this matters more than a labor dispute

It is tempting to read this as a standard contractor labor story, workers cutting corners on tedious piecework and getting fired for it once management noticed the pattern in their output. The more consequential angle is what it reveals about the fragility of the data pipeline underneath every frontier model's alignment process today. Reinforcement learning from human feedback is only as good as the human part of that phrase, and that assumption is doing more load-bearing work than most model documentation admits publicly. If a meaningful fraction of the raters generating that feedback are themselves using AI to produce or check their ratings, the model is partially learning from itself in a loop that nobody explicitly designed and that degrades signal quality in ways genuinely hard to detect after the fact, since the corrupted ratings look statistically similar to legitimate human judgment on the surface.

One fired contractor's own explanation captures the dynamic well: they described using AI as needing a little boost, framing it as an ordinary productivity shortcut rather than any kind of deliberate sabotage of the process. That is almost certainly how most violations actually happen in practice, workers under time and volume pressure reaching for tools that make repetitive work faster, without necessarily intending to corrupt anything downstream. The aggregate effect on model training data ends up the same either way, regardless of the individual contractor's intent going in.

The Project Lily context makes this worse, not better

This reporting builds directly on earlier revelations about Project Lily, in which hundreds of contractors were found to be reading actual user ChatGPT conversations containing personal information as part of a separate review workflow entirely, one built for a different purpose than the rating work described here but staffed through much of the same contractor pipeline. Taken together, the two stories describe a human-feedback pipeline facing scrutiny on two fronts simultaneously: whether contractor access to real user data is adequately controlled in the first place, and whether the ratings those same contractors produce can even be trusted once access controls are working exactly as intended.

Neither problem is unique to OpenAI in any meaningful sense. Every major AI lab runs some version of a human feedback pipeline at scale, using contractor workforces managed through intermediaries like Mercor to keep costs manageable. The specific failures reported here happen to be OpenAI's, but the structural vulnerability underneath them, a large distributed contractor workforce doing high-volume, low-oversight rating work that directly shapes model behavior, is shared across essentially the entire industry building frontier models today.

What enterprise AI buyers should ask vendors

For CIOs and CTOs evaluating frontier model vendors, this is a useful prompt to ask a question that rarely comes up in procurement conversations: what does your human feedback pipeline actually look like end to end, and what controls exist to prevent exactly this kind of quiet data corruption from creeping in over time as volume scales across thousands of contractors? Model cards and safety documentation rarely address the operational integrity of the human labeling workforce at all, even though that workforce directly shapes the values and behaviors the model ultimately exhibits once it reaches production traffic and real customers depending on it.

OpenAI declined to comment on the contractor terminations when 404 Media reached out, which is a fairly standard non-response but leaves open exactly how widespread the underlying problem was before enforcement finally caught up with it at scale. Enterprises building their own RLHF-style feedback loops internally, a growing practice as more companies fine-tune models on proprietary data for narrower use cases, should treat this reporting as a preview of a problem they will eventually face at smaller scale themselves, whether they staff the work with contractors or their own employees.

Tagged#news#ai-ml#ai#llm#agents#agentic-ai#openai#anthropic#regulation#rlhf#contractors#mercor#data-integrity#chatgpt