OpenAI's Safety Theater: Why 'Safe and Aligned' Deflects from the Real Privacy Tradeoff

OpenAI's Safety Theater: Why 'Safe and Aligned' Deflects from the Real Privacy Tradeoff

On August 25, 2026, TechCrunch published an exclusive interview with Thibault Sottiaux, OpenAI's head of product. The conversation covered agents, user experience, pricing strategy, and market leadership. ChatGPT Work had just hit 20 million users. Everything felt inevitable, triumphant.

Then the interviewer asked a direct question:

"What do you say to someone who is stressed out about giving it access to their e-mail, giving it access to iMessages, which you just rolled out today?"

Sottiaux's response:

"It's important to pick models that are safe and aligned. And a very, very big part of our investment is in the safety stack, the safety approach, publishing honest benchmarks on these things. And our models are world-class at these topics."

He did not answer the question.


What The Question Actually Was

The interviewer asked: Why do you need access to my email and iMessages?

This is a straightforward question about necessity and tradeoff. It's asking: What can ChatGPT do with email access that it couldn't do without it? Is that capability worth the privacy cost?

It's not asking whether OpenAI's models are safe. It's asking whether the access is necessary.


What Sottiaux Actually Answered

Sottiaux answered a different question: Are your models safe and aligned?

His answer: Yes, we invest heavily in safety, alignment, benchmarking, and our models are world-class.

All of that might be true. But it's irrelevant to the original question.

A safe model with unnecessary data access is still a privacy risk. A world-class safety stack doesn't answer why you need email access to accomplish your stated task.

This is deflection. Not evasion—the answer was direct and confident. But deflection nonetheless.


The Instinct Contrast (48 Hours Earlier)

On August 24, Instinct AI faced the exact same privacy criticism.

Instinct's response: They explained the architecture.

When asked why they need broad data access, they didn't claim to be "safe and aligned." They said: For an AI to autonomously book your meeting, it needs calendar context, past commitments, email patterns, location, availability. For it to send emails on your behalf, it needs to understand tone, relationships, communication style. Without this data, autonomy is impossible.

That's a real answer. Uncomfortable, but real.

Instinct acknowledged the tradeoff: maximum autonomy requires maximum data access. They chose autonomy. They defended the choice.

OpenAI took the opposite approach: Don't address the tradeoff. Pivot to safety.


Why This Is Deflection, Not Defense

A real defense would look like this:

"Yes, ChatGPT needs email access for certain tasks. Here's specifically what it accesses: email headers and message bodies for booking context. Here's how we protect that data: end-to-end encryption, deletion after 24 hours, no training data use. Here's your recourse: opt-out is one click. Here's the liability: we're covered under our insurance policy. It's a deliberate tradeoff—we chose capability over data minimization."

That would be a defense. Uncomfortable, but direct.

Sottiaux's answer addresses a different concern entirely. It's the conversational equivalent of saying: "Your question is about privacy, but let me tell you about safety instead. These sound related so you might not notice I changed the subject."

The technique works because safety and privacy are emotionally linked (both good things) but logically separate concepts.


What "Safe and Aligned" Actually Means

Safe: The model can't be jailbroken. It won't output harmful content. It resists adversarial prompts.

Aligned: The model follows user intent. It respects user values. It doesn't contradict the user's goals.

Private: The model uses minimal data. It accesses only what's necessary. It doesn't retain data unnecessarily.

These are three different dimensions. A model can be perfectly safe and aligned while still unnecessarily accessing data.

Example: A safe, aligned AI that reads your entire email inbox to answer one question about logistics. It can't be jailbroken (safe), it respects your intent (aligned), but it accessed unnecessary data (not private).

Sottiaux conflated these dimensions. He answered a question about safety and called it an answer to a question about access necessity.


The Corporate Playbook

Here's the pattern Sottiaux employed:

  1. Uncomfortable question: Why do you need this access?
  2. Emotional pivot: Let me talk about safety instead (also a good thing).
  3. Confidence signal: We're "world-class" (vague, but sounds authoritative).
  4. Measurable proof: We publish benchmarks (sounds rigorous).
  5. End result: Question feels answered, even though it wasn't. Founders will recognize this playbook. It's useful when you don't want to defend an actual choice. When transparency would highlight an uncomfortable tradeoff, deflection can work. When the deflection is subtle enough, audiences don't notice the pivot.

Instinct AI tried transparency instead. They got criticized anyway. So did OpenAI's strategy work better? Maybe. Or maybe it just delayed the reckoning.


The Market Baseline Effect

Here's the strategic implication:

OpenAI reaches 20 million users with a response that doesn't address privacy concerns. Most users don't notice the deflection. They hear "safe" and "world-class" and move on.

Market learns: Broad data access is acceptable if you claim to be safe.

Competitors watch: Sottiaux's playbook worked. The question disappeared. No controversy. Instinct got backlash for trying to explain.

Investors note: OpenAI's communication strategy is effective. Transparency created problems; deflection didn't.

Next company facing the same question learns the lesson: Don't answer. Pivot to safety.

Market baseline shifts. "Safe and aligned" becomes sufficient answer to "why do you need access?" It becomes normal. The question stops being asked.


What Founders Should Learn

Two lessons:

First: When asked an uncomfortable question, notice what question you're actually answering. Transparency is risky because it highlights real tradeoffs. Deflection is safer because it avoids the tradeoff entirely.

Second: Audiences are sophisticated enough to notice the pivot. Instinct got criticized for the privacy terms themselves, not for explaining them. OpenAI might get criticized for dodging the question, not just for having broad access.

Better approach: Answer directly. Acknowledge the tradeoff. Explain why you made that choice. Trust your audience to respect honest decision-making even if they disagree with it.

Instinct did this. They got criticized. But they also got respected for being direct.

OpenAI avoided the criticism. But they also trained their users not to expect direct answers.

Which builds better trust? That's a founder's strategic choice.


The Question That Matters

The real question—the one that still hasn't been answered—is this:

What can ChatGPT do with email access that it couldn't do with explicit user instructions? Is that capability worth the data cost?

If the answer is "We could build the same capabilities without email access, but it would require more user interaction," then the question becomes: Why did you choose convenience over privacy?

That's a legitimate business decision. But it's not a safety question. It's an incentive question.

And that's exactly what Sottiaux avoided answering.

For more on how AI companies make tradeoffs and what those choices reveal about their incentives, visit Bitroot.