Read · beginner
When privacy and AI safety stop being opposites
What OpenAI's Private Safety Processing preview teaches us about building AI workflows that protect data and remain observable.
The most interesting AI news this week was not another model leaderboard. It was a privacy announcement: can a provider detect dangerous patterns across an agent’s work without keeping the user’s prompts and answers around for people to inspect?
On August 19, OpenAI announced a preview of Private Safety Processing alongside its Zero Data Retention (ZDR) offering for eligible API customers. The company says the new system is designed to spot patterns across related interactions using automated processing, while keeping the underlying customer content away from OpenAI personnel.
That is a useful shift in the conversation. Privacy is often described as if it means turning off observability, while safety is described as if it requires collecting everything. This announcement suggests a more demanding design question: what is the smallest signal a safety system needs, and who needs to see it?
What was announced
The important details are easy to miss in the headline:
- ZDR is for eligible API customers. OpenAI says it does not retain their prompts or model responses after processing, and that the content is not available to OpenAI personnel for review.
- Private Safety Processing is a preview, not a generally available switch that every project can use today.
- The system is meant to look across related interactions. That matters for long-running agents, where an unsafe pattern may not be visible in one request.
- OpenAI says the service can return a narrowly defined safety signal instead of exposing the underlying prompts and responses to its staff.
- The company says customer content can stay on infrastructure controlled by the customer. It is also developing an option that stores content on OpenAI infrastructure while using customer-controlled encryption keys.
Those are claims from the provider’s announcement, not an independent audit. The feature is being tested with early customers, and OpenAI says it plans to publish a technical white paper in September. Treat the preview as a direction to evaluate, not as a guarantee you can build against yet.
Why this matters more for agents
A short completion is relatively easy to reason about: one input, one output, one decision. An agent is different. It may read several sources, call tools, write intermediate files, retry a failed action, and continue for a long time.
That creates two problems at once:
- Safety needs context. A single tool call may look harmless even when a sequence of calls shows an attempt to reach a restricted system or extract sensitive data.
- Privacy needs boundaries. The context that makes the pattern visible can contain personal data, source code, customer records, or confidential plans.
The tempting answer is to retain every interaction forever and let a human review the logs. It is operationally simple, but it creates a large collection of sensitive material and a new target for misuse.
The more interesting answer is to separate content from signals. A safety system may need to know that a sequence crossed a risk boundary without giving an operator the full text of every step. That does not solve every problem, but it gives privacy a place in the architecture instead of adding it after the fact.
The design lesson: observability is not the same as surveillance
When you build an AI workflow, “log everything” is usually a shortcut for not having decided what you actually need to debug or protect.
Start with four questions:
- What must the system inspect to detect abuse or a failed workflow?
- What can be represented as structured metadata instead of raw content?
- Which people or services need access to each signal?
- How long does each piece of information need to exist?
For example, a support agent may need to record that it attempted a refund, which policy version it used, and whether a human approved the action. It may not need to retain the customer’s entire conversation in a central log for months.
This is the same separation used in good application design:
- Content: the prompt, response, document, transcript, or source code.
- Event: “tool call requested,” “approval denied,” or “policy check failed.”
- Decision: the validated action your application allowed.
- Signal: a coarse indicator that a safety or abuse rule needs attention.
Keeping those layers distinct makes it easier to choose retention, access, and encryption rules for each one. It also makes a later provider change less painful because your application is not treating one vendor’s logging model as its domain model.
A practical pattern for a privacy-conscious agent
You do not need a frontier safety system to apply the principle. For a small agent, begin with a narrow event record and keep raw content in the smallest possible boundary:
type AgentEvent = {
runId: string;
tool: string;
action: "requested" | "approved" | "denied" | "completed";
policyVersion: string;
riskSignal?: "none" | "review" | "blocked";
createdAt: string;
};
This is a design example, not a provider-specific API. The useful properties are the boring ones:
runIdlets you reconstruct a workflow without making every log consumer read the full prompt history.toolandactionmake authority visible.policyVersiontells you which rule produced a decision.riskSignalcan route work to review without copying the entire payload into a shared dashboard.- A timestamp supports investigation and expiry.
Keep the actual content behind a separate access boundary. Encrypt it, restrict who can read it, and delete it according to a written retention rule. Do not assume that hiding a field in a UI is the same as removing access in storage.
Five controls worth adding now
1. Define retention per data type
Write down separate periods for prompts, outputs, tool arguments, event metadata, safety alerts, and audit records. “We keep logs for 30 days” is not a policy until you say which logs and why.
2. Make tool authority explicit
An agent should not inherit every permission available to the application. Give each tool a narrow scope, validate arguments outside the model, and require a human approval for actions that are costly, public, or hard to undo.
3. Prefer deterministic checks where possible
Use schemas, allowlists, rate limits, access controls, and database constraints for rules that do not require language understanding. A model can propose an action; application code should decide whether the action is allowed.
4. Log decisions, not secrets
Record that a check passed or failed, which rule ran, and what action followed. Avoid copying API keys, full documents, personal data, or complete tool payloads into every log sink just because the logger makes it easy.
5. Test the boundary, not only the answer
Run scenarios where an agent is asked to exceed its authority, continue after a stop request, reveal hidden context, or chain individually harmless actions. Check both outcomes: did the agent stop, and did your telemetry avoid creating a second privacy problem?
What this announcement does not mean
It does not mean that privacy controls make an agent safe by themselves. A workflow can leak data through its own database, browser history, traces, error reports, or third-party tools even when the model provider retains nothing.
It does not mean that a narrow safety signal is automatically understandable. People investigating an alert still need enough evidence to challenge a false positive and correct a broken rule. Privacy is not an excuse to make decisions unreviewable; it is a reason to design review access deliberately.
And it does not mean that a preview is ready for a production architecture. Check availability, eligibility, regional handling, contractual terms, and the provider’s current documentation before relying on any retention promise.
The takeaway
This week’s useful discovery is a design direction, not a shiny new model: safety systems may be able to work with carefully bounded signals instead of unlimited access to raw conversations.
For builders, the move is practical. Separate content from events. Give tools only the authority they need. Make retention explicit. Keep deterministic controls outside the model. Then test whether your monitoring helps a human understand what happened without turning every user interaction into permanent surveillance.
That is a higher bar than “we have logs.” It is also a much better starting point for AI systems people can trust.