What AI Safety Teams Actually Do (And Why It Never Makes Headlines)
When AI safety makes the news, it's usually about dramatic scenarios. Killer robots. Existential risks. Alignment researchers debating the fate of humanity. Meanwhile, actual AI safety teams are doing something far less cinematic: writing test cases, running red team exercises, and documenting edge cases in spreadsheets.
If you're a business leader trying to understand what an AI safety function actually does, you've probably noticed a gap. The public conversation is either theoretical philosophy or breathless hype. The reality is more mundane, more specific, and more useful.
Table of Contents
- What Safety Teams Actually Spend Their Time On
- Red Teaming Is Not What You Think
- Evaluation Design: The Unglamorous Core
- Incident Response Basics
- Why This Work Stays Invisible
- What Business Leaders Should Actually Care About
What Safety Teams Actually Spend Their Time On
Most AI safety work happens in three overlapping areas: evaluation design, red teaming, and incident response. None of these are glamorous. All of them are necessary.
Evaluation design means building tests that check if a model does what you think it does. Not theoretical alignment tests. Practical checks. Does it refuse harmful requests consistently? Does it hallucinate less often than the previous version? Does it maintain performance across different user groups?
Analogy: Think of it like car safety testing. Crash test dummies aren't exciting. But someone has to design the tests, run them hundreds of times, and document what breaks.
Red teaming means trying to break the system on purpose. Safety teams hire people to find failure modes before users do. This involves writing prompts designed to bypass guardrails, testing edge cases, and documenting every weird behavior.
Incident response means having a plan when things go wrong. Because they will. Every model ships with unknown failure modes. Safety teams build systems to detect, triage, and fix problems quickly.
Here's roughly how safety teams allocate their time:
| Activity | Approximate Time | Purpose |
|---|---|---|
| Evaluation design | 40% | Building reliable tests |
| Red teaming | 30% | Finding failure modes |
| Incident response | 20% | Fixing live issues |
| Documentation | 10% | Recording what works |
Red Teaming Is Not What You Think
When most people hear "red teaming," they imagine hackers in hoodies trying to jailbreak ChatGPT with clever prompts. That happens. But it's maybe 20% of the work.
Real red teaming is systematic. Safety teams build libraries of test cases organized by risk category. Hate speech. Misinformation. Privacy violations. Dangerous instructions. Each category gets hundreds of variations.
The boring part: most prompts don't work. A red teamer might write 50 variations of a harmful request and get 50 refusals. Then they try number 51, and something breaks. That one broken case gets documented, fixed, and added to the regression test suite.
The even more boring part: after the fix, they run all 51 tests again. Plus the other 2,000 tests in the suite. Because fixing one failure mode sometimes creates new ones.
Red teamers aren't just creative prompt writers. They're methodical testers who understand model behavior, track patterns across failures, and communicate clearly with engineering teams. Much of their time goes into documentation and coordination, not clever attacks.
Here's what makes red teaming effective:
| Element | Why It Matters |
|---|---|
| Systematic coverage | Random testing misses entire risk categories |
| Version tracking | Need to compare across model updates |
| Clear documentation | Engineers can't fix vague bug reports |
| Regression testing | Fixes shouldn't break existing safeguards |
Evaluation Design: The Unglamorous Core
Evaluations are the foundation of safety work. If you can't measure a problem reliably, you can't fix it. And measuring AI behavior is harder than it looks.
Say you want to test if a model refuses harmful requests. First question: what counts as harmful? Different cultures have different norms. Different use cases have different risk profiles. A medical AI and a creative writing AI need different safety boundaries.
Second question: how do you score responses? A flat "refuse or not" binary misses important details. Some refusals are clear and helpful. Others are vague or preachy. Some technically refuse but leak information anyway.
Safety teams spend weeks designing evaluation frameworks that capture these nuances. They write rubrics. They test inter-rater reliability. They argue about edge cases. It's the opposite of exciting.
But it matters. Without good evaluations, you're flying blind. You can't compare model versions. You can't track improvement. You can't make informed tradeoffs between capability and safety.
A typical evaluation pipeline looks like this:
Incident Response Basics
No model is perfect. Safety teams know this. So they build systems to handle problems when they appear in production.
Incident response starts with detection. Something goes wrong. A user reports unexpected behavior. An automated monitor flags unusual patterns. A news story surfaces a problem.
Then comes triage. Is this a new issue or a known edge case? How severe is it? How many users are affected? Can we reproduce it?
Next: containment. Sometimes that means pulling a feature. Sometimes it means updating prompts or filters. Sometimes it means adding the case to a blocklist while engineers work on a proper fix.
Finally: root cause analysis. Why did this happen? What can we change to prevent similar issues? How do we test for this going forward?
Most incidents are small. A weird edge case that affects 50 users. A prompt that bypasses a filter in a narrow context. A hallucination pattern that shows up in specific domains.
But small incidents add up. Safety teams track patterns across incidents. Five unrelated edge cases might reveal a systematic problem. Effective incident response isn't about heroic fixes. It's about systematic learning.
Why This Work Stays Invisible
Safety work is invisible when it succeeds. Nobody writes headlines about "Model Ships Without Major Safety Incidents." The absence of problems isn't newsworthy.
When safety work does make news, it's usually because something broke. A model said something offensive. A filter blocked legitimate content. A safety feature caused a backlash.
This creates a perception problem. Successful safety work looks like nothing happened. Failed safety work looks like incompetence. The actual daily work of building robust evaluations, running systematic tests, and responding to incidents stays hidden.
Safety teams also can't share much publicly. Detailed safety reports would be roadmaps for attackers. Specific failure modes become exploit guides. Most safety work happens behind closed doors because transparency would undermine effectiveness.
What Business Leaders Should Actually Care About
If you're building an AI function or evaluating AI vendors, here's what matters:
Do they have systematic evaluation processes? Not just vibes-based testing. Documented test suites. Version-to-version comparisons. Clear scoring rubrics.
Do they run regular red teaming? Internal teams trying to break things. External researchers invited to test. A clear process for handling findings.
Do they have incident response plans? Not just for catastrophic failures. For everyday edge cases. Detection systems. Triage processes. Clear ownership.
Do they document and learn? Safety work should improve over time. Each incident should feed into better evaluations. Each red team exercise should strengthen defenses.
The goal isn't perfect safety. That doesn't exist. The goal is systematic improvement. Fewer surprises. Faster response. Better understanding of where your models break.
AI safety teams won't eliminate risk. But they make risk manageable. They turn unpredictable systems into somewhat-predictable ones. They build the testing infrastructure that lets you ship with confidence.
That's not headline material. But it's the actual work. And if you're deploying AI at scale, it's the work that matters.