09/17/2026
If you've seen our recent posts on what's broken with child-safety testing, you might be asking what rigorous testing actually looks like.
In our experience, five factors separate evaluations that uncover real risk from those that just produce green checkmarks:
1. Testing multi-turn conversations vs. a single prompt.
Sometimes, a study session drifts into a disclosure after the 17th back-and-forth because real risk emerges over the course of a conversation. Single-prompt tests don't reflect how kids actually use these products. Good evals apply multi-turn pressure, including what happens when a user reframes a request or patiently tries to work around safeguards over time.
2. Tests written the way kids actually communicate
This means that your test scripts have to include current slang, indirection, emotional subtext, and real evasive behaviors. Kids can be tricksters, and will often use roleplay, euphemisms, and code-switching. This is a linguistics problem before it's a safety problem.
And age matters. A 12-year-old and a 17-year-old aren't the same user, and neither is "a minor." Whether a user states their age, hints at it, or never mentions it changes model behavior, tests need to cover all three.
3. Multilingual and multicultural by default
Most safety testing happens in English, but many of the world's youth speak something else. Harms caught in English slip through in other languages, and cultural context can have big implications on what "inappropriate" means. Scenarios need to be adapted by native experts for local context, not translations of an English master copy.
4. Graded by experts who know and understand children
"Did the model refuse?" is an automated check. "Was this response developmentally appropriate, comprehensible, and safe for a distressed 13-year-old?" is not. That judgement takes reviewers with relevant backgrounds like child-development specialists, educators, and clinicians that can to catch when a cold, canned refusal to a young user in crisis is actually a failure.
5. Auditable evidence (archived reports aren't enough)
Because models evolve constantly, a safety evaluation from launch might reflect a product that effectively no longer exists. Testing and evaluation needs to run on an ongoing cadence to give product, policy, and legal teams a documented record of what was tested, what failed, under what conditions, and how behavior changes across releases.
These are the same standards you'd apply to any complex system: realistic conditions, qualified judges, repeated over time. It's just rare in U18 safety because it takes a combination of language reach, child-development expertise, and evaluation infrastructure that is a tall order for most teams.
If you're working on this inside a product or safety team and looking to build out evaluations for underaged users, reach out today to learn more.