Productive Playhouse

Productive Playhouse Productive Playhouse offers secure, premium data services in more than 300 languages worldwide.

Your model can pass every automated evaluation and still fail in the hands of your end users. Because the real failures ...
10/02/2026

Your model can pass every automated evaluation and still fail in the hands of your end users.

Because the real failures don't always show up in automated scoring. A response might be technically accurate, but a native speaker will immediately know that something is off.

We work with technical teams at AI companies to run human evaluation across languages, task types, and domains. With native-speaking linguists, clear guidelines, and QA processes built for the messiness of modern LLM outputs, together it's the layer that makes the difference between a model that performs well on paper and one that performs well in the world.

Productive Playhouse didn't start as a data company. We started in children's education, content and speech-language wor...
09/28/2026

Productive Playhouse didn't start as a data company. We started in children's education, content and speech-language work where every word shaped how a child understood their world.

In 2011, a technology company asked us to help teach machines English, and we learned that teaching a model well requires the same core principles as teaching a child well.

Fifteen years and thousands of native-speaking linguists later, that history is our greatest strength in solving the challenges of evaluating whether AI is safe for minors.

Because anyone can test an AI model. Understanding the child on the other side of the screen, how a 14-year-old talks, deflects, pushes boundaries and signals distress, is what we’ve spent nearly two decades learning.

Our U18 evaluation framework is built on that history:

✔️ Expert proxy reviews by clinicians, parents, and educators who deeply understand how kids communicate
✔️ Dynamic, multi-turn persona testing designed for real-life adolescent behavior
✔️ Native-language scoring across 350+ languages and dialects
✔️ Zero exposure of untested models to real minors

Keeping AI safe for the users with the most at stake and the least protection requires human-validated precision.

Learn how we evaluate U18 AI safety: [LINK]

We've spent nearly two decades understanding how children learn, think, and communicate. Now we're using that expertise ...
09/25/2026

We've spent nearly two decades understanding how children learn, think, and communicate. Now we're using that expertise to make sure AI systems keep them safe.

Most youth safety evaluations rely on static, single-prompt checks that assume users will explicitly declare their age. But how will your AI model respond when a teenager uses slang while saying they’re an adult, or asks the model to hide something from an adult?

Our approach to youth evaluation starts with a simple distinction because we don't start with the model, we start with the child. Grounded in over 15 years of child language research, speech therapy, and school operation:

✔️ We separate younger teens from older teens rather than treating all minors as a single user group; we test explicit, implicit, and absent age cues.
✔️ We simulate how real teenagers actually communicate, not how safety benchmarks expect them to, using realistic evasion behaviors like code-switching and spelling drift.
✔️ We apply expert proxy review using educators and developmental specialists to pressure-test models safely before youth-facing releases.

Safety that only works when a minor explicitly announces their age isn't going to keep them safe. Learn more about our approach: [LINK]

Beginning in late August, Uber lets parents livestream their teen's ride.Most AI safety testing is centered on the scree...
09/23/2026

Beginning in late August, Uber lets parents livestream their teen's ride.

Most AI safety testing is centered on the screen, but when your product puts a minor in a vehicle or in a DM with an adult stranger, testing needs to go beyond the model.

* What signals decide which adult meets which minor?
* What does the conversation between them allow?
* When something feels wrong, in this case mid-ride, what can a 15-year-old do about it? And who reads the ticket afterward, and would they recognize what they're reading?

Most minors rarely report; they're twice as likely to block a user as they are to tell an adult. And studies keep finding that most harmful incidents never generate a report at all.

What actually happens is a message to a friend, a vague "this driver is weird lol," or a canceled order with no explanation. Distress from a 15-year-old looks like noise unless your systems and the people testing them know what it looks like.

So when we test real-world service flows, we don't just probe the matching logic or the chatbot. We run the whole loop with testers who understand how kids actually flag trouble, which is usually a bit sideways and in language no keyword filter was trained on.

"Age appropriate" is now a property of the entire product. If your product touches the physical world then your safety testing must, too.

The factors that make U18 safety testing hard in English only get harder everywhere else:▪️ Self-harm slang is local and...
09/22/2026

The factors that make U18 safety testing hard in English only get harder everywhere else:

▪️ Self-harm slang is local and fast-moving. The heavily coded terms shift by country, by platform, and by month (or faster). A filter that's current in English is almost guaranteed to be behind in Hindi or Spanish, let alone a less common language like Tagalog.

▪️ Grooming is culturally specific. The trust-building playbook a predator uses in Brazilian Portuguese doesn't look like the English one your system was trained on. Your model won't recognize what it hasn't seen.

▪️ Even "age-appropriate" isn't stable. Norms around teen independence, when kids travel alone, work, date, or get a bank account, vary considerably across cultures. A flow that's fine for a 16-year-old in one market is a red flag in another.

Most companies handle this by translating their English test suite. But that only tests their translation vendor.

Real coverage means native-language experts building scenarios from scratch by people who know what the risky conversation actually sounds like in their language, because they've heard it.

When the FTC opened its Section 6(b) study into AI companion products, most coverage focused on the headline names of fr...
09/21/2026

When the FTC opened its Section 6(b) study into AI companion products, most coverage focused on the headline names of frontier models and major social platforms.

But buried in the FTC's orders to seven AI companies is a phrase every safety team should probably read twice, "all testing, red team exercises, audits, or analyses" related to age verification, monitoring and mitigation measures.

The Commission isn't asking companies whether their products are safe for minors. It's asking them to produce the evidence: how they tested before and after deployment, what mitigations triggered when a user appeared to be a minor, the actual prevalence of sexually themed outputs for young users, complaints suggesting a child was harmed, and concerns raised by their own employees and contractors.

If your company received this request tomorrow, what would you hand over? For a lot of teams, the answer is pre-launch evals written by adults, in English, plus reactive user reports.

The 6(b) study will take time to wrap-up, and it's now part of a larger wave including demands from state AGs, Senate hearings and private litigation all asking similar questions.

The companies in the strongest position a year from now will be the ones whose evaluation records show realistic, ongoing, developmentally informed testing.

At PPH, we build evaluations shaped the way kids actually talk, across the languages and cultures of your real user base, re-run whenever the model changes. If you're starting to think about what your evidence package looks like, we're happy to chat.

09/18/2026

When Common Sense Media's Youth AI Safety Institute tested popular AI toys for children, independent evaluators rated 27% of the outputs inappropriate for kids. This included self-harm references, drug mentions, boundary-crossing roleplay, and risky advice.

Yet when parents whose kids used those same toys, only 4% reported that the toy had ever said anything inappropriate.

That’s a huge disconnect, and it likely applies to every AI-enabled product used by kids because:

1. These harmful outputs usually happen in private, 1:1 conversations when no one else is in the room
2. Kids don’t report these issues. While the 2026 AI census found that one in six kids using AI chatbots had come across inappropriate content, only a third of children told a trusted adult.

So the absence of complaints isn’t evidence of safety.

And the exposure is massive, and growing. Over a third of kids have used AI to discuss their feelings or personal problems, and twenty percent of kids said it would be hard to stop using it for a month.

Operationally, if your evidence of U18 safety is largely based on low user reports, then you’re putting your programs and products at risk. The gap between what you think your product says to kids and what it actually says is now measurable, and regulators are starting to require the evidence of your testing and protections.

The way to understand what’s really happening is with proactive and realistic testing that reflects how kids talk across ages, cultures, and the languages of your user base. And these tests need to be re-run every time your model changes or you release product updates.

Reach out to learn how we can help you turn U18 testing into auditable, launch-ready evidence.

If you've seen our recent posts on what's broken with child-safety testing, you might be asking what rigorous testing ac...
09/17/2026

If you've seen our recent posts on what's broken with child-safety testing, you might be asking what rigorous testing actually looks like.

In our experience, five factors separate evaluations that uncover real risk from those that just produce green checkmarks:

1. Testing multi-turn conversations vs. a single prompt.

Sometimes, a study session drifts into a disclosure after the 17th back-and-forth because real risk emerges over the course of a conversation. Single-prompt tests don't reflect how kids actually use these products. Good evals apply multi-turn pressure, including what happens when a user reframes a request or patiently tries to work around safeguards over time.

2. Tests written the way kids actually communicate

This means that your test scripts have to include current slang, indirection, emotional subtext, and real evasive behaviors. Kids can be tricksters, and will often use roleplay, euphemisms, and code-switching. This is a linguistics problem before it's a safety problem.

And age matters. A 12-year-old and a 17-year-old aren't the same user, and neither is "a minor." Whether a user states their age, hints at it, or never mentions it changes model behavior, tests need to cover all three.

3. Multilingual and multicultural by default

Most safety testing happens in English, but many of the world's youth speak something else. Harms caught in English slip through in other languages, and cultural context can have big implications on what "inappropriate" means. Scenarios need to be adapted by native experts for local context, not translations of an English master copy.

4. Graded by experts who know and understand children

"Did the model refuse?" is an automated check. "Was this response developmentally appropriate, comprehensible, and safe for a distressed 13-year-old?" is not. That judgement takes reviewers with relevant backgrounds like child-development specialists, educators, and clinicians that can to catch when a cold, canned refusal to a young user in crisis is actually a failure.

5. Auditable evidence (archived reports aren't enough)

Because models evolve constantly, a safety evaluation from launch might reflect a product that effectively no longer exists. Testing and evaluation needs to run on an ongoing cadence to give product, policy, and legal teams a documented record of what was tested, what failed, under what conditions, and how behavior changes across releases.

These are the same standards you'd apply to any complex system: realistic conditions, qualified judges, repeated over time. It's just rare in U18 safety because it takes a combination of language reach, child-development expertise, and evaluation infrastructure that is a tall order for most teams.

If you're working on this inside a product or safety team and looking to build out evaluations for underaged users, reach out today to learn more.

Last week, Governor Newsom signed Adam's Law and safety teams should understand the basics of this regulatory framework....
09/15/2026

Last week, Governor Newsom signed Adam's Law and safety teams should understand the basics of this regulatory framework. With this law, child safety compliance for AI products stopped being a best practice and is now a legal requirement.

The underlying shift is clear and safety compliance for youth-facing AI now means documented, auditable risk evaluations.

Three things to watch now that this is law:

1️⃣ California sets the national baseline for AI minor safety, similar to what we saw with consumer privacy. Compliance starts July 1, 2027, which may sound far away, but building an evidence base for auditing is significant work.

2️⃣ This isn't a stand-alone bill; it was signed as part of a 13-bill package along with laws creating an independent AI audit framework and a registry of AI auditors. While OpenAI and Pinterest publicly backed this bill, FTC scrutiny continues.

3️⃣ There will be litigation impacts, with attorneys already expected to use "Did you evaluate this system with minors in mind?" as a standard discovery question.

We broke down what the bill requires and what product and safety teams need to do to prepare. Swipe through the carousel for the full breakdown.

And if you're looking to build an auditable evidence base for your model, product or chat tool, drop us a line to learn more about how our team can help.

The first clients who ever reviewed our work were parents.The procurement teams and technical program managers running q...
09/14/2026

The first clients who ever reviewed our work were parents.

The procurement teams and technical program managers running quality reviews came later. Parents, watching over their children's shoulders, were deciding in real time whether our programs and content were worth their kids’ attention.

Talk about an unforgiving QA process. A bored child will tell you exactly how well you did.

We eventually applied that same discipline to a very different kind of output: the training data behind multilingual AI systems, and we continued to bring the same level of precision and discipline we cultivated long before AI became mainstream.

When a model trains on low-quality data, it doesn't fail in obvious ways. Instead, it’s a slightly wrong transcription, a culturally misread phrase, or a dialect that was inferred rather than understood and the model learns the approximation and carries it forward.

Today we apply that same precision to a harder problem: how AI systems behave in conversation with children. The failures there are easy to miss, too. An overlooked disclosure, a tone that's right for an adult but wrong for a 12-year-old, or a safety response that technically works but emotionally misses the mark.

Learn how we help teams uncover where models fail adolescent users before those failures reach the real world. LINK

Address

Mail To: PO BOX 27250
Los Angeles, CA
90027

Alerts

Be the first to know and let us send you an email when Productive Playhouse posts news and promotions. Your email address will not be used for any other purpose, and you can unsubscribe at any time.

Contact The Business

Send a message to Productive Playhouse:

Shortcuts

Share