3 AI Labs Drop a Safety Playbook—But It's Just Step One
Curated by the Inblix editorial team
Cohere, OpenAI, and AI21 Labs just put out a joint statement that’s less a finished manual and more of a stake in the ground. They’re calling it a “preliminary set of best practices” for deploying large language models, and the tone is refreshingly honest: these systems are powerful, they’re going to mess up in ways we can’t predict, and the guidelines will need to change fast. That candor alone is noteworthy in an industry prone to breathless hype.
The document is structured around seven principles that are sensible, if not revolutionary. Publish real usage guidelines that ban stuff like spam, fraud, and astroturfing. Build infrastructure to actually enforce those rules—think rate limits, content filters, and application approvals before anyone gets production access. Proactively test for harmful behavior using techniques like human feedback. Document what’s broken, because no amount of safety work eliminates every risk. Hire diverse teams to catch the biases homogeneous groups miss. Share lessons learned publicly so the whole industry iterates faster. And here’s the one that doesn’t get enough airtime: treat the humans in the supply chain—the labelers reviewing toxic outputs—with actual respect, including letting them opt out of tasks.
Anthropic’s one-line endorsement gets quoted directly: “While LLMs hold a lot of promise, they have significant inherent safety issues which need to be worked on.” John Bansemer from CSET chimed in too, calling the principles “an important contribution to the global conversation.” Google’s statement is the most corporate of the bunch—affirming the “importance of comprehensive strategies” without committing to anything specific. The real story here isn’t the principles themselves, most of which any responsible AI team should already be doing. It’s the public coordination. When three competitors co-author a document like this, they’re signaling that voluntary guardrails are preferable to the kind of regulation that Congress keeps threatening but can’t seem to pass. The question is whether “preliminary best practices” evolve fast enough to keep pace with models that are scaling at a breakneck clip.
💡 Key Takeaways
- The joint statement from OpenAI, Cohere, and AI21 Labs explicitly bans high-risk use cases like classifying people based on protected characteristics.
- The principles emphasize that no amount of preventative action can completely eliminate unintended harm, so documenting known weaknesses is non-negotiable.
- The document calls for enforceable infrastructure—rate limits and content filters—not just aspirational policies, which marks a shift toward operationalizing safety.
- A standout principle demands fair treatment of human labelers, including the right to opt out of reviewing disturbing content, addressing a long-ignored labor issue.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.