AI Pulse by Inblix

80 examples is all it takes to reshape a large AI model's values

OpenAI Blog · Jul 19, 2026 · 2 min read · Read original article →

Curated by the Inblix editorial team


Featured image for article: 80 examples is all it takes to reshape a large AI model's values

OpenAI has published new research demonstrating that a language model’s behavior can be significantly steered by fine-tuning it on a remarkably small, hand-curated dataset. The paper reveals that a set of just 80 text samples, formatted as question-and-answer pairs, was enough to improve adherence to specific behavioral values in GPT‑3 without degrading its performance on other tasks. The dataset, which is about 120KB in total size—a fraction of a percent of GPT‑3’s original training data—targeted categories with direct human impact, including violence, injustice, health advice, and political destabilization. The values were largely drawn from U.S. and international human rights law and Western social movements for equality. OpenAI found a clear scaling effect: the process became more effective as the models grew larger, from 125 million up to 175 billion parameters. This suggests operators of massive models could need surprisingly few examples to align behavior to their application’s specific needs. The company is now actively looking for API users to test the nascent technique in production, framing the selection of exact values not as a one-size-fits-all problem, but as a choice each developer must confront. Their qualitative probes showed the values-targeted models broadly adhered more to the desired behaviors, with statistically significant improvements measured through human evaluations, toxicity scoring via the Perspective API, and co-occurrence metrics for gender, race, and religion. The researchers are candid about the limitations, noting that toxicity scores carry known demographic biases, which is why they layered multiple evaluation methods. They also acknowledge that outlining values for large groups risks marginalizing minority voices, a tension they attempted to address by making the process scalable rather than requiring a full retraining from scratch.

💡 Key Takeaways

  1. Fine-tuning on a curated dataset of fewer than 100 examples can yield statistically significant changes in a large language model's behavior without hurting downstream task performance.
  2. The technique exhibits a scaling law in reverse: larger models require comparatively fewer data samples to adapt their behavior to a targeted set of values.
  3. OpenAI frames value selection as a user's responsibility, providing this tool so developers can constrain a model to a specific, application-appropriate behavioral standard.
  4. The evaluation strategy combined human assessment, automated toxicity scoring, and co-occurrence metrics to compensate for the known demographic biases within any single method.

Keep reading: See related articles below for more coverage on this topic.

Get smarter about AI

The sharpest AI news, curated daily. Delivered free to your inbox.

Learn more

Glossary terms

← Back to all articles