Models & Architecture
GPT (Generative Pre-trained Transformer)
A family of large language models developed by OpenAI based on the transformer architecture, known for their ability to generate coherent and contextually relevant text.
GPT (Generative Pre-trained Transformer) is a series of large language models created by OpenAI. The name reflects three key aspects: they are generative (produce new text), pre-trained (trained on massive unlabeled data before fine-tuning), and use the transformer architecture.
Major GPT model releases:
- GPT-1 (2018): 117M parameters, proved the effectiveness of pre-training
- GPT-2 (2019): 1.5B parameters, notable for text generation quality
- GPT-3 (2020): 175B parameters, demonstrated few-shot learning capabilities
- GPT-4 (2023): Multimodal capabilities, improved reasoning
- GPT-4o (2024): Omni model handling text, vision, and audio
GPT models power ChatGPT, Microsoft Copilot, and numerous third-party applications through the OpenAI API.
Related Terms
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.