"From recognition to creation: AI that reasons and acts."
OpenAI releases DALL-E, showing that LLMs can understand the relationship between text and visual pixels, creating images from scratch.
The release of RLHF-tuned models made AI conversational and useful for the general public, sparking a global race for "Generative" dominance.
Models begin to process "Native Multimodality"—they no longer just "read" text; they "see" video and "hear" audio in a single neural stream.
Shift from Chatbots to **Agents**. AI is given access to web browsers and APIs to perform tasks autonomously, such as booking travel or writing and executing code.
The methodology of "Self-Attention": the model looks at every word in a sentence simultaneously to understand context, rather than reading left-to-right.
The approach of using human "rankers" to tell the AI which answers are helpful and safe, aligning the model's behavior with human values.
A methodology where the AI is prompted to "think step-by-step," significantly improving its ability to solve math and logic problems.
Instead of relying only on its training memory, the AI "looks up" real-time information from external databases before answering, reducing hallucinations.
The Level Up: In 2023, the **Nvidia H100 GPU** became the most valuable commodity in the world.
The current methodology is Autonomous Planning. Instead of just generating a response, the AI uses a Reasoning Loop: