Beginner-Friendly Topic
Take your time - it's perfectly normal to re-read this topic 2-3 times. Try the interactive code editor below to run code yourself. Use the Q&A section to check your understanding before moving on. You've got this! 🚀
Types of Generative AI
Generative AI is not just one thing. There are different types, each specialized in creating a specific kind of content. Understanding the types helps you choose the right tool for the right task.
The Main Types of Generative AI
- Text generation: LLMs like GPT-4, Gemini, and Claude. They write text, answer questions, translate languages, and write code
- Image generation: Models like DALL-E, Midjourney, and Stable Diffusion. They create images from text descriptions
- Audio generation: Tools like ElevenLabs, Suno, and Adobe Podcast. They generate voice, music, and sound effects
- Video generation: Models like Sora, Runway, and Kling. They create video from text or image prompts
- Code generation: GitHub Copilot, Amazon Q, and Cursor. They write and complete code as you type
- Multimodal AI: Models like GPT-4o and Gemini 1.5. They work with text, images, and audio together
Tip
Tip
You do not need to learn all types of Generative AI at once. Start with text generation because it is the most versatile and most widely used. Once you understand how text generation works, the concepts transfer naturally to image, audio, and video generation.
Deep Learning ⊂ Machine Learning ⊂ Artificial Intelligence
Practice Task
Note
Think of a task you do regularly at work or school. For each type of Generative AI (text, image, audio, video, code), write one sentence explaining how that type could potentially help with your task. This exercise helps you understand which type of Generative AI is most relevant to your own work.
Key Takeaways
- Generative AI is not just one thing.
- Text generation: LLMs like GPT-4, Gemini, and Claude. They write text, answer questions, translate languages, and write code
- Image generation: Models like DALL-E, Midjourney, and Stable Diffusion. They create images from text descriptions
- Audio generation: Tools like ElevenLabs, Suno, and Adobe Podcast. They generate voice, music, and sound effects