OpenAI unveils DALL·E and CLIP
OpenAI announced two models on the same day: DALL·E, which generates images from text, and CLIP, which learns visual concepts from natural language supervision. DALL·E, a 12-billion-parameter version of GPT-3 trained on text-image pairs, could render prompts like "an armchair in the shape of an avocado" as images. CLIP performs zero-shot image classification given only the names of the target categories. Both works carried language-model techniques into vision, showing that text and images could be handled within a single framework.