@nearcyan I guess that they will try to train a single multimodal transformer with GPT, Codex, Dalle and VPT combined data.