Zhipu has introduced the GLM-5V-Turbo, a multimodal coding model designed for visual programming. According to Odaily, this model natively understands various multimodal inputs, including images, videos, design drafts, and document layouts. It also supports the use of multimodal tools such as framing, screenshotting, and web page reading, with an expanded context window of up to 200,000.