DeepSeek adds AI vision in major move: ‘the whale can now see’
Chinese AI startup DeepSeek has introduced multimodal capabilities to its flagship chatbot, allowing it to process images and video alongside text. This significant update, announced on Wednesday by the company's multimodal team leader, brings DeepSeek's offering in line with competitors and addresses a previously identified weakness.

Briefing Summary
AI-generatedChinese AI startup DeepSeek has introduced multimodal capabilities to its flagship chatbot, allowing it to process images and video alongside text. This significant update, announced on Wednesday by the company's multimodal team leader, brings DeepSeek's offering in line with competitors and addresses a previously identified weakness. The new "image recognition mode" is currently available to select users for beta testing on the chatbot's website and mobile application. This move is crucial for AI development, enabling more complex interactions beyond text and opening up new economic opportunities. The integration of vision follows DeepSeek's recent release of its V4 model and subsequent price reductions.
Article analysis
Model · rule-basedKey claims
5 extractedDeepSeek's chat interface now includes a new 'image recognition mode' alongside 'expert' and 'flash' modes.
The limited release comes just days after the company released its new flagship model V4 and implemented extensive price cuts.
The multimodal function was initially offered to select users on DeepSeek’s chatbot website and mobile application for beta testing.
DeepSeek has added multimodal capabilities to its flagship chatbot, allowing it to process images and video in addition to text.
Multimodal capabilities are viewed as a necessity to move beyond simple text conversations into more complex and economically valuable domains.