NEWSAR
Multi-perspective news intelligence
SRCSouth China Morning Post
LANGEN
LEANCenter-Right
WORDS223
ENT7
WED · 2026-04-29 · 14:00 GMTBRIEF NSR-2026-0429-72435
News/DeepSeek adds AI vision in major move: ‘the whale can now se…
NSR-2026-0429-72435News Report·EN·Technology

DeepSeek adds AI vision in major move: ‘the whale can now see’

Chinese AI startup DeepSeek has introduced multimodal capabilities to its flagship chatbot, allowing it to process images and video alongside text. This significant update, announced on Wednesday by the company's multimodal team leader, brings DeepSeek's offering in line with competitors and addresses a previously identified weakness.

Vincent ChowSouth China Morning PostFiled 2026-04-29 · 14:00 GMTLean · Center-RightRead · 1 min
DeepSeek adds AI vision in major move: ‘the whale can now see’
South China Morning PostFIG 01
Reading time
1min
Word count
223words
Sources cited
2cited
Entities identified
7entities
Quality score
100%
§ 01

Briefing Summary

AI-generated
NEWSAR · AI

Chinese AI startup DeepSeek has introduced multimodal capabilities to its flagship chatbot, allowing it to process images and video alongside text. This significant update, announced on Wednesday by the company's multimodal team leader, brings DeepSeek's offering in line with competitors and addresses a previously identified weakness. The new "image recognition mode" is currently available to select users for beta testing on the chatbot's website and mobile application. This move is crucial for AI development, enabling more complex interactions beyond text and opening up new economic opportunities. The integration of vision follows DeepSeek's recent release of its V4 model and subsequent price reductions.

Confidence 0.90Sources 2Claims 5Entities 7
§ 02

Article analysis

Model · rule-based
Framing
Technology
Economic Impact
Tone
Measured
AI-assessed
CalmNeutralAlarmist
Factuality
0.85 / 1.00
Factual
LowHigh
Sources cited
2
Limited
FewMany
§ 03

Key claims

5 extracted
01

DeepSeek's chat interface now includes a new 'image recognition mode' alongside 'expert' and 'flash' modes.

factual
Confidence
1.00
02

The limited release comes just days after the company released its new flagship model V4 and implemented extensive price cuts.

factual
Confidence
1.00
03

The multimodal function was initially offered to select users on DeepSeek’s chatbot website and mobile application for beta testing.

factualChen Xiaokang
Confidence
1.00
04

DeepSeek has added multimodal capabilities to its flagship chatbot, allowing it to process images and video in addition to text.

factual
Confidence
1.00
05

Multimodal capabilities are viewed as a necessity to move beyond simple text conversations into more complex and economically valuable domains.

prediction
Confidence
0.80
§ 04

Full report

1 min read · 223 words
Chinese artificial intelligence start-up DeepSeek has added multimodal capabilities to its flagship chatbot for the first time – meaning that it can process images and video in addition to text – bringing it in line with rivals that already offer the function.The limited release to select users comes just days after the Hangzhou-based company released its new flagship model V4, which was followed by extensive price cuts.According to DeepSeek multimodal team leader Chen Xiaokang, who made the announcement on Wednesday on social media, the function was initially offered to select users on DeepSeek’s chatbot website and mobile application for beta testing.“Come try out the incredible work from our genius multimodal colleagues!” senior researcher Chen Deli wrote on social media shortly after, adding that “the little whale can now see”, a reference to DeepSeek’s whale logo.On DeepSeek’s chat interface, a new “image recognition mode” had been added alongside the “expert” and “flash” chat modes, which were introduced earlier this month.As AI continues to rapidly progress, multimodal capabilities are viewed as a necessity to move beyond simple text conversations with users into more complex and economically valuable domains.While DeepSeek’s breakout moment in January 2025 made it a household name internationally due to its model’s powerful reasoning capabilities and cost-efficiency, the start-up’s lack of a multimodal offering since then has been seen as an Achilles' heel.
§ 05

Entities

7 identified
§ 06

Keywords & salience

8 terms
multimodal capabilities
1.00
ai vision
1.00
deepseek
0.90
artificial intelligence
0.80
chatbot
0.70
image recognition
0.60
ai model
0.50
beta testing
0.40
§ 07

Topic connections

Interactive graph
Network visualization showing 21 related topics
View Full Graph
Person Organization Location Event|Click node to navigate|Edge numbers = shared articles