DeepSeek's top-ranked V4 Flash stumbles on real agent tasks as its prices surge
DeepSeek's V4 Flash has topped model leaderboards and been hailed by developers as a "total monster" since its rollout. But in real-world testing, it completed just 53.8% of a batch of complex agent tasks. Composio ran the model through eight different agent harnesses, including Claude Code, Codex, and OpenCode, on 30 deliberately difficult, multi-step tasks spanning live tools like Gmail, GitHub, Slack, and Google …
AI brief
Pulse reads the full article- What happened
- Why it matters
- What to watch
- Who's exposed
- GOOGLAlphabet Down
Companies named in the story.
Sign in to get the AI brief. Pulse explains what happened, why it matters, what to watch and who's exposed. Free for members.
Sign in to read the brief

