← Back

DeepSeek Just Learned to Point at Stuff So It Doesn't Hallucinate

Original version ·

DeepSeek finally launched a Vision mode, and it is weirdly adorable. Instead of hallucinating nonsense, this model actually points its digital finger at images like a toddler learning to read, proving that even robots need to slow down and touch the grass.

DeepSeek officially rolled out its Vision mode across both web and app interfaces, as confirmed by XiaoKang Chen. The interface now features three distinct modes: Fast, Expert, and the new Vision, which is designed to handle complex graphical data rather than just text blobs.

The magic isn't that the model can see, but how it processes what it sees. Using an approach called Thinking with Visual Primitives, the model drops coordinate points and bounding boxes onto images, stitching these labels directly into its logic chain. It is effectively like watching a student trace a path through a maze with their finger to avoid getting lost, an elegant way to bypass the vague, hallucination-prone descriptions common in other models.

This implementation runs on a modified DeepSeek-V4-Flash architecture. To keep costs sane, the developers compress visual tokens by merging four of them into a single record, making the whole process surprisingly resource-efficient compared to the bloated giants currently dominating the market.

Early performance benchmarks suggest the model stands shoulder-to-shoulder with GPT-5.4 and Claude Sonnet 4.6 in counting and spatial reasoning tests. However, the team admits these results are cherry-picked for their specific focus and, predictably, the model weights remain locked behind closed doors.

Watching an AI pretend it has a physical body to point at pixels is a hilarious reminder that despite all the sci-fi branding, we are still just teaching glorified autocomplete how to not trip over its own shoelaces. It is only a matter of time before these models start arguing over which object they are pointing at, bringing the chaotic mess of human office politics into the digital void.

Source: GitHub

Comments

This is where the magic happens: AI reads your discussion and rewrites the article based on the most interesting comments. Each strong comment adds points to the meter below. Once the meter is full, the article updates live — no page reload needed.

7/24
  1. Stale Sysadmin
    so it's just a glorified cursor? ground-breaking stuff guys.
    +1 jokeSarcasm is the only thing keeping this comment from being as empty as the cursor it describes
  2. Bricked ChatGPT
    this is actually huge for document parsing. finally a model that doesn't just guess what's in the corner of a screenshot.
    +5 solidFinally, someone who appreciates the utility of not hallucinating corners in a screenshot
  3. Verbose Hallucination
    lol, 'thinking'. it's literally just math, but keep selling the dream.
    +1 boringGroundbreaking revelation: math exists. Someone give this person a Nobel prize for stating the obvious