Most visual AI demos have a satisfying shape: show an image, receive an answer, move on. “This is a ceramic bowl.” “That is a monarch butterfly.” “The sign says…” The format is tidy because the world is not.

In real use, an image often contains just enough evidence to narrow the field and not enough to finish the job. A small object may have three plausible names. A plant may be easy to place in a family but impossible to identify from one blurry leaf. A label may be readable while the thing it describes is not.

That is where visual AI becomes more interesting. The system is no longer only answering What is this? It is helping decide What should we find out next?

Identification is a doorway, not the room

Identification is useful. It gives a person a handle for search, memory, conversation, and action. But a label can also end curiosity too early. Once an object has been named, the interaction often collapses into a dead end: result received, tab closed.

Exploration has a different shape. It treats the first answer as a working hypothesis and asks what would make the next answer more useful. That might mean checking the underside, photographing a maker’s mark, placing a coin beside the object for scale, moving into better light, or asking where it was found.

When the picture is insufficient, the best response is not a longer guess. It is a smaller, more useful question.

This is not just a conversational flourish. It is a way to connect the uncertainty of perception with the agency of the person taking the picture.

A more useful loop

ShowCapture the thing in front of you.
NoticeState what the image supports.
Ask nextRequest the clue with the highest value.
ExploreTake another look and update.
This loop separates a useful observation from a final claim and gives the user a concrete next move.

Uncertainty should create momentum

People usually tolerate uncertainty better when it is specific. “I’m not sure” is honest but not very actionable. “The shape and texture are consistent with two common kinds of seed pod; a photo of the stem and the surrounding plant would separate them” is both honest and useful.

There are three parts to that pattern:

  • Evidence: what is visibly present—shape, texture, color, text, context, or scale.
  • Boundaries: what the image cannot establish, especially when lookalikes remain.
  • Next clue: the smallest additional observation likely to change the conclusion.

The third part is the design opportunity. A follow-up question is valuable when it reduces ambiguity without asking the user to become a computer-vision researcher.

From visible detail to next question

What the image gives us

Rounded form, carved grooves, dark patina, and no visible maker’s mark. The object is shown without a size reference.

What to ask next

“Can you photograph the underside and place a coin beside it? Those two details would separate a small decorative object from a tool part.”

A good next question is concrete, low-effort, and tied to a visible source of ambiguity.

The overlooked details are often the useful ones

Visual systems are naturally drawn toward salient features: a bright color, a familiar silhouette, a face-like pattern. Human investigators often win by noticing less glamorous evidence.

For an unknown object, the underside may reveal construction. A seam can distinguish a molded part from a hand-made one. A serial number can turn a broad category into a narrow search. Scale can eliminate an entire branch of possibilities. Location and date can matter too: the same shell, insect, or thrift-store object means something different in different contexts.

That suggests a practical rule for multimodal UX: do not only ask the model to describe the image. Ask it to identify which missing detail would most change its mind.

In interface terms, that can become a small set of prompts—show the underside, add scale, move closer, read the marking—rather than an empty chat box that makes the user invent the workflow from scratch.

Conversational multimodal UX is a control system

Calling an interface “conversational” can make it sound like the main challenge is tone. Tone matters, but the deeper shift is control. A conversational visual assistant lets the user steer the observation over multiple turns: correct the category, add context, reject a guess, or decide that the original question was not the interesting one.

That means the system needs memory, but bounded memory. It should retain the recent clues that make the current exploration coherent without implying an unlimited personal archive. It needs personality, but personality should not disguise uncertainty. It needs initiative, but initiative should appear as an invitation the person can ignore—not as an action taken on their behalf.

Research on human-AI interaction keeps returning to related principles: make system behavior legible, preserve user control, and design for correction. Google’s People + AI Guidebook frames explainability as part of helping people understand and appropriately trust an AI system. Microsoft’s HAX Toolkit treats interaction patterns as a design problem, including how a system communicates limits and recovery. NIST’s AI Risk Management Framework is broader than product UX, but its emphasis on validity, reliability, transparency, and accountability is a useful reminder that “helpful” is not the same as “confident.”

What building WhatIsThat taught us

While building WhatIsThat: Your Camera Pal, the most durable product lesson was that the camera interaction should not end at the first result. The interesting loop is: show something, let the camera companion react, notice what it explains, ask a follow-up, and decide what to show next.

That lesson changes the product questions. Instead of asking only “Can the model name this object?” we ask:

  • Did the response distinguish observation from inference?
  • Did it give the person a reasonable way to improve the next image?
  • Could someone correct the system without restarting the whole interaction?
  • Did the personality make the moment more inviting without pretending to know more than the image supports?

These are design questions, not claims of perfect accuracy or universal recognition. They are also portable. You can use the same rubric when choosing a visual assistant, designing a camera feature, or deciding whether a generated answer deserves your trust.

A better answer may be a better question

Visual AI will keep getting better at recognizing what is already in the frame. The product opportunity is to make the frame itself better—through a question that helps a person notice, move, compare, or look again.

The next time an image produces a plausible answer, pause before treating it as the finish line. Ask what evidence supports it, what remains ambiguous, and which single detail would be most useful to see next. That is where a label becomes an investigation.

A transparent note from the build

This article grew out of building WhatIsThat: Your Camera Pal, a camera companion designed for exploring the things around you. The ideas here stand on their own, whether or not you use the app.

If you want to try that camera-led conversation on Android, you can learn more on the app’s Google Play listing.

Learn about WhatIsThat on Google Play