Multimodal Search in Search Console: What It Is and How to Read It
Table of contents
Multimodal search in Search Console has its own report since September 24, 2026: Google split the web view into text-based and multimodal. For the first time you can see which pages image searches reach, instead of only the pages typed text reaches.
This article covers what the report counts, where the filter lives, what it will not tell you (query data does not exist for this search type) and how to turn the numbers into decisions about your images and landing pages. Honest warning: if your site gets no visual traffic, you will see the filter but the table will be empty. That is not a bug you can fix, it is a documented limitation.
What multimodal search is (and why it finally got its own report)
The official definition is short: web search results where an image was used as part of the search. In other words, searches that do not begin with text typed into the search bar, but with an image: a photo, a screenshot or an object framed with a phone camera. The Search Console help centre phrases it as Web: multimodal, web search results where an image was used as part of the search.
What multimodal search is not
- It is not voice search: it does not feed this report.
- It is not the Images tab: the image search type still exists separately and measures that tab. Multimodal measures searches started with an image that end up in web results.
- It is not image only: it is a mix of an image and, often, some context. What changes is where the search starts, not the kind of result.
Why Google splits it from the text-based web view
The key detail is in the announcement: this traffic was not previously counted in the regular web view. It is new information, not a redistribution of what you already saw. The practical consequence: do not read the arrival of multimodal impressions as a drop in text traffic, and do not read the reverse either. They are two separate buckets by design.
To put the scale in context, stick to official, dated figures: in February 2025 the Google blog said Lens handled more than 20 billion visual searches a month, and in May 2026 the company confirmed AI Mode passed one billion users in its first year.
What multimodal search in Search Console includes
The Search Central announcement explains that the data lands in two reports: the Search results performance report and the generative AI performance report (the one measuring impressions for AI-created views and AI Mode). The rollout is global and progressive from September 24, 2026.
The four surfaces Google counts
- Google Lens, from the camera and from the screen.
- Circle to Search on Android, when a user circles something with a finger.
- Images uploaded to Google Search, meaning a photo or file submitted directly.
- Search this image in Chrome, the right-click option on a web image.
Where the filter is and how to reach it in 4 steps
- Open Search Console and go to Performance, for Search results.
- At the top of the chart, open the Search type filter.
- Inside Web there are now two options: text-based and multimodal.
- Select multimodal and click Export if you want to analyse it elsewhere.
The official help is explicit about the Export button: today that is the route to take the data with you.
The limits of the data: no queries, no API and only with visual traffic
There is no query data, and that is not your fault
The dimensions page puts it literally: because multimodal searches mostly use images rather than text, specific text query data is not available for this traffic. As a result, the queries dimension is not available when this search type is selected.
In practice: picking multimodal removes the queries view and leaves you with pages, plus breakdowns by country, device and date. If your dashboard showed a keyword table, it goes empty. Change the question: instead of what people searched for, ask which of your pages help someone looking at something similar to what you published.
The Search Analytics API does not expose it yet
The searchanalytics.query reference documents the type parameter with the values discover, googleNews, news, image, video and web (the default). There is no multimodal value. If you built your SEO dashboard with Data Studio and Search Console, it is not there either: today the only route is the Export button.
When data starts showing, and what you will not see
- Requirement: real traffic from one of the four surfaces. Without it, filter present and table empty.
- No Search Labs experiment traffic appears here.
- 16-month history: what you do not export is gone.
- Consolidation: the latest days can shift; do not conclude from a single day.
How to read the data without drawing false conclusions
Multimodal versus text-based, page by page
The most profitable use is the cross-check: take the pages with impressions and clicks under multimodal and compare them with the same URLs under the text-based filter. That comparison reveals whether a page works visually even with little optimised text: a product page, a guide with original photos or a tutorial with screenshots perform better than their text traffic suggests. You can conclude that the URL interests people searching with an image; you cannot conclude which query it came from, or attach a keyword to it.
The breakdowns that do exist: country, device and date
Multimodal is mobile by nature: Circle to Search and Lens live on the phone. If the device breakdown comes out almost entirely mobile, that is not a measurement error. Use it to prioritise where you invest in product photos or explanatory screenshots, and compare long periods (28 days and earlier): with little data, daily swings are noise.
What to do with this data: priorities for visual content
Optimize the landing page, not just the image
Google repeats it in its image SEO best practices: beyond getting the image discovered and indexed, you have to optimise the landing page with surrounding content, title, context and useful text. Rule of thumb: if the image attracts the click, the page converts.
Make your images crawlable: quick checklist
- Real
<img>elements withsrc: Google does not index CSS background images. - Descriptive alt text: it describes the image to search engines and screen readers.
- Image sitemap: it discovers images Google would not find. It can include URLs from another domain, such as a CDN, and you should verify ownership of that domain.
- Supported formats: BMP, GIF, JPEG, PNG, WebP, SVG and AVIF, with an extension matching the real format.
- Responsive with a fallback: if you use
srcsetor<picture>, always keep an<img src="...">fallback. - Quality and speed together: review image lazy loading and LCP, and add structured data where it applies.
Measure, change, measure again
Export the report and mark the five pages with the most multimodal impressions; improve their alt text, context and speed; compare again in 28 days. If your site has not a single multimodal impression, the conclusion is not that this is useless: it is that there is no visual surface to capture because your images are decorative or set as backgrounds.
Five mistakes when reading the multimodal report
- Looking for the keyword. That dimension does not exist; any multimodal search keyword you see in a third-party report is invented.
- Adding multimodal to the web view and announcing growth. They are separate views by design.
- Confusing it with the Images tab. image measures that tab; multimodal measures searches started with an image.
- Concluding from four impressions. Accumulate weeks first.
- Forgetting to export. The window is 16 months.
Conclusion: you can finally measure visual search
Three ideas: the data exists and it is new; it is limited on purpose (no queries and no API, it measures which pages, not which searches); and it is acted on in the landing page, not only in the image file.
Your next step: open Search Console, filter by Search type, select multimodal and export. Then cross those pages with the image optimisation for SEO you already apply. And if you see odd movements in overall traffic, do not mix them with this report: diagnose them separately with the guide on a search traffic drop in Google.
The report also matters for sites that want to show up in AI assistants: better contextualised visual information is a better base for GEO and citations in ChatGPT, Perplexity and Gemini, and for a site ready for AI agents.