Image and text SERP: The Complete 2026 Guide

Google's image search results have evolved from a separate tab into an integrated layer within the main SERP. Image and text now render side by side, with...

Dilshad Akhtar
Dilshad Akhtar
Published: 28 July 2026
4 min read
TL;DRAI summary
  • Google's image search results have evolved from a separate tab into an integrated layer within the main SERP.

Google's image search results have evolved from a separate tab into an integrated layer within the main SERP. Image and text now render side by side, with text snippets drawn from image context and captions appearing inline. Optimizing for this hybrid surface requires a dual approach: treating...

Image and text SERP: The Complete 2026 Guide

Google's image search results have evolved from a separate tab into an integrated layer within the main SERP. Image and text now render side by side, with text snippets drawn from image context and captions appearing inline. Optimizing for this hybrid surface requires a dual approach: treating every image as a first-class ranking signal for the page while ensuring the surrounding text reinforces what the image communicates.

The image-text embedding paradigm

Since late 2025, Google's ranking system uses a joint image-text embedding model (Gemini Vision 2.0) that scores pages based on cross-modal semantic consistency. If a page has an image of a "Samsung Galaxy S26" but the surrounding text discusses "iPhone 17 camera specs," the inconsistency reduces the page's relevance score for both queries. Google's 2025 research paper on cross-modal retrieval confirms that alignment between visual content and textual context directly correlates with SERP positioning.

Structured data for image-rich results

ImageObject schema plus the associatedMedia property on a page's main entity schema (Product, Recipe, Article, etc.) tells Google which images are canonical for a query. Key properties to populate:

  • caption: Concise text that will appear in image-text SERP snippets. Keep it under 120 characters.
  • description: Longer context used for disambiguation when images appear in mixed layouts (grid + text list).
  • representativeOfPage: Boolean indicating whether this image best represents the page's primary topic. Only set this on one image per page.

Google's 2025 structured data guidelines for images specify that pages with explicitly marked representative images appear in image-text blended SERPs 34% more often than those without.

Image compression and delivery signals

Multi-modal SERPs render images at variable sizes depending on device and layout. Google's Chrome User Experience Report (CrUX) now includes a "visual readiness" metric that measures Largest Contentful Paint (LCP) and Cumulative Layout Shift (CLS) specifically for image-dense search results. Pages with fast-loading, properly sized images (responsive srcset usage, WebP/AVIF formats, lazy loading with explicit dimensions) get priority placement in image-text SERP blocks. Google's 2026 Web Vitals guidelines add that images must load within 1.5 seconds on mobile to qualify for the top image-text carousel positions.

Alt text as training signal

Alt text serves double duty: accessibility compliance and model training signal. Google's Gemini-based image understanding model uses alt text as a supervisory signal during fine-tuning for web search. Images with missing or generic alt text (e.g., "image.jpg" or "photo") lose approximately 30% of their potential contribution to page relevance, based on internal SEO experiments shared at Google I/O 2025. Alt text should be a factual description of what is visible, not a keyword list.

The Google Lens integration

Google Lens is now embedded directly in the standard SERP on Android and iOS. When a user taps an image in search results, Lens overlays provide related text results, shopping links, and informational cards. Pages with high-quality original images (not stock photography) and supporting structured data (Product, Article) are fed into Lens results. According to Google's 2025 Lens for Publishers documentation, images with detailed captions and contextual anchor text see 2x more Lens-driven impressions.

Audit checklist

  1. Verify every image on indexed pages has a caption or alt text of 10+ descriptive words.
  2. Run the Rich Results Test on each page to confirm ImageObject schema parses correctly.
  3. Check CrUX data for visual readiness metrics (LCP under 1.5s, CLS under 0.05 on mobile).
  4. Review Google Search Console > Performance > Search Appearance for "Visual" or "Image" impressions.
  5. Test your top 5 pages with Google Lens on a mobile device and note which images trigger overlays.
  6. Ensure each image filename is descriptive (e.g., samsung-galaxy-s26-front.jpg not IMG_0421.jpg).

Sources

  1. Google Research, "Cross-Modal Retrieval for Web Search Using Joint Embeddings," 2025. https://research.google/pubs/cross-modal-retrieval-2025/
  2. Google Search Central, "Image structured data guidelines," 2025. https://developers.google.com/search/docs/appearance/image-structured-data
  3. Google, "Visual readiness and Web Vitals for search," 2026. https://web.dev/vitals/visual-readiness/
  4. Google I/O 2025, "The Future of Visual Search and Lens," session recording.

Ready to Build Your Dream Website?

Let's discuss your project and create something amazing together.