Building with AI

Cloudflare AI Search goes GA with native image embeddings and OCR

Cloudflare’s AI Search is now generally available, adding native image embeddings, OCR for PDFs, and support for larger files. Billing starts November 1, 2026.

Cloudflare has moved its AI Search product from preview to general availability, marking a significant step for developers building retrieval-augmented generation pipelines. The update introduces native multimodal capabilities, including direct image embeddings and optical character recognition for scanned documents. Billing for the service will commence on November 1, 2026, though a free tier remains available for smaller workloads.

What happened

The launch expands the platform’s ability to handle complex data types beyond simple text. Previously, image search relied on generating text captions from visual content and then embedding those captions. This approach often lost fine-grained visual details such as texture, color palettes, or spatial relationships within charts and diagrams. The new system preserves both textual understanding via captions and direct visual signals through native image embeddings.

Alongside multimodal support, Cloudflare increased the maximum file size for ingestion from 4 MiB to 10 MiB. This change accommodates larger documentation sets and richer media assets. The company also enabled optical character recognition for PDFs that consist of scanned images rather than selectable text. When enabled, the system extracts text from each page before chunking and embedding, ensuring that legacy scanned documents become searchable.

The transition to general availability also brings finalized pricing structures. During the August 2026 Agents Week, Cloudflare announced preview pricing, but the final model includes a slight adjustment to the free tier. Instead of a shared pool of queries, users now receive separate allotments for semantic and full-text searches. The billing engine calculates costs based on ingestion tokens, storage volume, and query types, removing the need for upfront capacity planning.

How it works

The core improvement in image handling relies on Matryoshka Representation Learning. This technique allows the system to create embeddings that retain useful information at varying levels of granularity. By doing so, it keeps storage requirements manageable while maintaining search speed. The Qwen3-VL-Embedding model powers native multimodal retrieval, placing image pixels and text into the same vector space.

When a user submits a query, the system checks if the selected embedding model supports images. If it does, the query image is embedded directly. If the model is text-only, the system converts the image to text using a tool called ToMarkdown and searches using the resulting caption. This dual-path approach ensures compatibility across different model configurations while offering enhanced precision for those using multimodal models.

For scanned documents, the optical character recognition feature acts as a pre-processing step. It reads text from image-based PDFs before the standard indexing pipeline begins. This process consumes image processing ingestion tokens, which are billed separately from base text ingestion. The rest of the pipeline, including parsing, chunking, and reranking, remains included in the base ingestion cost.

Key details

  • Native image embeddings use the Qwen3-VL-Embedding model to preserve visual details like texture and layout.
  • Maximum file size for ingestion has increased from 4 MiB to 10 MiB for text files and PDFs.
  • Optical character recognition is available for scanned PDFs and billed as image processing ingestion tokens.
  • Billing starts on November 1, 2026, with no instance hours or monthly minimums required.
  • The free tier includes 5M ingestion tokens, 10 GB of storage, 1,000 semantic queries, and 1,000 full-text queries per month.
  • Base ingestion costs $0.75 per 1 million tokens, with an additional $0.50 per 1 million tokens for image processing.

Why it matters

For engineering teams building search interfaces, the shift from caption-based to native image retrieval reduces the friction of curating metadata. Developers no longer need to write exhaustive descriptions to capture visual nuances, which is particularly valuable for e-commerce catalogs, technical diagrams, or screenshot libraries. The ability to search by visual similarity or combined text-and-image queries opens new use cases that were previously difficult to implement without custom computer vision pipelines.

The pricing model’s transparency helps teams forecast costs more accurately. By charging only for ingestion, storage, and queries, Cloudflare removes the operational overhead of managing server instances or scaling capacity units. This pay-as-you-go structure lowers the barrier to entry for startups and side projects, while the generous free tier allows for substantial experimentation before committing to paid usage.

Support for larger files and scanned documents addresses common pain points in enterprise knowledge bases. Many organizations rely on legacy PDFs that are essentially images of text. Enabling OCR out of the box means these documents can be integrated into modern AI search workflows without external preprocessing tools. This simplifies the architecture required to build comprehensive internal search engines.

What you can do

  • Audit your current document repository to identify scanned PDFs that would benefit from the new OCR capability.
  • Test the Qwen3-VL-Embedding model with a subset of your image assets to evaluate the precision of native visual retrieval.
  • Calculate your expected monthly costs using the published rates for ingestion, storage, and queries to budget for the November billing start.
  • Update your ingestion pipelines to handle files up to 10 MiB, taking advantage of the increased size limit for richer content.
  • Review the free tier limits to determine if your current development and testing workload fits within the 5M token and 10 GB storage allotment.
  • Explore hybrid search queries that combine text descriptions with image inputs to improve relevance for visual-heavy datasets.

Tools from the Bytechap store

Keep reading

All stories