Your privacy choices

Optional Google Analytics cookies help us understand which guides are useful. They stay off unless you accept. You can change your choice in the footer. Privacy policy

Google’s EmbeddingGemma 2 Brings Offline AI Search to Photos, Video and Audio

Google’s new embedding model supports local search across text, photos, video and audio. Here’s what its memory figures mean—and what developers should test.

By Ubedulla · 6 min read
Editorial illustration of a smartphone with floating photo cards, an audio waveform and video frames connected by light.
AI-generated editorial illustration of on-device media search.

You remember what was in the video. You cannot remember its filename. Somewhere among hundreds of clips is the moment someone sketched the new kitchen layout, and searching for “kitchen” returns nothing.

That is the kind of problem Google’s EmbeddingGemma 2 is designed to help developers solve. Announced on October 6, the 740-million-parameter model turns text, images, video and audio into numerical representations that software can compare. Google released it under Apache 2.0 for applications that can run on consumer devices. Google’s launch announcement describes a move beyond the original model’s text-only focus.

The appeal is straightforward: find something by describing it, with the search processing happening on your device. The harder question is how well that works on your actual files.

What does EmbeddingGemma 2 actually do?

An embedding is a list of numbers representing an input. A search application compares those lists to rank potentially relevant items. If you are new to the terminology, our AI glossary explains the surrounding concepts.

Think about searching a folder of renovation photos. A conventional filename search needs matching words. A semantic search tool could instead look for images related to “the room before the cabinets were installed.” That is an illustrative use case, not a result we have tested.

EmbeddingGemma 2 produces vectors rather than written answers. A developer can use it to retrieve material for another model to explain, or simply show the matching files. The distinction matters: a relevant-looking search result still needs inspection.

The small memory figures need context

Google reports roughly 191MB of active RAM for text-only weights and 567MB for the full multimodal model with quantization on a Pixel 11 Pro. Those are specific deployment figures, not a promise that a complete search app will occupy 567MB on every phone. The announcement specifies the device and configuration.

For anyone building an app, the useful budgeting exercise is broader. Count the model, the search index, the original files, temporary processing, and the interface. Also decide when indexing happens. A fast search after preparation does not tell you how long the first import takes.

Imagine handing the app 20,000 photos. Before judging the search box, ask what happens during that first evening: does the app show progress, pause politely, and recover after being closed? Those product details may matter more to a user than a small difference in benchmark scores.

A storage calculation worth doing before you build

Google’s model card lists a native output of 768 dimensions, with supported shorter representations of 512, 256 and 128. It says 128 dimensions is best suited to text-only workloads.

Here is our calculation for 100,000 vectors stored as 32-bit numbers. Each number occupies four bytes; these decimal megabyte totals exclude metadata, index structures and source files.

DimensionsBytes per vectorRaw storage for 100,000 vectors
7683,072307.2MB
5122,048204.8MB
2561,024102.4MB
12851251.2MB

For the first row, the arithmetic is 100,000 × 768 × 4. Moving to 128 dimensions divides this particular storage requirement by six. It does not divide the whole application’s size by six.

The decision should depend on what gets lost. If the smaller representation saves space but repeatedly misses the clip someone needs, the saving is a poor trade. Keep a set of known searches and compare the results before choosing the smallest setting.

The catch: the benchmark and the phone configuration are different

The model card reports a code-retrieval benchmark score of 78.68, compared with 68.76 for the predecessor—a gain of 9.92 points. It also states that its reported evaluation results use the full-precision checkpoint. Those evaluation details matter when comparing them with the launch’s quantized memory figures.

Neither number answers the question a small business might ask: “Will this find the right maintenance recording when the room is noisy and the filename is useless?” A score across benchmark tasks cannot settle that specific case.

Our suggested first evaluation would be deliberately modest: collect 30 real questions and mark which files should answer each one. Include awkward cases—similar-looking rooms, background noise, and requests for things that are absent. Record whether the right item appears among the first five results, along with indexing time and memory use.

That last category is easy to overlook. An app should have a sensible response when it cannot find something. Returning five plausible files for every query can make an unreliable system look busy.

Where you can try it

Google’s AI Edge release post introduces Instant Media Search and Video Moments Finder in its Gallery showcase app, plus local retrieval demonstrations in AI Edge Foresight for Mac. The same post says Android access through an ML Kit service is planned for the coming weeks; that part should not be described as already generally available.

These are useful starting points for exploration. We have not run the demonstrations or independently measured their accuracy, speed or battery use.

For a first trial, use a small folder of non-sensitive files whose contents you know well. Try literal searches, then descriptions that do not reuse the filenames. Keep track of failures as carefully as successes. A demonstration becomes much more informative when you know the answer beforehand.

Does local search automatically make an app private?

No. Local processing is one part of the design. Our practical recommendation is to check the full path: where imported files are stored, whether indexes are backed up, what analytics collect, and whether retrieved material is subsequently sent to an online assistant.

The release gives developers another way to build useful search without making remote processing a prerequisite. Whether a particular app delivers that benefit depends on its implementation. For readers, the test is refreshingly ordinary: can it find the file you remember, show enough context to confirm it, and explain where your data went?

Source note: Researched October 7, 2026, from Google’s launch announcement, its model card and its AI Edge developer post. Performance figures are Google-reported. Storage totals are The Bot Post’s calculations; scenarios and the suggested evaluation are our analysis. This is an AI-assisted, source-checked explainer, not a hands-on review.

About the author

Ubedulla

Founder & Editor

Founder and editor of The Bot Post, covering AI news and technology.

Related Articles