Run EmbeddingGemma 2 Locally

... PLUS: Vercel Ship 26 Is Happening October 15

In today’s newsletter:

  • Run EmbeddingGemma 2 Locally

  • Vercel Ship 26 Is Happening October 15

Reading time: 5 minutes.

If you're building agents and want to see how teams actually run them in production, Vercel Ship 26 is on October 15 at the Palace of Fine Arts in San Francisco.

Vercel’s Andrew Barba will lead a hands-on workshop on eve, Vercel’s framework for durable backend agents. He’ll build an agent from an empty directory, then walk through the pieces that let it survive a restart: file-based authoring, the channel and workflow runtime, and persisted state.

Other sessions walk through real production setups:

  • Notion on the agent platform it built on Vercel for both developers and agents

  • SpaceXAI on how Grok's .grok.me surface serves more than 300,000 subdomains and 125 million edge requests a month

  • AWS on agent patterns shipping across engineering teams, including running Vercel in your own AWS account with full VPC control

Speakers include Guillermo Rauch, Tobias Lütke from Shopify, Katelyn Lesse from Anthropic, and Diogo Almeida from TypeSafe AI.

Run EmbeddingGemma 2 Locally

Google recently released EmbeddingGemma 2, a 740M-parameter open embedding model for text, code, images, video, and audio.

Most retrieval systems start by choosing one kind of data. Text search indexes documents. Image search needs a vision model. Audio search often starts by turning speech into text. If your product contains all three, you usually end up maintaining several models and several indexes.

EmbeddingGemma 2 maps text, code, images, video, and audio into the same 768-dimensional vector space.

One Vector Space for Every Input

EmbeddingGemma 2 uses separate encoders for text, vision, and audio. Each encoder turns its input into a compatible vector, which means your application can compare a text query against a code file, image, audio recording, or video without converting everything into text first.

A support team could search call recordings with a typed question. A developer could search code and screenshots with the same query. A local RAG pipeline can index PDFs, diagrams, and audio without adding an embedding model for each format.

Load Only What Your Application Uses

The full model has 740M parameters, but the vision and audio encoders are optional. You do not need to load them for a text-and-code search tool.

Workload

What loads

Parameters

Text and code

Text model

270M

Text, code, images, and video

Text + vision

440M

Text, code, and audio

Text + audio

570M

All input types

Text + vision + audio

740M

Google reports that the quantized text-only model used about 191MB of active RAM on a Pixel 11 Pro. The full multimodal model used about 567MB.

Smaller Vectors Can Shrink Your Index

Every input produces a 768-dimensional vector by default. That is useful when quality matters most, but embedding storage adds up quickly once you index millions of files.

EmbeddingGemma 2 supports 512, 256, and 128 dimensions through Matryoshka Representation Learning. A 256-dimensional vector uses one-third of the storage of a 768-dimensional vector, and Google reports that it keeps most of the full model’s quality across text, code, vision, video, and audio benchmarks.

Reducing the vector size changes two parts of the setup:

  • Query and document dimensions must match. A 256-dimensional query cannot be compared with a 768-dimensional document vector.

  • Truncated vectors need L2-normalization. Cutting dimensions changes vector length, which otherwise distorts cosine-similarity scores.

For a new index, start at 768 dimensions and measure retrieval quality on your own queries. Lower the dimension only after you know what accuracy you are giving up for cheaper storage.

Run It With Unsloth

Unsloth supports EmbeddingGemma 2 in two ways. Unsloth Desktop is for indexing files and searching them locally through the app.

The GGUF build is for developers who want a local embedding endpoint inside their own application.

To use Unsloth Desktop:

  1. Download and open Unsloth Desktop.

  2. Go to Settings → General → Documents & RAG. Select EmbeddingGemma 2 from the embedding-model dropdown, then download it.

  3. Start a chat, enable Chat with Files, and add the documents you want to search.

Unsloth creates vectors when each file is uploaded. If you later switch to another embedding model, upload the files again so the new model can index them.

Run the GGUF with llama.cpp:

Use this route when you want EmbeddingGemma 2 to run as a local embedding service for your own retrieval pipeline. The Unsloth GGUF supports text and code embeddings.

First, build a recent version of llama.cpp and download Unsloth’s 176MB Q4 model:

Then start a local embedding server:

Your application can now request vectors from the local endpoint:

The server returns normalized 768-dimensional vectors. Store the document vectors in your vector database, then embed each incoming query through the same endpoint before retrieval.

The GGUF route covers text and code. Image, video, and audio embeddings need a runtime that loads EmbeddingGemma 2’s vision and audio encoders.

That’s all for today. Thank you for reading today’s edition. See you in the next issue with more AI Engineering insights.

PS: We curate this AI Engineering content for free, and your support means everything. If you find value in what you read, consider sharing it with a friend or two.

Your feedback is valuable: If there’s a topic you’re stuck on or curious about, reply to this email. We’re building this for you, and your feedback helps shape what we send.

WORK WITH US

Looking to promote your company, product, or service to 200K+ AI developers? Get in touch today by replying to this email.