EmbeddingGemma 2 and Qwen3-Embedding vs bge-m3: Why I Kept bge-m3
EmbeddingGemma 2 and Qwen3-Embedding 0.6B against bge-m3 on one workload, clustering news articles from Vietnamese newspapers into stories. Quality was level, each model has its own cosine scale, and the one real gap covers under one percent of pairs.
Google released EmbeddingGemma 2 on 2026-10-06. I measured it, and Qwen3-Embedding 0.6B, against bge-m3, the 2024 model a pipeline I run uses in production to cluster news articles from Vietnamese newspapers into stories. On quality the three came out level, so bge-m3 stays.
This is not a scientific benchmark. It is a comparison made for one practical need, on one workload, to answer "should I swap the model I run?". It does not rank these models in general, and none of it transfers to English, to retrieval, or to someone else's pipeline.
The workload: clustering articles into stories
The data is news: articles published by Vietnamese online newspapers, in Vietnamese. Many papers report the same event, and one event is reported many times as it unfolds. The pipeline's job is to turn a day of articles into stories, where a story is every article about one event.
It does that with clustering. A model turns each article (title and opening paragraphs, about 1,400 characters) into a vector, and the similarity of two articles is the cosine between their vectors. Starting from the closest pair, two groups are merged into one story when their cosine clears a cut-off: 0.80 always, or 0.70 when both also name the same rare person or place. So the embedding model decides everything: score two reports of one event below the cut-off and the story splits in two, score two different events above it and they are glued into one.
The same vectors also feed a category classifier and match a bare link title to its article, so those were measured too. All three models ran as 8-bit GGUF files on one llama.cpp build, on one 16 GB Apple-silicon desktop, over the same 19,192 articles.
One day of articles as bge-m3 places them, flattened to three dimensions.
How to read it. Each ball is one article, and balls close together are articles the model finds similar: the coloured clump is one story reported in three stages, the black clump a different story with similar wording. It is an illustration, not evidence, because squeezing a thousand dimensions into three bends distances.
Results
| Measure | bge-m3 | Qwen3 0.6B | EmbeddingGemma 2 |
|---|---|---|---|
| Same story or not, AUC on 120 labelled pairs | 0.953 | 0.957 | 0.887 to 0.960 by prefix |
| Category classifier, held-out accuracy | 0.885 | 0.885 | 0.881 to 0.891 |
| Link title finds its article | 94.2% | 94.0% | 91.6% to 92.3% |
| Articles a second | 10.1 | 5.0 | 15.1 |
| Vector size | 1024 | 1024 | 768 |
| Resident memory | 1,227 MB | 1,167 MB | 1,011 MB |
The first row is the clustering question. I took 120 pairs of articles labelled in advance, 75 about the same story and 45 about different stories, and asked how well each model's cosine separates the two groups. AUC is the chance that a same-story pair scores above a different-story pair: 1.0 is perfect, 0.5 is a coin toss. Its 95% intervals run from about 0.92 to 0.98 for each model at its best, so the first row is a tie. So is the second: one standard error on 1,493 test articles is 0.8 points. EmbeddingGemma 2 swings with its prompt prefix, and the one named "clustering" was its worst.
The same 120 labelled pairs under each model.
How to read it. Every dot is one pair, placed by the cosine the model gave it: violet (up) is a same-story pair, yellow (down) a different-story pair. Look at how much violet and yellow share the same columns, since no cut-off can get those right, and it is about the same in all three panels.
Three models, three scales
A cosine of 0.80 means nothing by itself. Across 21 million pairs of articles published the same day, an ordinary pair scores 0.31 under bge-m3, 0.22 under Qwen3 and 0.61 under EmbeddingGemma 2. The cut-offs tuned for bge-m3, 0.70 and 0.80, land near 0.83 and 0.90 for EmbeddingGemma 2, which has about half the room between "unrelated" and "same story".
Every threshold in a pipeline belongs to the model it was tuned on.
How to read it. The hump is where unrelated pairs score, and the two black lines are the cut-offs for joining a story. Look at the distance between them: EmbeddingGemma 2 has to make the same decisions in half the space.
The one real difference
Averages hide the interesting case: pairs where two models disagree. One day held a long trial reported in stages: it opened, a sentence was asked for, a verdict came. For clustering, all of it is one story.
One story in three stages beside a different story with similar wording.
How to read it. Every small square is one pair of articles, darker means more similar, and the number in each block is its average: the first three rows and columns are the stages of the trial, the last is the other story. Look at stage 1 against stage 3: 0.70 for bge-m3, right on its looser cut-off, so 44% of those pairs fall under it and the story risks splitting, while Qwen3 scores them 0.82 and loses 10%.
On that day I drew 20 pairs only Qwen3 was sure of and 20 only bge-m3 was sure of, and had them judged blind. Qwen3's were one story 20 times of 20, bge-m3's 10 of 16. That looked decisive. On the next day, with no staged story, such pairs barely existed: ten of them, among 5,886 that bge-m3 joins.
When only one model would join a pair, how often is it right to?
How to read it. Each bar is the share of pairs only that model would join that a blind judge called one story, with the count above it. The left pair of bars in each panel is the day of the trial and the right pair is the day after, where the bars stand on fewer than ten pairs each.
So I drew samples built to favour neither model: four other days, pairs bucketed by how sure each model was, each article used once per bucket so no story could fill one. Weighted by bucket size, the share of pairs above the strict cut-off that really are one story was 0.998 for bge-m3, 0.982 for Qwen3 and 0.976 for EmbeddingGemma 2. The strong disagreements were under one percent of those pairs. Inside them Qwen3 was right a little more often than bge-m3 (10 of 14 against 8 of 13), and EmbeddingGemma 2 clearly less often (5 of 12 against 10 of 12).
Why the old one stays
The model file is the cheap part. A swap means embedding every stored article again, retraining the classifiers, and measuring every cosine threshold again, because the scales do not match. Qwen3 0.6B, the only candidate with a lead, would also make every embedding run twice as long. That is a lot to pay for a gain in under one percent of pairs.
Limits
- 120 labelled pairs and two neutral samples of about 130 pairs each are small.
- Labels came from language models that saw the two articles and nothing else, not from people. A second blind pass agreed with the first on 110 of the 111 pairs both judged firmly.
- The cut-offs were measured, not the clusters: I did not run the full clustering under each model and compare the stories it produced.
- One language, one kind of text (news), one machine, one llama.cpp build, one quantisation.
TL;DR
- Question: should a pipeline that clusters Vietnamese news articles into stories swap
bge-m3for EmbeddingGemma 2 or Qwen3-Embedding 0.6B? No. - Quality is a tie. All three separate same-story pairs from different-story pairs about equally well (AUC about 0.95), and a classifier trained on their vectors scores the same.
- Scales differ. An unrelated pair scores 0.31, 0.22 or 0.61 depending on the model, so no threshold survives a swap, and EmbeddingGemma 2 leaves half the room to work in.
- One real gap, and it is small. Qwen3 0.6B keeps the stages of one long story together better than
bge-m3. It shows on a day with such a story and covers under one percent of pairs otherwise. - The swap is the expensive part. Every stored vector, classifier and threshold has to be redone, and Qwen3 0.6B runs at half the speed.
- If you run your own comparison: measure the job you use the model for, and test a striking result on a second day before believing it.
