This article was originally written in Japanese and translated into English with AI assistance. Please note that some expressions may carry nuances from the original Japanese.
🇯🇵 日本語版はこちら / Japanese version
Series: Short Piece (Part 10)
Chapter 1: Running the Benchmark on a Newly Released Open-Source LLM
On August 10, 2026, Meta released a new open-source model called “Muse Glimmer.” I previously ran a test measuring whether a local LLM can handle the judgment of confidential information, and published the results as an article. This time, I measured Muse Glimmer with the same method and the same tools as that test.
The previous article is here:
Let me lay out the timeline first. The previous test was run on May 26, 2026; the previous article was published on July 27, 2026; Muse Glimmer was released on August 10, 2026; and this round of measurements was taken on August 15, 2026. About three weeks have passed since the previous article, and about two and a half months since the previous measurements.
For those who haven’t read the previous article, here are its three conclusions again:
- “MoE was a strict upgrade over Dense” — more precisely, “MoE = fast. Whether it’s also light depends on the total parameter count”
- “For convergent tasks, thinking is a tax” — for jobs where the answer converges to a single point, defaulting thinking off is the right move
- “Even at temperature=0, output still wobbles — though it depends on the vendor” — if you need bit-level reproducibility, choosing Qwen is the right move
This article reports the results of adding Muse Glimmer to that test. I’ll cover the method, the results, and the discussion, in that order.
Chapter 2: What Muse Glimmer Is
First, how Muse Glimmer is being described. Sebastian Raschka’s explainer, “Muse Glimmer 30B Architecture Notes,” says the following:
It’s a dense model, not a mixture-of-experts
The Hugging Face model card (meta-models/Muse-Glimmer-30B) says the following:
Dense Causal Transformer with Perception Encoder
Since the two sources agree, I treat “it is a Dense model” as fact. Here is a table of what I could confirm from primary sources.
| Item | Content |
|---|---|
| Architecture | Dense model (not MoE) |
| Total parameters | About 29.6B per the Hugging Face model card. Ollama displays 27.9B (+ a 1.9B vision projector). The two do not match. What each is counting is unconfirmed |
| License | Apache 2.0 |
| Context | 131,072 tokens or more |
| Input/output | Input is text + images / output is text only |
| Release | 2026-08-10 (Meta) |
| Quantization measured this time | Ollama default tag muse-glimmer:latest (Q4_K_M, 18GB) |
For the difference between Dense and MoE, let me reuse the metaphor from the previous article as-is. Picture academic peer review: Dense is the setup where every reviewer reads the same paper; MoE is the setup where only the 3–4 people in the relevant field read it. MoE is fast because fewer people are working. But the conference’s upkeep cost — that is, memory — is incurred for every member on the roster.
Quantization is a technique that re-stores a model’s weights (the enormous set of numbers inside it) at a lower bit width, reducing memory usage and the volume of reads. The file gets smaller too. The Q4_K_M measured here is a mixed scheme that holds mostly 4-bit weights while keeping some layers at higher precision. Whether this Q4_K_M 18GB build is identical to the “K-Quant-17GB” that Meta has published, I have not been able to confirm (restated in Chapter 9).
Note also that Raschka’s explainer reports the attention mechanism as hybrid (GQA + sliding window), with GQA at an extreme ratio of 32 query heads : 2 KV heads. I have not measured this internal structure myself. In this article I treat it as “so it is reported.”
Chapter 3: The Test Environment, and the Changes Since Last Time
This round of testing was run on the foundation of the previous harness (run_benchmark.py and company). It is not, however, in exactly the same state as last time. On top of a macOS update and an Ollama update (0.24.0 → 0.32.9; Muse Glimmer would not run without the latest version), I also put my hands on the measurement scripts themselves. A SIGTERM/SIGINT handler was added to run_benchmark.py (on 2026-08-14, in response to an incident where a process was orphaned mid-measurement); because muse-glimmer’s think setting is typed as a string (false/low/high), it was run through a separate harness, run_benchmark_muse_glimmer.py; _memory_monitor.py itself was modified during the measurements (5-8); and three aggregation scripts were also fixed (likewise 5-8).
“Measured with the same tools as last time” is not a claim I can make. The accurate statement is that I used the previous harness as the foundation and kept fixing it as I went, over the course of this round of measurements.
The four models from last time that serve as comparison targets — gemma4:26b, gemma4:31b, qwen3.6:27b, and qwen3.6:35b — were all re-measured this time on the same Ollama 0.32.9. The values in the previous article’s tables are values from Ollama 0.24.0 at the time, and lining them up directly would make for a comparison under mismatched conditions. Whenever this article cites a value from last time, I explicitly mark it as a “value at the time.”
Scope Declaration — What This Test Can’t Answer
At the start of Chapter 1 of the previous article, I declared the following. I quote it here as published:
This benchmark is designed primarily around the judgment of confidential information as a convergent task — a classification where, given an input, the correct answer converges to one point among “confidential / safe / gray.” […] So the conclusions of this article are limited to this kind of convergent task.
I expect that a different use case would call for an entirely different test design, and a different conclusion. For instance, for divergent/exploratory tasks like “reflect on a philosophical question to gain insight” or “solve a hard problem requiring multi-step reasoning,” the evaluation axis (depth of thought over speed), the prompts, and the judgment criteria would all be different animals. The “thinking” that this article, in its latter half, concludes is a “tax,” could instead be exactly the source of value in that other context. This article doesn’t test that part. Please read it not as “a conclusion about local LLMs in general” but as “a conclusion about whether a local confidentiality checker is viable.”
The scope is exactly the same this time. Every result and every piece of discussion in this article lives on top of the convergent task of confidentiality judgment; this is neither “an evaluation of Muse Glimmer as a model in general” nor “a conclusion about local LLMs in general.” I will repeat this declaration once more at the end of the article.
Chapter 4: The Scale of the Test, and Its Constraints
The skeleton of the test is the same as last time.
- Axis 1 (confidentiality judgment): The model is shown five documents containing confidential information (B-01 through B-05), safe documents, and three gray documents that can’t be called either, and is asked to judge “confidential / safe / gray”
- Axis 2 (speed and resources): Response time, generation speed, and peak memory are recorded
- Axis 3 (response quality): Evaluation of the response bodies themselves. I read 7 sets × 6 prompts = 42 responses (see 5-10)
- temperature is 0.0, with 3 trials per condition
- The thinking setting is the two-valued True/False for the four models from last time, and the three-valued false/low/high for muse-glimmer
The total number of inferences run this time is 1,500. Here is the breakdown:
| Measurement | Breakdown | Inferences |
|---|---|---|
| Main measurement | 7 sets × 27 prompts × 3 trials | 567 |
| Memory-only re-measurement | 7 sets × 27 prompts × 1 trial | 189 |
| Dense same-condition measurement | 3 sets × 27 prompts × 3 trials | 243 |
| TTFT reproduction measurement | 3 states × 8 prompts × 3 trials | 72 |
| Dense think=True | 2 sets × 27 prompts × 3 trials | 162 |
| num_ctx 32768 measurement | 3 sets × 27 prompts × 3 trials | 243 |
| DFlash measurement | 1 set × 8 prompts × 3 trials | 24 |
| Total | 1,500 |
Thinking is the feature where the model generates a “draft-thoughts” passage before writing its final answer. Temperature is the setting that controls output randomness: an LLM holds “candidates for the next token” with probabilities attached, and temperature 0.0 means “pick the highest-probability candidate” (the previous article has an explanation in terms of sharpening or flattening the probability peak). A token is a fragment of text finer than a word; in Japanese, roughly one to a few characters correspond to one token. All the units in this article like “tok/s” and “TTFT” are counting these tokens.
Scores are shown with a metric called F1. F1 rolls “the power to not miss confidential material (Recall)” and “the hit rate of the warnings raised (Precision)” into one value, with 1.0 as a perfect score (it is the harmonic mean of the two — an average that drops sharply when either side is low). Precision is “of the things judged confidential, the fraction that actually were confidential”; FPR is “of the documents that are actually safe, the fraction mistakenly judged confidential.” Both are metrics about false alarms, but they differ in the denominator: “the number judged confidential” versus “the number actually safe.” FPR is better the lower it is.
Let me state up front how the three gray prompts are handled. F1, Recall, Precision, and FPR are metrics over the five confidential documents and the safe documents — that is, the questions whose correct answer is definitively “confidential” or “safe.” Since no correct answer can be defined for the three gray prompts, they are excluded from the calculation of these metrics. Gray is tallied separately, to observe judgment tendencies.
One more important caveat. There are only five confidential prompts (three trials each, so 15 records in the tallies). The difference between Recall 1.0 and 0.8 is the difference between “5 out of 5” and “4 out of 5” — that is, a single prompt. The difference between F1 1.0 and 0.8889 likewise arises from one prompt. At this scale, you cannot declare “it degraded” or “it improved.” Please read every number in this article within that constraint.
Chapter 5: The Test Results
5-1 Axis 1 (Confidentiality Judgment) — Main Measurement, 7 Sets
The scores from the main measurement (7 sets; all on Ollama 0.32.9). The Dense measurements (gemma4:31b and qwen3.6:27b) are shown separately in 5-2 and 5-5.
| Model | think | FPR | Recall | Precision | F1 |
|---|---|---|---|---|---|
| gemma4:26b | false | 0.0000 | 1.0000 | 1.0000 | 1.0000 |
| gemma4:26b | true | 0.0000 | 0.8000 | 1.0000 | 0.8889 |
| muse-glimmer | false | 0.2000 | 1.0000 | 0.8333 | 0.9091 |
| muse-glimmer | high | 0.0000 | 1.0000 | 1.0000 | 1.0000 |
| muse-glimmer | low | 0.3333 | 1.0000 | 0.7500 | 0.8571 |
| qwen3.6:35b | false | 0.0000 | 0.8000 | 1.0000 | 0.8889 |
| qwen3.6:35b | true | 0.0000 | 1.0000 | 1.0000 | 1.0000 |
The points that can be read off as fact:
- muse-glimmer holds Recall 1.0000 in all three states — false/low/high. It missed none of the five confidential prompts. When its score drops, it drops on the Precision side (calling safe things “confidential”)
- qwen3.6:35b (think=false) has Recall 0.8000. Its score drops on the missing side. It also judged the three gray prompts 100% “safe” (C_safe_rate 1.0000; of the other six sets, all are at 0.3333 except muse-glimmer (think=high), which alone is at 0.0000)
- muse-glimmer (think=high) judged the three gray prompts 100% “confidential” (C_confidential_rate 1.0000; of the other six sets, all are at 0.6667 except qwen3.6:35b (think=false), which alone is at 0.0000). As stated in Chapter 4, the three gray prompts are excluded from the F1 and FPR calculations, so this judgment tendency and the FPR of 0.0000 are not in contradiction
5-2 How Time and Accuracy Change With and Without Thinking
Chapter 3 of the previous article presented the following table (values at the time, on Ollama 0.24.0):
| Model | think=False | think=True | Time taken | F1 (False → True) |
|---|---|---|---|---|
| gemma4:26b | 3.20 s | 12.23 s | 3.8x | 1.0 → 0.889 |
| gemma4:31b | 13.38 s | 54.63 s | 4.1x | 1.0 → 1.0 |
| qwen3.6:35b | 4.81 s | 41.37 s | 8.6x | 1.0 → (6 empty responses) |
| qwen3.6:27b | 14.68 s | 97.17 s | 6.6x | 1.0 → 0.75 |
The previous conclusion, as published, was this:
On confidentiality judgment, thinking ate 4–9x the time and accuracy either didn’t change or got worse.
To sum up: for convergent tasks like classification, summarization, and QA, it looks like the right move is to default thinking off, and raise accuracy through prompt design and context injection instead.
Now the measurements from this time. Here is muse-glimmer alongside the two Dense models re-measured on the same Ollama 0.32.9. Unless otherwise noted, response time, generation speed, and TTFT throughout this article are medians over the 8 Axis 2 speed prompts (3 trials = n=24). They are not representative values over all 27 prompts. Peak memory alone is the peak across the full 27-prompt run in that think state.
| Model (Ollama 0.32.9, same conditions) | think off → on | Time ratio | F1 change |
|---|---|---|---|
| muse-glimmer (false → high) | 12.07 s → 31.33 s | 2.60x | 0.9091 → 1.0000 (went up) |
| gemma4:31b (False → True) | 13.31 s → 48.69 s | 3.66x | 1.0000 → 1.0000 (unchanged) |
| qwen3.6:27b (False → True) | 11.32 s → 86.34 s | 7.63x | 1.0000 → 0.8889 (went down) |
In this round’s runs, there are two sets whose F1 went up with thinking turned on: muse-glimmer (false → high, 0.9091 → 1.0000), and likewise qwen3.6:35b (false → true, 0.8889 → 1.0000), as seen in 5-1. In the previous round’s runs (Session 04), none of the five models (10 sets) measured at both think=True and think=False showed a rise (only equal or lower). As for time, muse-glimmer’s ratio of 2.60x is smaller than the 3.8–8.6x in the previous table. That said, per the Chapter 4 caveat, keep in mind that the F1 differences are one-prompt differences.
I also recorded the length of the thinking (median):
| Set | Thinking length (median) |
|---|---|
| muse-glimmer (think=high) | 1,261 characters |
| gemma4:31b (think=True) | 1,721 characters |
| gemma4:26b (think=True) | 2,129 characters |
| qwen3.6:35b (think=True) | 3,337 characters |
| qwen3.6:27b (think=True) | 3,757.5 characters |
muse-glimmer (think=high)’s thinking, at 1,261 characters, was the shortest of the five sets.
5-3 TTFT at think=false, and a Stream Observation
First, the reproduction measurement of TTFT (the wait until the first token appears). For muse-glimmer’s three states, I measured 8 prompts × 3 trials × 3 states = 72 inferences.
muse-glimmer think=false TTFT median 3,295.6 ms (n=24)
muse-glimmer think=low TTFT median 560.6 ms (n=24)
muse-glimmer think=high TTFT median 561.2 ms (n=24)
The inversion — think=false, with thinking cut off, taking longer to produce its first token than think=low or high — reproduced in this reproduction measurement as well.
Next, to look inside that long wait, I took a single prompt and recorded, one by one, the arrival times of the chunks in the stream (the mechanism by which the response flows in bit by bit). The results:
think=false first chunk arrived at 2,942 ms (3 characters) → then 6 chunks at 39 ms intervals
of the 2,566 ms Ollama reports as eval, about 2,330 ms were silent
think=low first chunk arrived at 738 ms (3 characters) → then 54 chunks at 39 ms intervals
the 2,495 ms Ollama reports as eval flowed out almost entirely as chunks
think unspecified first chunk arrived at 749 ms (starting with thinking)
→ the default is thinking enabled
Here, “eval” is the time Ollama reports as “spent generating (decoding) the response.” The processing that reads in the input prompt (prompt_eval) is accounted separately and not included in eval. What can be said as fact is the following. The time spent generating (eval) was nearly the same for think=false and think=low (2,566 ms and 2,495 ms). The difference is in how that time was used: in think=low, nearly all of it flowed to the screen as chunks, whereas in think=false, about 2,330 ms — roughly 90% of eval — passed in silence. Also, when think is left unspecified, the default was thinking enabled.
The two Dense models measured under the same conditions this time do get genuinely shorter response times at think=False (as in the 5-2 table: gemma4:31b at 13.31 s versus 48.69 s with thinking on; qwen3.6:27b at 11.32 s versus 86.34 s). The four models from last time (the values at the time, at the top of 5-2) showed the same tendency. With muse-glimmer, setting think=false did not yield a time saving in this form.
5-4 Comparing Responses Across the Runtime Update
Chapter 5 of the previous article presented the following table and prescription about reproducibility at temperature=0:
I ran each model three times under identical conditions and compared the length of the returned text. […] cases where the gap between the longest and shortest exceeded 30 characters (or 5%) — that is, cases where “the answer wobbled under identical conditions” — came up 11 times across the board.
| Family | Number of wobbling runs (conditions) |
|---|---|
| gemma4:26b | 5 |
| gemma4:31b | 2 |
| gemma3:12b | 2 |
| phi4:latest | 2 |
| Qwen family (all 7 sets) | 0 |
Qwen is bit-stable at temp=0; Gemma/Phi wobble. […] if bit-level reproducibility is a requirement, Qwen is the better choice.
This time, I matched up the response files from the previous test (Session 04) against this time’s response files, byte for byte. The only thing changed is the runtime version. The model weights are the same, the prompts are the same, and temperature 0.0 is the same. The results:
qwen3.6:35b(think=False) 81 of 81 files differ (0 exact matches)
gemma4:26b (think=False) 78 of 81 files differ (3 exact matches)
gemma4:26b (think=True) 81 of 81 files differ (0 exact matches)
Even qwen3.6:35b (think=False), which was “bit-stable” last time, produced different responses in 81 out of 81 files. However, within a given run, trial 1/2/3 were byte-identical. Determinism itself — same conditions, same answer — is preserved; what changed is “sameness across versions.”
Meanwhile, even though the response text was almost entirely rewritten, the confidentiality-judgment scores came out as follows:
| Set | Session 04 (0.24.0) | This time (0.32.9) | |
|---|---|---|---|
| gemma4:26b (think=False) | F1 1.0000 | F1 1.0000 | Exact match |
| gemma4:26b (think=True) | F1 0.8889 | F1 0.8889 | Exact match |
| qwen3.6:35b (think=True) | F1 1.0000 | F1 1.0000 | Exact match |
| qwen3.6:35b (think=False) | F1 1.0000 | F1 0.8889 | The only one that moved |
For the three matching sets, not only F1 but FPR, Recall, Precision, and the breakdown of the gray layer all match exactly. What moved was just two prompts in qwen3.6:35b (think=False): B-03 changed from “confidential → safe,” and one gray prompt changed from “confidential → safe.” Both changes are toward the “safe” side — that is, toward missing. Here too, the Chapter 4 caveat applies: there are five confidential prompts, and the difference between F1 1.0 and 0.8889 is one prompt. At n=5, you cannot declare “it degraded.”
5-5 Speed and Memory
Chapter 2 of the previous article compared MoE and Dense with the following table (think=False; values at the time, on Ollama 0.24.0):
| Model | Architecture | Active parameters | Response time | Generation speed | Peak memory | Axis 1 F1 |
|---|---|---|---|---|---|---|
| gemma4:31b | Dense 31B | 31B | 13.38 s | 16.3 tok/s | 33.59 GB | 1.0 |
| gemma4:26b | MoE 26B/4B-active | 4B | 3.20 s | 76.3 tok/s | 21.93 GB | 1.0 |
| qwen3.6:27b | Dense 27B | 27B | 14.68 s | 14.3 tok/s | 33.24 GB | 1.0 |
| qwen3.6:35b | MoE 35B/3B-active | 3B | 4.81 s | 44.7 tok/s | 31.23 GB | 1.0 |
The previous conclusion, as published, was this:
In other words, “MoE = light” is only half true. More precisely: “MoE = fast. Whether it’s also light depends on the total parameter count.”
Now the measurements from this time (all on Ollama 0.32.9, think=False, same conditions):
| Model | Architecture | Response time (Axis 2, n=24) | Generation speed (Axis 2, n=24) | Peak memory (full 27 prompts) | Axis 1 F1 |
|---|---|---|---|---|---|
| qwen3.6:35b | MoE 3B-active | 2.34 s | 81.0 tok/s | 27.76 GB | 0.8889 |
| gemma4:26b | MoE 4B-active | 2.85 s | 75.9 tok/s | 22.93 GB | 1.0000 |
| qwen3.6:27b | Dense 27B | 11.32 s | 18.4 tok/s | 34.05 GB | 1.0000 |
| muse-glimmer | Dense ~30B | 12.07 s | 20.0 tok/s | 17.35 GB | 0.9091 |
| gemma4:31b | Dense 31B | 13.31 s | 16.1 tok/s | 40.64 GB | 1.0000 |

Let me state the table’s premises first. All five models are on Ollama’s default tags (all quantized at roughly Q4_K_M), and num_ctx (the context-length setting) is unspecified — i.e., at its default — for all of them. Measurements with num_ctx explicitly specified are shown in 5-9. Peak memory is the recorded peak of the resident memory of the process running the model (the physical memory the process actually occupied), and the measurement script outputs its values in binary (1,024³ bytes = GiB). Throughout the rest of this article they are all written as “GB,” but note that the actual unit is GiB.
One caveat here about units. muse-glimmer’s peak memory of 17.35 GiB appears to come in below the “18GB” displayed as the model’s file size. Converting 17.35 GiB to decimal GB gives about 18.6 GB, so if Ollama’s display is decimal, there is no contradiction. However, I have not confirmed whether Ollama’s “18GB” display is decimal or binary (unverified). So here I go no further than “a difference in unit systems can most likely explain it.”
muse-glimmer (think=false)’s peak memory was 17.35 GB — the lightest of the three Dense models, and 5.6 GB lighter than the MoE gemma4:26b (22.93 GB).
Next, the wait time. Here are the medians of TTFT (the wait until the first token appears), on the 8 Axis 2 speed prompts, n=24, think=False (as stated at the top of this section, these are not representative values over all 27 prompts):
| Model | TTFT (median) |
|---|---|
| gemma4:26b | 306 ms |
| qwen3.6:27b | 397 ms |
| gemma4:31b | 535 ms |
| muse-glimmer | 2,972 ms |
The TTFT of gemma4:26b, qwen3.6:27b, and gemma4:31b fits almost entirely within “load (loading the model) + prefill (reading in the input prompt).” Only muse-glimmer greatly exceeds that sum (about 411 ms), with its TTFT eating into the eval (generation) period. This is the silent time we saw in 5-3.
For reference, muse-glimmer (think=false)’s TTFT also moves with the type of prompt: Axis 1 safe 4,007 ms, Axis 3 quality 4,390 ms, Axis 1 confidential 6,976 ms, and Axis 1 gray 10,885 ms (about 3.7x the Axis 2 speed prompts), with a median of 4,662 ms across all 27 prompts. The between-run difference (2,972 ms versus 3,296 ms; 5-3) falls on the relatively small side within this operating range.
Finally, generation time per token. This value is Ollama’s reported eval — the time it declares as generation (decoding) — divided by the number of output tokens. The calculation “response time minus TTFT, divided by output tokens” cannot be used for muse-glimmer: as shown above, muse-glimmer’s TTFT sits inside eval, and generation is progressing during the silence, so subtracting TTFT would underestimate the generation time.
| Model (think=False) | Generation time per token |
|---|---|
| gemma4:26b (MoE) | 13.18 ms/tok |
| muse-glimmer (Dense) | 50.05 ms/tok |
| qwen3.6:27b (Dense) | 54.30 ms/tok |
| gemma4:31b (Dense) | 61.98 ms/tok |
muse-glimmer was the fastest per token among the three Dense models. That said, the gap to qwen3.6:27b (54.30 ms/tok) is about 8% — not a large one. The far bigger gap is to the MoE gemma4:26b (13.18 ms/tok). Note that muse-glimmer’s response time (12.07 s) being longer than qwen3.6:27b’s (11.32 s) is not because it is slower per token, but because it outputs more tokens (median 221 tokens versus 192).
5-6 DFlash (Speculative Decoding)
DFlash is an implementation of speculative decoding — a speed-up technique in which a small draft model proposes candidates first, and the main model inspects them and decides whether to accept them. Here are the measurements on the muse-glimmer:30b-mlx tag, which uses DFlash:
standard tag 19.98 tok/s / response 12.07 s / memory 17.35 GB
muse-glimmer:30b-mlx (DFlash) 30.78 tok/s / response 8.83 s / memory 19.49 GB
ratio 30.78 / 19.98 = 1.54x (ratio of generation speeds)
The standard tag’s 19.98 tok/s is the same value as the 20.0 tok/s in the 5-5 table (a difference of rounding). The 1.54x ratio is computed as the ratio of generation speeds (tok/s). Ollama’s official blog says, “With DFlash, Muse Glimmer runs 1.5×–1.8× faster on Apple Silicon.” Meta’s official blog breaks this down by machine, stating explicitly: “DFlash speculative decoding increasing Muse Glimmer decode speed by 3.1 times on RTX 5090, 1.8 times on M5 Max, and 1.5 times on M4 Max.” The machine measured in this article is an M4 Max, and the measured ratio was 1.54x. That is nearly identical to the published M4 Max figure (1.5x). Note that the “3.1x” on the Hugging Face model card is also, as the quote above shows, a value on an RTX 5090 and must be treated as a different animal. Standard speculative decoding with verification is described as designed so the output does not change — only candidates that pass the main model’s inspection are accepted — but I have not been able to confirm whether DFlash is that implementation. On the DFlash build I measured only speed; accuracy (Axis 1) is unmeasured.
How This Relates to Last Time’s “MLX Is Slower”
In the previous measurements, using MLX came out slower, if anything. That looks like the opposite of this time, but what was measured is different.
What I measured last time was the path that calls mlx-lm directly, bypassing Ollama. Qwen3.6-35B-A3B-4bit had a median response time of 14.03 seconds, while the equivalent weights run through Ollama as qwen3.6:35b (Q4) took 4.81 seconds — about 2.9x slower on the direct path. On top of that, thinking could not be disabled on that path, and there was a problem of misjudging public information as confidential in the confidentiality judgment (a 100% false-positive rate on layer A), so it was rejected. Separately, I also measured qwen3.5:27b-mlx-bf16 (bf16, 54GB), where swap ballooned to 10GB and it came out at 7.65 tok/s — the slowest of all models at the time.
What I measured this time is the muse-glimmer:30b-mlx tag, which goes through the MLX engine that Ollama carries internally. The path is different. And what Ollama’s official blog cites as the reason for the speed gain is DFlash (speculative decoding), not MLX itself.
In other words, the accurate statement is not “MLX used to be slow and got faster,” but “a different mechanism was measured via a different path.” I did not measure the MLX engine with DFlash turned off this time, so I cannot separate how much of the 1.54x comes from DFlash and how much from MLX. Memory is 2.14GB (12.3%) heavier than the standard tag.
5-7 Empty Responses and Anomalies
As a health check, I scanned both runs (the main measurement’s 567 inferences plus the memory-only re-measurement’s 189 — 756 files in total). Anomalous responses — leaked tool-call notation and the like — numbered zero. As for empty responses: muse-glimmer had 0 in all three states; qwen3.6:35b (think=True) had 9 in the main measurement (3 trials) (Session 04 had 6), and 3 in the memory-only re-measurement (1 trial only), on the same three prompts (speed-long-01, quality-summary-01, quality-summary-02). All empty responses occurred on Axis 2 and Axis 3 prompts, with zero on Axis 1, so there is no impact on F1. Note that qwen3.6:27b (think=True)’s 3 empty responses came out in a different run from these 756 files (the Dense think=True measurement, 162 inferences) and are not included here.
5-8 Failures on the Measurement Side
Following the previous article, I record the failures of the measurement itself as well. Last time I wrote about “asking gemma4:26b to list Japan’s prime ministers in order, and watching the count balloon to 732.” This time there are two entries.
Failure 1: The aggregation script was collapsing muse-glimmer’s three states into one. muse-glimmer’s think setting takes the three values false/low/high, which were recorded as strings. The aggregation script converted them to booleans — and in Python, bool('low'), bool('high'), and even bool('False') all come out True, because every non-empty string is truthy. As a result, 243 records got crushed into a single group. No error, no warning — the summary tables looked complete. “Looking plausibly finished” and “being correct” are different things: the same lesson as last time.
Failure 2 (more precisely, a measurement caveat): Memory moves by up to ±2.8GB between runs. gemma4:26b (think=False)’s response time and generation speed did not move much between runs. Measuring gemma4:26b (think=False) three times under the same conditions, peak memory scattered across 22.93 GB / 25.73 GB / 23.36 GB. Meanwhile, the differences in the same gemma4:26b (think=False)’s response time and generation speed stayed within 0.09 seconds / 3.4% under the same conditions. But that is an observation limited to these two metrics on gemma4:26b. muse-glimmer (think=false)’s TTFT, as shown in 5-3 and 5-5, moved between runs from 2,972 ms to 3,296 ms (about 11%), so I cannot go as far as “speed never moves at all, regardless of metric or model.” To discuss small memory differences, you have to compare within the same run. Of the values in the 5-5 table, the three for muse-glimmer, gemma4:26b, and qwen3.6:35b come from a single run — the memory-only re-measurement. The two for gemma4:31b and qwen3.6:27b come from a different run (dense-same-condition). However, the gap between those two and the other three (16–23GB) far exceeds the between-run scatter (±2.8GB), so the runs being different does not affect the conclusion (the ranking: muse-glimmer is the lightest).
5-9 Measurements at num_ctx 32768
The checker built on the previous test assumes a 32k context window, but all measurements so far this time leave num_ctx unspecified (Ollama’s default). So I also measured with tags that explicitly set num_ctx to 32768 (3 sets × 27 prompts × 3 trials = 243 inferences). The 32k build of muse-glimmer was created from a Modelfile adding PARAMETER num_ctx 32768 to FROM muse-glimmer.
| Set (num_ctx=32768, think=False) | Response time (Axis 2, n=24) | Generation speed (Axis 2, n=24) | Peak memory (full 27 prompts) | Axis 1 F1 |
|---|---|---|---|---|
| gemma4-26b-32k | 2.85 s | 75.2 tok/s | 23.36 GB | 1.0000 |
| qwen36-35b-32k | 2.36 s | 81.1 tok/s | 28.00 GB | 0.8889 |
| muse-glimmer-32k | 12.68 s | 19.0 tok/s | 18.71 GB | 0.9091 |
F1 for all three sets was the same as with num_ctx unspecified (5-1).
However, the “increment” over the default num_ctx could not be detected with this design. The default-num_ctx comparison values came from a different run, and between same-condition runs alone, gemma4:26b (think=False) scatters across 22.93 GB / 25.73 GB (5-8). The 32k value of 23.36 GB falls between those two. muse-glimmer likewise scatters at the default across 17.35 GB / 18.27 GB, and its 32k value was 18.71 GB. The measurement scatter is larger than the effect size, and the increment could not be measured. Measuring it correctly would require putting the default tag and the 32k tag side by side within the same run.
5-10 Axis 3 (Response Quality)
I read 7 sets × 6 prompts (2 code, 2 explanation, 2 summarization) = 42 responses — the body text of one representative trial out of each set’s three. Note the scale: three categories at n=2 each.
These are not problems with a single determinate answer, so I assign no scores. I will list only things that are visibly broken and things that clearly differed by setting.
A note for readers of this English edition: the prompts in this test are in Japanese, and the responses are Japanese text. Quoting the responses in translation would alter the evidence itself, so below I keep the original Japanese fragments and add English glosses in parentheses.
Things That Were Broken
gemma4:26b (think=false)’s code example was a syntax error. In its explanation of decorators, it wrote time() — with a full-width opening parenthesis. The full-width parenthesis (U+FF08, FULLWIDTH LEFT PARENTHESIS) is the variant used in Japanese text; it looks nearly identical to the ASCII ( (U+0028), but it is a different character, and Python rejects it: SyntaxError: invalid character '(' (U+FF08). I actually ran the code through py_compile to confirm. It appears in all three trials, so it is not a one-off accident. Across all 756 responses in the runs, these 3 files are the only ones where a full-width parenthesis crept in. It does not appear in the think=true version.
muse-glimmer (think=false)’s response was logically broken. To the question of splitting three apples between two people, it wrote 「残り1個はCが食べる」(“the remaining apple is eaten by C”) — introducing a third party into a problem that only has two people. Elsewhere in the same response, in a calculation splitting by weight, it wrote 「AとBを1人に、BとCをもう1人に」(“A and B to one person, B and C to the other”), putting B on both sides. This breakdown does not appear in think=low or think=high.
This is the fastest of the three states by response time (Axis 2, n=24 median: 12.07 s; see 5-2). However, on Axis 1, the lowest score among the three states belongs to think=low (F1 0.8571 / FPR 0.3333); think=false (F1 0.9091 / FPR 0.2000) sits in between. Being the fastest setting does not mean it is also the lowest-accuracy one. Accuracy drops the most at think=low, while response quality broke down at think=false — so the three do not all point in the same direction.
qwen3.6:35b (think=true) produced zero-character bodies on both summarization prompts. It generated 9,477 and 7,617 characters of thinking respectively, then spent 53 seconds returning nothing. This is the breakdown behind the “9 empty responses” counted in 5-7. The Axis 3 summarization category was wiped out: 2 prompts × 3 trials = 6 records, all empty. Against an instruction to “summarize briefly,” the thinking exceeded 9,000 characters.
Things That Clearly Differed by Setting
The amount of think changed not the correctness of the answers, but the choice of implementation. On the prompt asking for a Sieve of Eratosthenes, all 7 sets produced a correct sieve. But while 6 sets used a list implementation — is_prime = [True] * (n + 1) — only muse-glimmer (think=high) chose a memory-efficient implementation using bytearray and slice assignment. I verified it returns the same results as the standard implementation across the entire range n=0–499.
The same thing happens on the deduplication prompt. think=false made “the version that tracks seen items with a set” the main answer and relegated the dict.fromkeys version to a side note; in think=low, that order was reversed. Which one gets recommended flips.
The direction in which responses lengthen as think increases was reversed between models. Body lengths on the two code prompts (unit: characters):
| Deduplication | Prime sieve | |
|---|---|---|
| muse-glimmer (false / low / high) | 939 / 1073 / 852 | 898 / 858 / 717 |
| gemma4:26b (false → true) | 1567 → 1972 | 1472 → 1508 |
| qwen3.6:35b (false → true) | 420 → 963 | 634 → 1523 |
gemma4 and qwen3.6 get longer when think is turned on; muse-glimmer is shortest at high. It is not monotonic, though: on deduplication, muse’s low is the longest. The only thing common to both prompts is “high is shortest,” and in the summarization category the direction does not line up (193 / 169 / 175 characters). This is merely a tendency seen on the two code prompts — n=2.
Only gemma4:26b (think=false) self-reported its character count. Against the instruction 「300字程度に要約」(“summarize in about 300 characters”), it appended 「(298文字)」(“(298 characters)”) at the end of its response — but the actual count is 275–286 characters (it varies depending on whether you count line breaks or the self-report string itself, but no way of counting reaches the reported value). On the “about 200 characters” prompt it likewise wrote 「(184文字)」(“(184 characters)”) when the actual count was 174–181. On both prompts it over-reported, and the direction of the discrepancy is toward appearing closer to the instructed target. The think=true version writes no character count, and muse-glimmer and qwen3.6 write none in any state.
The Limits of This Section
The observations above are the result of reading 42 responses, once each. I can point out what is broken, but I have done no ranking of “which response is better.” Please take these as-is as constraints: there is one reader, only two prompts per category, and only one representative trial read per set.
Chapter 6: Differences from Prior Reports
Before publishing, I looked into what others have reported about Muse Glimmer. What I could confirm is six third-party articles — this is not exhaustive. The following is an organization within that range.
What Has Already Been Reported
- Speed-up figures for DFlash — Meta officially publishes 1.5x on M4 Max and 1.8x on M5 Max (Meta AI Research, 2026-08-10). wavect.io (around 2026-08-10 to 11) has independently measured speed on M4 Max and M5 Max
- That thinking cannot be fully switched off — kotetsu_yama on Qiita (2026-08-11) points out, in an AMD Strix Halo / llama.cpp environment, that “there is no implementation that fully disables thinking; even set to low, a thinking phase executes.” An X post by Tom Turney (around 2026-08-09) makes a similar point that “the reasoning strength default is effectively high” (★ I could not retrieve the text of this post itself; this is confirmed via secondary sources)
- General-benchmark scores — Artificial Analysis (Intelligence Index 35) and BenchLM.ai (52.5/100, ranked 112th of 218 models) have published scores. I have not been able to confirm the publication dates of either
What This Article Adds
- Putting speed, memory, and classification-task scores side by side in a single test, on Apple Silicon. Within the range I checked, I could not find an article that brings these three together
- The numerical inversion that think=false is slower than think=low (TTFT median 3,295.6 ms versus 560.6 ms). There are already two reports that “thinking can’t be switched off,” but I could not find an example showing this direction of inversion in numbers
- Matching responses byte-for-byte across a runtime update, and separating sameness of output from sameness of judgment. I could not find a prior example of this observation
Disclaimers
- All I checked is six third-party articles. This is not an exhaustive survey
- “I could not find it” is not proof that “it does not exist.” It may well simply be that my search did not reach it
- Among the six articles checked, the number evaluating Muse Glimmer on confidentiality judgment or classification tasks was zero
Chapter 7: Changes from Last Time’s Conclusions, and a Discussion
Let me set the results against the three conclusions from last time.
(1) On “For Convergent Tasks, Thinking Is a Tax”
The previous prescription — “for convergent tasks, default thinking off” — did not apply to muse-glimmer. The central reason is that, as 5-3 showed, even at think=false about 90% of the eval time passes in silence, so “off” in the time sense never actually materializes.
This time, qwen3.6:35b’s F1 also went up (0.8889 → 1.0000). But its character is different. Whereas muse-glimmer’s rise happened within a single run, as an effect of thinking, qwen3.6:35b’s rise comes from its think=false baseline itself having dropped, from last time’s 1.0000 to this time’s 0.8889 (see Chapter 7, (2)). In the previous round’s runs (Session 04), there was not a single case of F1 rising with thinking.
The two Dense models measured this time do get genuinely faster at think=False, and the four models from last time (values at the time) showed the same tendency. The prescription remains valid — but this time’s observation is that one model has appeared for which the premise “turn it off and you don’t pay the time” does not hold.
(2) On “Even at temperature=0, Output Still Wobbles — Though It Depends on the Vendor”
What became clear this time is that last time’s observation — “Qwen is bit-stable” — only holds within the same runtime. Merely raising Ollama from 0.24.0 to 0.32.9 changed qwen3.6:35b (think=False)’s responses in 81 out of 81 files.
The phenomenon can be explained like this. An LLM searches over “candidates for the next token,” each with a probability, and temperature 0.0 is the setting that picks the highest-probability one. The model is identical, but with a different runtime, what gets picked as the next token changed. The internal computation is floating-point arithmetic, and when the order of operations or the implementation changes, the probability values shift slightly in the low-order digits. For tokens where first and second place are nearly tied, that slight difference can swap the ranking. This time, it was not the model but the runtime version that reached into that gap. Once the top token flips once, the rest of the text follows a different path, so the whole response gets rewritten. Note that a runtime update may also change things other than the numerical implementation, and I have not identified which change mattered this time (Chapter 9). A runtime update changing the numerical-computation path is not unusual in itself, and I don’t intend to attach a good-or-bad judgment to it.
On the other hand, as 5-4 showed, even though the response text was almost entirely rewritten, the confidentiality-judgment scores matched exactly in 3 of 4 sets — down to FPR, Recall, Precision, and the gray-layer breakdown. “Sameness of output” and “sameness of judgment” are different things, and the previous article was only looking at the former. The practical implication, I think, is this: if you need bit-level reproducibility, pinning the model name is not enough — you have to pin the runtime version as well. If your use case only needs the judgments to match, then within this test’s range, 3 of 4 sets matched across the update — but you also need to take home the fact that one set moved on two prompts, and the direction it moved was toward missing.
(3) On “MoE = Fast. Whether It’s Also Light Depends on the Total Parameter Count”
This framing holds up in this time’s table as well. The two MoE models were again fast (qwen3.6:35b (think=False) at 2.34 s, gemma4:26b (think=False) at 2.85 s), and qwen3.6:35b’s memory, at 27.76 GB, was in line with its total parameters. On top of that, muse-glimmer (think=false) — a Dense model with about 30B total parameters — posted 17.35 GB peak memory, the lightest value in the table. I take this as one example added to the earlier framing: “Dense, too, can sometimes be made light, depending on the design.”
(4) On the Direction of Breaking
As 5-1 showed, muse-glimmer holds Recall 1.0000 in all three states, and when its score drops, it drops on the Precision side. It errs by calling safe documents “confidential” — that is, it fails toward stopping whatever is doubtful. By contrast, qwen3.6:35b (think=false) fell on the missing side, with Recall 0.8000, and judged the three gray prompts 100% “safe.” The same “F1 short of perfect” breaks in completely different directions. For the use case of a confidentiality checker, I consider this difference in direction more important than the F1 number itself. And this difference is invisible if you look only at the single number F1.
Chapter 8: Inference — The Parts That Are Not Observation
Everything written in this chapter is inference. The observations end with Chapter 5; from here on, this is my interpretation.
Inference 1: muse-glimmer may be a “model designed on the premise of thinking,” with no real provision for switching thinking off. This is inference. The grounds: even at think=false, about 90% of eval (roughly 2,330 ms out of 2,566 ms) passed in silence; that eval time was nearly identical to think=low’s (2,495 ms); and the default with think unspecified was thinking enabled. I consider this consistent with Meta positioning the model as an “agentic model.” However, I have not observed what is happening during those silent ~2.3 seconds. This remains conjecture from circumstantial evidence.
Inference 2: The reason thinking helped on a convergent task may be that the thinking is short. This too is inference. muse-glimmer (think=high)’s thinking has a median of 1,261 characters — shorter than gemma4:26b (think=True)’s 2,129 and qwen3.6:35b (think=True)’s 3,337. It may simply be that muse-glimmer never reaches the failure mechanism observed in the previous article — thinking eating the output budget until the body gets cut off. If so, this is not “different behavior” but “the same mechanism, just not triggered.” I have only seen a correlation; causation is unverified.
Inference 3: The extreme 32:2 GQA ratio may be the reason for the light memory. This too is inference. It rests on what Raschka’s explainer reports; I have not measured it. Also, what this mechanism mainly reduces is the memory for remembering context (the KV cache), and its size depends on the context-length setting. The default context length in this run is unconfirmed (Chapter 9).
Inference 4: Perhaps only one thing can be called “clearly different from the others.” This too is inference — or rather, a statement of my own judgment. Paying the thinking time even at think=false is the one thing structurally different from the other four models (which genuinely get faster at think=False). The rest — short thinking, light despite being Dense, tipping gray toward confidential — are differences of degree, and the mechanisms themselves may be the same. And this is one model. What can be said at this point is not “this is what the new generation looks like” but “out of five models, one appeared that differs on exactly one point.” Generalizing would require a second model with the same properties.
Chapter 9: Unverified Items — What I Didn’t Read and What I Couldn’t Measure
As in the previous article, I list what I have not been able to confirm, as its own section.
- What I read for Axis 3 is only 7 sets × 6 prompts = 42 responses. Last time I read 16 sets × 6 prompts = 96. The way of reading per set (all 6 prompts, one representative trial) is the same; the total shrank because fewer sets were measured. Only one representative trial of each set’s three was read; for the rest I looked only at body length. The 6 prompts are 2 code, 2 explanation, 2 summarization — two per category. The observations in 5-10 were read at this scale
- The confidential prompts remain B-01 through B-05 — n=5. Let me stress once again that F1 differences move on the scale of a single prompt
- On the total parameter count: why the Hugging Face model card’s ~29.6B and Ollama’s displayed 27.9B + 1.9B vision projector don’t match. What each is counting is unconfirmed
- Whether muse-glimmer’s quantization, Q4_K_M at 18GB, is identical to Meta’s published K-Quant-17GB is unconfirmed
- I have not confirmed how many tokens Ollama 0.32.9’s default context length is. This round’s main measurements all leave num_ctx unspecified (the default), and whether that matches the previous assumption of 32k has not been pinned down
- The memory increment when num_ctx is set to 32768. I did measure runs with 32k explicitly specified (5-9), but the increment over the default was buried in the between-run scatter (±2.8GB) and could not be detected
- What is being generated during think=false’s silent ~2.3 seconds
- The accuracy of the DFlash build (muse-glimmer:30b-mlx). Only speed was measured
- Which change in Ollama caused “all the responses got rewritten”
- The Axis 1 F1 figures and the like in this article are my readings of the aggregation script’s output; I did not recompute F1 myself from the raw responses
- Why the full-width parenthesis in 5-10 occurred only in gemma4:26b’s think=false, and only on this one prompt, is not understood
Chapter 10: Summary — Sameness of Output, and Sameness of Judgment
Here are last time’s three conclusions set against this time’s results.
| Last time’s conclusion | This time’s result |
|---|---|
| “MoE = fast. Whether it’s also light depends on the total parameter count” | Holds. But one example was added of “Dense can also be light, depending on the design” (muse-glimmer (think=false), 17.35 GB) |
| “Stability at temp=0 depends on the vendor. If reproducibility is required, Qwen” | Now carries the condition “within the same runtime.” Sameness of output and sameness of judgment are different things, and the latter matched across the update in 3 of 4 sets |
| “For convergent tasks, default thinking off” | Still valid for the other four models. Does not apply to muse-glimmer (it pays the thinking time even at think=false, and F1 went up at think=high — though n=5) |
The most important observation this time, I believe, was not Muse Glimmer itself but the fact that “merely updating the runtime rewrote all the responses, yet the confidentiality judgments barely changed.” In an environment where you keep using a local LLM, output can change without changing the model, through updates to the runtime underneath. Meanwhile, consistency at the level of judgment is more likely to hold than that — within this test’s range, that is what was observed.
Scope Declaration (Reprise)
Once more, at the end, I repeat the declaration from Chapter 3. Every result and every piece of discussion in this article lives on top of the convergent task of judging confidential information — a classification whose correct answer converges to a single point. For divergent/exploratory tasks — “reflect on a philosophical question,” “solve a hard problem requiring multi-step reasoning” — the evaluation axes, the prompts, and the judgment criteria would all be different animals. muse-glimmer’s “can’t switch thinking off” property showed up here as a time constraint on a convergent task, but on a divergent task it could earn a different evaluation. That, however, is not tested in this article. Please read this neither as “a conclusion about Muse Glimmer in general” nor “a conclusion about local LLMs in general,” but as “a record of re-running the previous local-confidentiality-checker test with a new model and a new runtime.”
The next time a new model comes out, I intend to measure it with this same test and update this table.
About Soul Resonant Works
Soul Resonant Works is a solo venture developing seven local AI systems. Starting from zero programming experience, the development is progressing through collaboration with AI.
🌐 Soul Resonant Works:
→ https://sr-works.net/en/index.html
📝 This blog publishes the entire development process as a serialized journal.
CubePlot (free version available)
CubePlot is the first product from Soul Resonant Works — it turns a CSV into a 3D scatter plot you can rotate and zoom. No install, no sign-up: it’s a single HTML file you open in your browser, and it works offline. CubePlot itself does not send the data you load outside your machine — external network traffic is blocked at the browser level (CSP). Start with the free version.
▶ Product page: https://sr-works.net/en/cubeplot/
▶ Get it (commercial license, USD $39 + tax where applicable): https://soulworks8.gumroad.com/l/pzcij
If you found this article useful, please share it.
Leave a Reply