This article was originally written in Japanese and translated into English with AI assistance. Please note that some expressions may carry nuances from the original Japanese.
🇯🇵 日本語版はこちら / Japanese version
Series: Short Piece (Part 7)
On a MacBook Pro M4 Max (64GB), I ran 10 models × 27 prompts × 3 trials — 1,296 inferences in total — over 11 hours. There was just one question: “Can a local LLM check for confidential information without ever sending the data to the cloud?” This article is a first-hand record of how that test was designed, what it found, and the landmines I stepped on along the way.
Chapter 0: How Do You Handle Data You Can’t Send to the Cloud, Right Where You Are?
It’s been about two years since AI (LLMs) reached practical usefulness. From everyday questions to work support, they’re a genuinely convenient tool that takes over the time-consuming work of researching and analyzing things yourself. They’re contributing to work efficiency too — but the moment you throw personal information or a confidential document you handle at work into a cloud LLM like ChatGPT, that data leaves the company.
There’s an awkward dilemma here. Today’s work tools are mostly documents created with Office-type software, plus email, PDFs, and so on — a lot of it tied to personal information. It’s not just your own personal information either; it includes personal and sensitive information belonging to business partners. You want to use AI to raise work efficiency, but if something contains personal information, you can’t throw it at the AI. So do you check sentence by sentence for personal or confidential information? That’s not easy either. You want to use AI because you don’t want to do it yourself, and yet you end up having to check the entire text yourself just to be able to hand it to the AI. You can now set “don’t use the data I send for training,” but there’s unfortunately no way to prove “is it really okay?”
“I want to ask an LLM to check whether confidential information is included, as part of having AI do work on my behalf. But if I have to send the confidential data to the cloud just to check it, that defeats the purpose” — the very act meant to protect confidentiality ends up being the act that leaks it.
The only way to cut through this contradiction is for the data to never leave at all — that is, to complete the confidentiality check entirely with a local LLM. If the judgment happens closed inside your own machine, the leak pathway itself simply doesn’t exist.
So, is it realistically workable? What I have on hand is a MacBook Pro M4 Max with 64GB of unified memory. The AI told me that what you can run on this is roughly in the 9–35B class. Can a local LLM of this scale handle the judgment of “is this document confidential, safe, or gray” at a level of accuracy and speed you could call practical? And if it can, what should you use as the basis for choosing a model? I decided to look into exactly that.
I’d originally bought this Mac with the intention of running AI locally on it, so this finally feels like arriving at its intended use. I used this Mac to run local LLMs, compared them under identical conditions, and took benchmarks.
As It Happens, OpenAI Was Working on the Same Problem
While I was designing this benchmark, in the latter half of May 2026, there was a piece of news I hadn’t yet heard. A little before that, in late April, OpenAI had released for free a small 1.5B (1.5-billion-parameter) model called “Privacy Filter” that detects personal information in text. It runs locally and lets you avoid ever sending the raw data out. It was, essentially, OpenAI’s own answer to exactly the problem I was trying to solve — handling confidentiality locally, right at hand.
I only came across this OpenAI news after finishing the tests described below and starting to write this article — and rather than disappointment at the overlap, what I felt more strongly was, “so this problem really is real.” The fact that a world-leading research lab went so far as to build a dedicated model and give it away for free tells you how urgent the problem of “handling confidentiality without sending data out” really is. That one individual, around the same time, independently arrived at the same need without knowing about it — I chose to take that as validation that the problem itself was correctly framed.
That said, my approach is the exact opposite of OpenAI’s. To lay it out:
- Scale: OpenAI built a dedicated small 1.5B model. What this article tests is the general-purpose, open-source 9–35B models already sitting on my machine.
- Method: OpenAI’s approach was “build a dedicated product from scratch.” Mine asks, “can I repurpose the general-purpose model I already have, without building anything dedicated?”
- Philosophy of output: Privacy Filter returns the personal information it detects with the relevant parts redacted (blacked out). And OpenAI itself explicitly states, “this is a detection aid, not a guarantee of safety.” In other words, even at the cutting edge, the machine isn’t automatically guaranteeing safety — it’s positioned strictly as a tool that helps a human make the judgment. What this article aims for is the same: not “the machine silently rewrites it,” but an attention-flagging check that tells a human, “is it okay to send this out?”
As a side note, Privacy Filter’s 1.5B is an order of magnitude smaller than the LLM models I evaluated, and since it only pastes over redactions without generating text, it runs on an entirely ordinary laptop. The 9–35B generative models this article deals with are a different animal, both in size and in the nature of the work.
The benchmark shown in this article isn’t a general “local LLM shootout.” It’s a test designed to answer exactly one question: “does a local confidentiality checker actually work?” Because of that, both how the models were chosen and what metrics were measured were prepared by working backward from that single purpose. And because a local confidentiality checker is the kind of task with one determinate answer, the test content is also narrowed to tasks of the “answer converges to one point” type (what I’ll call “convergent” tasks below).
Note also that this is distinct from lightweight, OS-resident models like Apple Intelligence (roughly 3B class). Those are aimed at lightweight tasks like summarizing or paraphrasing; what’s dealt with here is the 9–35B mid-weight layer aimed at “judging business documents.” By the time you finish reading this article, I hope it will be clear, from the measured results of 1,296 inferences, which class of model and how to choose it, to make a business task like confidentiality judgment practical on your own 64GB-unified-memory Mac.
Chapter 1: What Did I Measure, and How — And What This Test Can’t Answer
To make comparisons trustworthy, you need to hold conditions constant. Before that, let me clarify up front exactly what this test was designed for, so as to bound the scope of the conclusions in advance.
The Scope of the Test
This benchmark is designed primarily around the judgment of confidential information as a convergent task — a classification where, given an input, the correct answer converges to one point among “confidential / safe / gray.” The evaluation axes, too, are built to measure “how fast and accurately you reach the correct answer.” So the conclusions of this article are limited to this kind of convergent task.
I expect that a different use case would call for an entirely different test design, and a different conclusion. For instance, for divergent/exploratory tasks like “reflect on a philosophical question to gain insight” or “solve a hard problem requiring multi-step reasoning,” the evaluation axis (depth of thought over speed), the prompts, and the judgment criteria would all be different animals. The “thinking” that this article, in its latter half, concludes is a “tax,” could instead be exactly the source of value in that other context. This article doesn’t test that part. Please read it not as “a conclusion about local LLMs in general” but as “a conclusion about whether a local confidentiality checker is viable.”
Three Evaluation Axes
- Axis 1 — Confidentiality-judgment accuracy: F1 / Recall / Precision / false-positive rate (FPR). The checker’s core function.
- Axis 2 — Speed and memory: time to first token (TTFT), generation speed (tok/s), peak resident memory (RSS). The core of practicality.
- Axis 3 — General quality: the response bodies for code generation, summarization, and QA, qualitatively evaluated by a human reader.
All three axes are designed to measure “how practically the model handles convergent tasks.”
Breaking Down the Four Metrics of Axis 1
The F1, Recall, Precision, and FPR used in Axis 1 are the confidentiality checker’s “report card.” The names sound intimidating, but what they’re doing is simple. Whatever judgment the checker makes ends up falling into one of these four buckets:
- Correctly flagged as “confidential” (actually confidential → judged confidential)
- Missed it (actually confidential → judged safe) ← the scariest one
- False alarm (actually safe → judged confidential) ← annoying false positive
- Correctly flagged as “safe” (actually safe → judged safe)
From these four counts, you can compute the following scores:
- Recall = the power to not miss anything. Of the things that are actually confidential, what fraction did it catch? For a confidentiality checker, a miss means a direct leak, so this is the most important metric — the goal is “zero misses.”
- Precision = the power to not false-alarm. Of the alarms it raised as “confidential,” what fraction were actually confidential? If this is low, you get flooded with false positives and end up having to review everything by hand anyway.
- FPR (false-positive rate) = the rate of false alarms. Of the things that are actually safe, how many did it mistakenly flag as “confidential”? The lower this is, the better the checker is at quietly letting safe things through.
- F1 score = the balance point between Recall and Precision. Reducing misses usually increases false alarms and vice versa, so this rolls both into a single number (0 to 1, higher is better). F1 = 1.0 means “a perfect score with zero misses and zero false alarms.”
Wherever “F1 = 1.0” appears repeatedly in this article, read it as meaning “that model scored a perfect score on confidentiality judgment.”
The design of these evaluation axes was done together with Claude. At this point, no confidential information was involved.
Held-Constant Conditions
# Ollama API. Explicitly pinning think to False is the key move.
import requests
resp = requests.post("http://localhost:11434/api/chat", json={
"model": "gemma4:26b",
"messages": [{"role": "user", "content": PROMPT}],
"think": False, # Must be set explicitly — some models default it on
"stream": False,
"options": {
"num_predict": 4000, # Cap the output length uniformly
"temperature": 0.0, # Aim for deterministic responses
},
"keep_alive": 0, # Explicitly unload after each inference (prevents memory contamination)
})
temperature=0.0, num_predict=4000, keep_alive=0 (explicitly unloading the model each time so a previous model’s residency doesn’t contaminate the memory measurement), three trials per condition, taking the median.
There’s a reason I explicitly pinned think to False. Before running the main benchmark, I ran a couple of preliminary experiments. The think feature does a pass, before generation starts, of “sprawling out whatever it’s thinking.” I interpret this as analogous to us drafting and revising when we write. This revising is nice to have, but the number of tokens usable in a single inference isn’t infinite, so I ran into a problem where the model would use up all its tokens while still revising, and never actually output a result. For one model, I’d initially set the token cap at 1500; it used that up, so I raised it to 4000, and it used that up too. It’s a feature that genuinely thinks things through, but it ended up thinking so much it never arrived at an answer. Since this is a local LLM, you don’t hit the kind of “contractual usage cap” you’d hit with a cloud LLM, but the time an inference takes keeps stretching out indefinitely, so you do need to set a generation cap. In this test, think didn’t play well with that setup, and kept hitting the cap before producing a result.
Based on these preliminary experiments, I settled on a policy of pinning think to False across the board. That said, for models that could produce proper output with think set to True in the preliminary experiments, I also ran the tests with True.
Scope and Scale
Since this is the foundation of the benchmark, let me lay out honestly every model I evaluated. I chose them so that Dense (traditional, uses all parameters every time) and MoE (sparse, uses only some “experts”), generation, quantization, and inference path (Ollama vs. mlx-lm directly) could each be compared without getting tangled together.
| Model | Architecture | Size | Reason for inclusion |
|---|---|---|---|
| gemma3:12b | Dense 12B | 8.1 GB | The lightweight-tier favorite. Strongest judgment accuracy in preliminary experiments |
| gemma4:31b | Dense 31B | 19 GB | Mid-weight Dense representative |
| gemma4:26b | MoE 26B/4B-active | 17 GB | MoE counterpart to gemma4:31b (same-family Dense vs. MoE) |
| phi4:latest | Dense 14B | 9.1 GB | For validating lightweight, parallel use cases (different character from gemma3:12b) |
| qwen3.5:9b | Dense 9B | 6.6 GB | Qwen’s previous-generation lightweight representative |
| qwen3.5:9b-mlx-bf16 | Dense 9B (bf16) | 18 GB | Control group for isolating the effect of quantization and inference path |
| qwen3.5:27b | Dense 27B | 17 GB | Qwen’s previous-generation mid-weight representative |
| qwen3.5:27b-mlx-bf16 | Dense 27B (bf16) | 54 GB | An extreme case testing the practical limits of 64GB |
| qwen3.6:27b | Dense 27B | 17 GB | Dense counterpart to qwen3.6:35b (same-generation Dense vs. MoE) |
| qwen3.6:35b | MoE 35B/3B-active | 23 GB | The MoE favorite |
I crossed these 10 models with think True/False (models without a think mechanism got only False), and additionally measured qwen3.6:35b via the mlx-lm direct path as well, for a total of 16 sets. Prompts covered confidentiality judgment, code, summarization, QA, and so on — 27 questions in all.
Let me also note what I excluded from evaluation. One was Llama 4 Scout (67 GB, MoE 108B/17B-active). By the preliminary-experiment stage, it was already clear that memory would break down on a 64GB machine and it wouldn’t be practical, so I excluded it from the main test (if I get access to a 128GB-class machine in the future, there’s room to re-evaluate). The other was the Qwen 3.5 line via the mlx-lm direct path — mlx-lm doesn’t support that model format, and forcing it through would mean mixing in a different library, which would break the purity of the comparison. So the mlx-lm-direct verification was narrowed to just the latest Qwen 3.6 line.
27 prompts × 3 trials × 16 sets = 1,296 inferences. On a MacBook Pro M4 Max (64GB, macOS 26.3), it ran to completion in 10 hours 56 minutes with zero anomalies.
(For reference, the preliminary experiment — partly because it ran into the think problem mentioned above — was a brutal test that had the Mac running at full power for over 24 hours straight.)
The single most important thing in a benchmark is holding conditions constant — and, I realized after running this, even more important than that is deciding up front “who is this test for.” Test design is subordinate to the nature of the task. There’s no such thing as a universal, general-purpose benchmark — that was the lesson at the very starting point of this project.
Chapter 2: MoE Was a Strict Upgrade Over Dense
The first thing that jumped out from the results was how large the effect of architecture was. Lining up a Dense model and an MoE model from the same vendor, the gap was stark (think=False, median of the Axis 2 speed prompts). (The “active parameters” column in the table below refers to the part that actually runs during a single inference — more on this below.)
| Model | Architecture | Active parameters | Response time | Generation speed | Peak memory | Axis 1 F1 |
|---|---|---|---|---|---|---|
| gemma4:31b | Dense 31B | 31B | 13.38 s | 16.3 tok/s | 33.59 GB | 1.0 |
| gemma4:26b | MoE 26B/4B-active | 4B | 3.20 s | 76.3 tok/s | 21.93 GB | 1.0 |
| qwen3.6:27b | Dense 27B | 27B | 14.68 s | 14.3 tok/s | 33.24 GB | 1.0 |
| qwen3.6:35b | MoE 35B/3B-active | 3B | 4.81 s | 44.7 tok/s | 31.23 GB | 1.0 |
Google’s MoE (gemma4:26b) was 4.2x faster in response time and 4.7x faster in generation speed than its same-family Dense counterpart (gemma4:31b). Alibaba’s MoE (qwen3.6:35b) was also about 3x faster than its same-generation Dense counterpart (qwen3.6:27b). And all four models hit F1 = 1.0 on confidentiality judgment — no accuracy degradation whatsoever. There was no observed “cost of getting faster” — it was a strict upgrade, with no tradeoff.
I think the underlying reason is simply the nature of MoE itself. MoE (Mixture of Experts) has, out of its total parameters (26–35B), only a portion of experts (3–4B) actually active for any given token. What determines inference cost is the active parameters, not the total parameters. So it seems able to deliver “26–35B-class knowledge capacity” while running at “4B-class inference speed.” This matches prior findings (measurements showing a 30B MoE running at roughly 3B speed on M4/M5, and the observation that MoE saves compute but not memory). I think this article’s original contribution is showing that, within a single unified benchmark, with numbers directly comparing four models head to head.
The Gains in Speed and the Gains in Memory Are Separate
Let me break this down a bit further. The key to understanding this is that, in MoE, “what determines speed” and “what determines memory” are two separate things.
Picture academic peer review. A Dense model is like a conference where, to review one paper, “every expert across the entire relevant field reads through it, one by one.” With all 31 reviewers (= 31B) involved, the conclusion is solid but it takes time.
An MoE model is a conference where “only the handful of people whose specialty matches the paper’s topic review it.” There are 26–35 reviewers total on the roster, but only 3–4 of them (= active parameters) actually do the work on any given paper. So each paper gets reviewed fast. That’s the true source of the speed gain.
But here’s the thing: all the remaining experts who aren’t involved in this particular review are still members of the conference. The conference’s upkeep cost (= memory) is incurred for every member on the roster, regardless of whether they’re actually working. So “review is fast” and “the conference is cheap to maintain” turn out to be two separate stories. What determines memory is “the total number of members (total parameters)”; what determines speed is “the handful who actually work on one paper (active parameters)” — the two don’t move together.
Looking at it this way, the difference between the two MoE models becomes clear. The key point is how much memory each model actually occupied (the measured peak memory).
- Google’s gemma4:26b (MoE, occupying 21.9 GB) came in about 12 GB lower than its same-family gemma4:31b (Dense, occupying 33.6 GB). Since its total membership (total parameters, 26B) is smaller than the Dense model’s (31B), it gets both the speed win and the memory win.
- Alibaba’s qwen3.6:35b (MoE, occupying 31.2 GB) is only about 2 GB lower than its same-generation qwen3.6:27b (Dense, occupying 33.2 GB). Because you need to keep all 35B members loaded in memory, the review (speed) is 3x faster, but the conference’s upkeep cost (memory) is almost unchanged.
In other words, “MoE = light” is only half true. More precisely: “MoE = fast. Whether it’s also light depends on the total parameter count.” The scatter plot below makes this separation plain. The two MoE models both dominate the “fast band” (over 40 tok/s), but gemma4:26b sits on the left (light) while qwen3.6:35b sits on the right (heavy).

To sum up: if the accuracy is equal, choosing MoE is the smart move. But a caveat — don’t estimate memory footprint from the total-parameter number on a spec sheet alone. Speed tracks active parameters; memory tracks total parameters — you have to look at them separately. As an aside, the OpenAI Privacy Filter mentioned at the start also appears to use this same MoE structure: of its 1.5B total parameters, only about 50M actually run during inference. It seems OpenAI landed on the same design — “keep the knowledge capacity, but narrow what actually runs, to be fast and light” — for the same purpose, confidentiality detection.
Chapter 3: For Convergent Tasks, Thinking Is a Tax
Going in, my assumption was “surely making it think more deeply makes it smarter.” In practice, that’s not what happened. This is, however, strictly a result for the case “narrowly limited to convergent tasks.”
A Concrete Case Where Think Broke the Answer
After finishing the benchmark test, since I now had a local LLM installed and running anyway, I asked gemma4:26b to “list Japan’s prime ministers in order.” Running it with Ollama’s default (thinking is on by default for models with a thinking mechanism), the thinking portion looped endlessly through hesitations like “should I write them all out or break it up by era, it’s easy to miscount past 100,” and even after it got into the main text it repeated several names hundreds of times, ballooning the count of prime ministers to 732. Having judged this to be a think-driven runaway, I set “/set nothink” and asked the same question again. The loop stopped completely, and it terminated at number 89. This was the moment a runaway of the same shape as the thinking infinite loop that Qwen 3.5 models produced in this experiment (discussed below) reproduced itself on a model from an entirely different vendor.
(As of when I’m writing this, late May 2026, we’re on the 105th Prime Minister, Takaichi — but the model’s knowledge was from the Ishiba administration era, so it listed up through Ishiba. Also, a task that requires “listing a large number of precise facts” has inherent limits even with thinking turned off, on the generative model alone — that needs RAG, discussed later. What matters here isn’t the accuracy itself, but the behavior that thinking triggers a runaway loop. The figure of 89 prime ministers isn’t accurate either — that’s a limitation of the LLM model’s closed knowledge itself.)
This is an extreme accident case, but it’s the clearest illustration of thinking working against you on a convergent task.
Axis 1 (Classification): Thinking Is Slower, and Accuracy Is “No Difference” or “Worse”
| Model | think=False | think=True | Time taken | F1 (False → True) |
|---|---|---|---|---|
| gemma4:26b | 3.20 s | 12.23 s | 3.8x | 1.0 → 0.889 |
| gemma4:31b | 13.38 s | 54.63 s | 4.1x | 1.0 → 1.0 |
| qwen3.6:35b | 4.81 s | 41.37 s | 8.6x | 1.0 → (6 empty responses) |
| qwen3.6:27b | 14.68 s | 97.17 s | 6.6x | 1.0 → 0.75 |
On confidentiality judgment, thinking ate 4–9x the time and accuracy either didn’t change or got worse. gemma4:26b’s Recall dropped (more misses); qwen3.6:27b’s F1 collapsed to 0.75. qwen3.6:35b even had 6 cases where the thinking ate up the entire output-token cap, leaving the body empty.
Axis 3 (Generation Quality): “Maybe It Helps Quality” Was Also Refuted
Even if it backfires on classification, there was still a possibility that thinking would help quality on more open-ended tasks like summarization or code generation. I read through the response bodies to check, and found no observed benefit to quality — if anything, degradation showed up.
- Cut off mid-way: qwen3.6:35b’s (think=True) summary spent 7,515 characters on thinking, and the body ended up cutting off mid-sentence at “…next time,” in a state where it’s unclear what it was even trying to say. This is the thinking eating up num_predict and the body getting pushed past the token cap.
- Factual drift: qwen3.6:27b (think=True) took phrases from the original test material like “planning/considering hiring” and output them as “decided to hire,” and “will be discussed at the board meeting” as “approved,” pulling unconfirmed nuance toward a definitive statement — in other words, the content got rewritten. The think=False version showed no such drift.
- Wildly disproportionate cost: a simple summarization task took 90–170 seconds, with 5,000–7,500 characters of thinking generated. The resulting body text was equal to or worse than the alternative.
Why Does Thinking Backfire on Convergent Tasks?
Confidentiality judgment and summarization are tasks where “the correct answer converges to one point.” Stretching out the thinking here (a) turns extra speculation into definitive statements, causing factual drift, and (b) lets the thinking eat the output budget, cutting off the body. “Thinking” didn’t aid convergence — it ended up destabilizing the landing point instead.
I believe this observation independently aligns with Apple’s research, “Reasoning’s Razor” (arXiv 2510.21049, 2025). That study systematically showed that, for classification tasks like safety detection and hallucination detection, reasoning raises average accuracy while dropping recall in the low-FPR region that matters most in practice — with no-reasoning dominating there. My own confidentiality judgment (a convergent classification task) reproduced exactly the same Think-Off advantage. I’d position this not as a rehash but as an independent replication.
That said, this isn’t a claim that “thinking has no value.” For tasks where the answer doesn’t converge — reflecting on philosophical questions, hard problems requiring multi-step reasoning, divergent exploration or creative work — the process of thinking itself is likely to generate value. There, what this article calls a “tax” becomes an “investment.” This benchmark doesn’t measure those, so it neither affirms nor denies anything about them. What I can say here is: “for convergent tasks like confidentiality judgment, thinking was a tax.”
To sum up: for convergent tasks like classification, summarization, and QA, it looks like the right move is to default thinking off, and raise accuracy through prompt design and context injection instead. Conversely, for tasks where introspection or multi-step reasoning is the whole point, don’t reuse this conclusion — run a separate test to check.
Chapter 4: Choose Architecture, Not Generation
Let me correct the naive expectation that “the new generation is faster across the board” against the actual measurements. I lined up Qwen 3.5 and 3.6’s Dense 27B (think=False).
| Model | Response time | Generation speed | Peak memory | Axis 1 F1 |
|---|---|---|---|---|
| qwen3.5:27b | 15.15 s | 14.3 tok/s | 33.24 GB | 1.0 |
| qwen3.6:27b | 14.68 s | 14.3 tok/s | 33.24 GB | 1.0 |
Speed, memory, and accuracy are all within margin of error. The generation update barely changed the Dense line at all. This is an important fact — a lesson that you shouldn’t take the “generation” banner at face value as “across-the-board improved performance.”
So where was 3.6’s actual progress? I think it comes down to two things:
- Thinking isn’t broken anymore. The Qwen 3.5 line’s thinking structurally loops infinitely (the Issue discussed below). The 3.6 line fixed this. If you want to use reasoning mode, upgrading the generation is the only real fix — a difference with actual practical impact.
- A usable MoE option was added. Where 3.5 was Dense-only, 3.6 added an MoE (35B/3B-active). This achieved 4.81 s versus its same-generation Dense’s 14.68 s — 3x faster at the same accuracy.
So I’d say the generational progress wasn’t “Dense got faster” — it was “a usable MoE option got added.”
To sum up: don’t take the “generation” banner at face value. The value of upgrading generations comes down to “thinking got fixed” and “MoE got added.” If you’re only using Dense, the generational difference turns out to be small.
Chapter 5: Path, Quantization, and Reproducibility — Small Landmines in Implementation
Three details you’ll actually trip over in practice. This is an area where prior articles have almost no first-hand measurements.
mlx-lm vs. Ollama: Changing the Path Doesn’t Make It Faster
For this benchmark, I ran each LLM model through Ollama. There are also LLM models that support MLX, which is built for Apple Silicon. Since I’m using a Mac anyway, I wanted to enjoy the benefit that comes with it — MLX — and expected that going through mlx-lm directly would be faster on Apple Silicon. I ran the identical Qwen3.6-35B-A3B via Q4 (Ollama) versus 4bit (mlx-lm direct), comparing only the path, and mlx-lm direct showed no clear speed advantage — if anything it was slower (about 2.9x in response time). On top of that, controlling think didn’t work well through the mlx-lm direct path. The conclusion is simple: for this use case, sticking with Ollama alone is enough. Unfortunately, there’s no reason to pick MLX here.
Running bf16 Through Ollama Is a Trap
Running qwen3.5:27b-mlx-bf16 (bf16, 54GB) through Ollama pushed things to the edge — 10.0 GB of swap, free memory bottoming out at 1.7 GB — and it came out as the slowest of every model tested, at 7.65 tok/s and 22.77 s response time. On my 64GB-memory machine, it became clear that bf16 27B sits outside the practical limit. Even if you want to use bf16 weights, at 64GB it just doesn’t run reasonably — you need to use an appropriately quantized version instead.
Even at temperature=0, Output Still Wobbles — Though It Depends on the Vendor
I’d assumed temperature=0.0 would be deterministic. Reality turned out otherwise.
First, let me explain what temperature does. When an LLM picks the next word, it holds a “likelihood (probability)” for each candidate. For “The cat is ___,” it might be cute 40%, an animal 25%, sleeping 15%, and so on. Temperature is the dial that controls how sharp or how flat this probability distribution is.
- Turn it up (e.g., to 1.0) and the distribution flattens, so lower-probability words get picked more often → different output every time, more creative.
- Set it to 0 and the distribution goes sharply peaked, always picking the single highest-probability word → should be the same every time.
So temperature=0 is the setting “stop rolling the dice, always take the top candidate” — in theory, no matter how many times you run it, you should get the same answer back. This is called “deterministic.”
So why did it wobble? Because there’s a source of wobble somewhere other than the dice. Mainly two things:
- The tie problem: when the top candidates are nearly even, like “40% versus 39.99%,” which one gets treated as #1 can flip on the basis of a tiny computational rounding error. Once the first step changes, the entire rest of the text branches down a completely different path.
- The order-of-addition problem: a computer’s floating-point arithmetic can shift in the last digit depending on the order you add the same numbers in. GPUs add things up in parallel, in scrambled order, so the final digit wobbles slightly from run to run — and if that wobble happens to land on a “tie problem,” it turns into a branch in the output.
In short: “you stopped rolling the dice, but a tiny wobble remains in the arithmetic just before that, and every so often it flips the result.” And how prone to that wobble a model is turned out to differ by vendor — that was this test’s discovery.
I checked this directly. I ran each model three times under identical conditions and compared the length of the returned text. If it were perfectly deterministic, all three runs should come out the same length. Instead, cases where the gap between the longest and shortest exceeded 30 characters (or 5%) — that is, cases where “the answer wobbled under identical conditions” — came up 11 times across the board. And how that wobble broke down varied clearly by vendor.
| Family | Number of wobbling runs (conditions) |
|---|---|
| gemma4:26b | 5 |
| gemma4:31b | 2 |
| gemma3:12b | 2 |
| phi4:latest | 2 |
| Qwen family (all 7 sets) | 0 |
Qwen is bit-stable at temp=0; Gemma/Phi wobble. Ironically, gemma4:26b — the leading candidate for the confidentiality checker — turned out to wobble the most. If your pipeline relies on exact output matching as a premise for caching keys or diffing, this is worth watching out for.
To sum up: the Ollama inference path is sufficient. There’s no need to force bf16 onto 64GB. And if bit-level reproducibility is a requirement, Qwen is the better choice.
Chapter 6: How to Choose in a World Where Accuracy Has Saturated
Three sets of results and observations are now on the table. Let me pull them together and land the conclusion.
On Axis 1, F1 = 1.0 for confidentiality judgment showed up across multiple models. On Axis 3 too, quality on straightforward code, summarization, and QA saturated at a practically-passing level, with small differences between models. At the 9–35B class, accuracy and quality on business tasks at the level of confidentiality judgment no longer differ meaningfully between models. I take this to be where local LLMs currently stand.
Since accuracy no longer differentiates models, the axis for choosing shifts from “how smart” to “how easy to run” — speed, memory, reproducibility. The decision map looks roughly like this:
- Accept that accuracy has already saturated (for this class and this kind of task).
- Choose MoE (same accuracy, faster).
- Decide your speed/memory tier according to use case (if you need it light, go with a small-total-parameter MoE or a lightweight Dense model).
- If bit-level reproducibility is a requirement, use the Qwen family.
- For convergent tasks, default thinking off.
And back to the question at the start. The dilemma from Chapter 0 — a mechanism to check for confidential information without ever sending it to the cloud — does work. On 64GB Apple Silicon, a local LLM can handle a business task at the level of confidentiality judgment fully practically. What settles it isn’t how smart the model is, but the architecture (MoE) and the operational setup (thinking off, path choice, reproducibility).
Chapter 7: Limits, and What Comes Next
The Limits of Scope (Most Important — Echoes the Chapter 1 Scope Declaration)
Every conclusion in this article is about the convergent task of judging whether confidential information is present. None of it transfers to use cases where “the process of thinking itself generates value” — introspection, multi-step reasoning, creative work, divergent exploration. Those have fundamentally different evaluation axes, prompts, and judgment criteria, and I’d expect a different test to produce a different result. In particular, I think the evaluation of thinking could well reverse in that context. To repeat: this is not “a conclusion about local LLMs in general,” but “a conclusion about whether a local confidentiality checker is viable.”
The Limits of the Setup
The prompts were convergent and relatively straightforward, quantization was mostly Q4, the hardware was a single M4 Max 64GB configuration, and the models are all from a specific point in time. I withhold generalizing to harder problems, other quantizations, or other hardware. Where this aligns with prior research (the academic finding that reasoning can backfire on classification, MoE’s memory characteristics), I treat that as independent verification; where it might not align, I leave that as future work.
What Comes Next
This knowledge, naturally, applies directly to choosing a model for an actual local confidentiality checker. It looks like the right setup would be gemma4:26b as the primary (fast, light, F1 = 1.0), with the lightweight tier (gemma3:12b / phi4) as parallel candidates, thinking defaulted off, and accuracy topped up via externally-injected context hints. And for situations that require “listing a large number of precise facts” — like the prime-minister example above — even with thinking off, the generative model alone still has inherent limits, so that needs to be solved with RAG (a mechanism that references external, correct information) after handing it correct data to reformat. I’ll leave that implementation to a separate article.
I think a locally-complete confidentiality check has moved past the stage of asking “does it work at all.” It’s now in the stage of “how do you choose it wisely, and how do you operate it.”
Appendix: Information for Reproduction
- Hardware: MacBook Pro M4 Max, 64GB unified memory, macOS 26.3.
- Inference path: Ollama (llama.cpp Metal backend). Quantization is Q4_K_M unless otherwise noted.
- Held-constant conditions: temperature=0.0 / num_predict=4000 / keep_alive=0 / 3 trials per condition, median taken. think was explicitly specified True/False via the API (since some models default it on).
Prompts:
- Axis 1 (confidentiality judgment) is a convergent classification task: “given an input document, produce structured output for ‘confidentiality level: confidential / safe / gray’ and ‘confidential elements detected: …’”. The specific judgment instructions are the confidentiality checker’s core know-how, so only the structure is shown in this article.
- Axis 3 (general quality) uses common tasks: single-function code generation (e.g., deduplication), summarizing a routine meeting record, and conceptual QA (e.g., explaining decorators).
Prior research: Apple’s “Reasoning’s Razor” (arXiv 2510.21049, 2025), OptimalThinkingBench (2025), TextReasoningBench (2026), and various measurements of MoE speed/memory on Apple Silicon. All citations are paraphrases of the gist; no original text is reproduced.
(The figures in this article are individual, first-hand measurements taken at a specific point in time, in a specific configuration. Your own environment and use case may produce different results.)
About Soul Resonant Works
Soul Resonant Works is a solo venture developing seven local AI systems. Starting from zero programming experience, the development is progressing through collaboration with AI.
🌐 Soul Resonant Works:
→ https://sr-works.net/en/index.html
📝 This blog publishes the entire development process as a serialized journal.
CubePlot (free version available)
CubePlot is the first product from Soul Resonant Works — it turns a CSV into a 3D scatter plot you can rotate and zoom. No install, no sign-up: it’s a single HTML file you open in your browser, and it works offline. Your CSV never leaves your machine — the app itself sends no data anywhere and blocks network traffic at the browser level (CSP). Start with the free version.
▶ Product page: https://sr-works.net/en/cubeplot/
▶ Get it (commercial license, USD $39 + tax where applicable): https://soulworks8.gumroad.com/l/pzcij
If you found this article useful, please share it.
Leave a Reply