Think of three cars sharing one lane: the deployment that scored 78.4% was not the first to finish. How should we pick a daily deployment on 16 GB?
On 2026-09-28, one Windows workstation and one NVIDIA RTX 4060 Ti 16GB served three deployments that each fit in GPU memory. They took turns on the same frozen list of 700 items. Ridge got 549 right, while OrcaSAQ reached 543 and Ternary finished with 522. Wall clock ran the other way. Ternary took 34.01 minutes, Ridge 47.55, OrcaSAQ 62.59. The accuracy ranking and the arrival-time ranking are not the same line.
State the contract first
With only one 16 GB card, the accuracy winner and the speed winner are often different deployments.
This note answers a narrow contract. Seven public benchmarks, 100 frozen items each, and all three arms saw the same list. The judge extracts a choice letter, or the final number on a numeric item. It does not score prose, and it does not score tool use. A 2026-09-19 note asked whether short tasks, a long-context needle, and tool loops were enough to point a daily alias at Ternary. That is a different question. This run does not rewrite that conclusion, and these 700 items are not the whole of daily work.
The three runtimes took turns on one card. They cannot all sit in 16 GB at once. Fitting one weight file is not the same as co-residency. The item hashes matched across arms, so the gap is the deployment, not a different draw of questions.
These are three deployments
All three fit in 16 GB, but the quant format, runtime, and packaging are not the same.
The shared upstream is Qwen3.8-27B. The model card lists Apache 2.0 and a context length of 262,144.[1] This run did not serve the official unquantized checkpoint. It served three instances that can already run on a 16 GB card.
The Ridge arm was served as Qwen3.8_Ridge through Ollama. The public Empero Ridge 3.7 bpw file is the registered counterpart of that name: a mixed GGUF, about 11.73 GiB (12.59 GB). The card calls 16 GB a practical starting point. That is their hardware guidance, not a VRAM measurement on this card.[2] This run did not recompute a weight hash, so it cannot claim byte identity with that file.
The OrcaSAQ arm is OrcaSAQ-2-27B, EXL3, served with vLLM. The card lists about 12.3 GB, 3.21 bpw, thinking on by default, and a suggested temperature of 1.0. The SWE-bench Pro score of 70.0 on that card is their vLLM agent eval, not these 700 items.[3]
The Ternary arm used the file Ternary-Bonsai-2-27B-PQ2_0-MTP-Q8_0.gguf under Prism's llama.cpp MTP runtime. That filename matches the ProCreations MTP file. They describe an experimental MTP head, not an official Prism release, about 7.658 GB, and they say stock llama.cpp is not enough.[5] Prism's own PQ2_0 is a 7.21 GB, 2.13 bit/weight ternary pack. The card says stock llama.cpp will refuse the file or emit garbage, recommends temperature 1.0 in thinking mode, and reports 84.78 as the average of 14 thinking-mode benchmarks.[4] That 84.78 is another exam and another sampler. We do not subtract 74.6% from 84.78. This run also did not check the weight hash.
Packing a coat into a carry-on does not measure the keys in the pocket. A smaller file does not automatically keep the same score on these 700 items. The three kinds of "small" are not the same small. Ridge is a mixed quant that treats hybrid-architecture tensors differently. OrcaSAQ is EXL3. Ternary is ternary main weights plus an MTP head. These are not three labels on one weight file.
The method only locks the objective answer
Temperature was 0, thinking was off, each item was capped at 512 tokens, and the judge only extracted the final answer.
The protocol was temperature 0, max_tokens 512, thinking disabled, judge name objective-v1. Each item was a separate request. Wall clock is that HTTP round trip, including prompt processing and generation, summed across 700 sequential items. It is not llama-bench tg128, and it is not concurrent throughput. Context allocation was not recorded. Each arm has 700 result rows, and the summary request-error count is 0. This article recounts the passed flag in those files. It does not reimplement the judge.
This protocol disagrees with the vendor recommendations, and that disagreement has to sit in front of the scores. Prism's published numbers use thinking mode at temperature 1.0.[4] The OrcaSAQ card also defaults to thinking at temperature 1.0.[3] We turned thinking off and used temperature 0 so the three arms could be compared on extraction, not so we could reproduce their headline numbers. If you compare this 74.6% with 84.78, you are comparing two exams, not two deployments.
The 512-token cap cuts long answers. Ridge hit the cap on 32 items, OrcaSAQ on 26, Ternary on 92. Most capped items were marked wrong. That cap is entangled with the accuracy gap, and the suite split below separates the cases we can actually see. Empty predictions were rare: 4 for Ridge, 6 for OrcaSAQ, 8 for Ternary, all failed. Those are extraction misses, not transport errors.
The seven names match public benchmarks: MMLU, MMLU-Pro, CMMLU, GSM8K, ARC-Challenge, HellaSwag, and MathQA.[6][7][8][9][10][11][12] We used a frozen 100-item slice of each, not the full official test set. The percentages in the table do not belong on those papers' leaderboards.
The accuracy gap is 27 items
The totals are Ridge 549, OrcaSAQ 543, and Ternary 522, with no request errors.
| Suite | Ridge | OrcaSAQ | Ternary |
|---|---|---|---|
| MMLU | 80 | 83 | 77 |
| MMLU-Pro | 63 | 68 | 48 |
| CMMLU | 78 | 77 | 72 |
| GSM8K | 85 | 87 | 85 |
| ARC-Challenge | 96 | 94 | 92 |
| HellaSwag | 92 | 91 | 86 |
| MathQA | 55 | 43 | 62 |
| Total | 549 | 543 | 522 |
As percentages, that is 78.4%, 77.6%, and 74.6%. Against Ridge, OrcaSAQ is short 6 items, about 0.86 percentage points. Ternary is short 27 items, about 3.86 percentage points. A 6-item total gap is thin when each suite has only 100 items. Twenty-seven items is thicker, but it is not a thin loss spread across all seven suites. It concentrates.

MMLU-Pro is the hole. Ternary scored 48, Ridge 63, OrcaSAQ 68. That is 15 items behind Ridge and 20 behind OrcaSAQ. Those 15 items cannot be booked entirely as "the model is worse." On this suite Ternary hit the 512-token cap on 37 of 100 items, and 34 of those capped items failed. Ridge averaged 2 completion tokens here and hit the cap 0 times. OrcaSAQ averaged 27.5 completion tokens and hit the cap 5 times. If a long explanation spends the budget before the letter is extracted, the judge records a miss. This run did not repeat the suite with a higher cap, so it cannot estimate how many of those 34 items would have passed if the answer had finished. What it can say is that the accuracy gap on this suite is entangled with output length and with the 512 ceiling.
MathQA runs the other way. Ternary scored 62, Ridge 55, OrcaSAQ 43. Ternary is 7 ahead of Ridge and 19 ahead of OrcaSAQ. Ternary hit the cap on 41 MathQA items: 34 failed, 7 still passed. Ridge also hit the cap on 20 items. Both were cut, Ternary more so, and Ternary still scored higher. OrcaSAQ averaged only 109 completion tokens on MathQA and hit the cap 16 times. It wrote shorter answers and got more of them wrong. That gap looks less like a 512 artifact.
The other suites are closer. GSM8K is 85, 87, 85: Ternary ties Ridge and trails OrcaSAQ by 2. ARC-Challenge is 96, 94, 92, all at or above 92. HellaSwag is 92, 91, 86: Ternary is short 6, not collapsed. MMLU is 80, 83, 77. CMMLU is 78, 77, 72. The total makes Ternary look uniformly thinner. The suites do not.
Read wall clock, not only tok/s
Ternary finished 700 items in 34.01 minutes, Ridge in 47.55, OrcaSAQ in 62.59.
Why can't you read only tok/s?
Completion tokens divided by wall seconds give 46.74 tok/s for Ternary, 17.45 for Ridge, and 10.85 for OrcaSAQ. That looks like 2.68x and 4.31x. Wall clock is only 1.40x and 1.84x: Ridge took 1.40 times Ternary's time, OrcaSAQ 1.84 times. Ternary used 28.5% less wall time than Ridge and 45.7% less than OrcaSAQ. A speedometer is not an arrival time. tok/s measures how fast tokens leave. Wall clock measures when these 700 items were done.
The reason is volume. Ternary wrote 95,395 completion tokens, mean 136.3. Ridge wrote 49,792, mean 71.1. OrcaSAQ wrote 40,742, mean 58.2. More tokens enlarge the tok/s numerator. Extra text piles onto the clock, but not evenly. On short multiple-choice items, Ridge and OrcaSAQ often emitted 2 tokens, one letter. Ternary averaged 49.3 tokens on MMLU and 220 on MMLU-Pro. One arm hands in a letter. The other hands in a paragraph. The rate table calls the paragraph fast. Arrival time does not have to agree.

| Suite | Ridge | OrcaSAQ | Ternary |
|---|---|---|---|
| MMLU | 2.13 | 2.94 | 3.35 |
| MMLU-Pro | 1.43 | 8.35 | 7.54 |
| CMMLU | 3.27 | 1.07 | 1.87 |
| GSM8K | 21.58 | 32.66 | 7.32 |
| ARC-Challenge | 0.98 | 0.73 | 0.85 |
| HellaSwag | 1.27 | 1.28 | 1.20 |
| MathQA | 16.89 | 15.57 | 11.88 |
| Total | 47.55 | 62.59 | 34.01 |
Totals are raw seconds summed, then converted to minutes. If you round each row first and then add, OrcaSAQ is off by 0.01 minute. Use the total row.
The overall speed win is not spread evenly. On GSM8K the mean completion lengths are close: 255, 261, and 219 tokens. Wall clocks are 21.58, 32.66, and 7.32 minutes. Ridge took 2.95 times Ternary's time on this suite. That gap looks like decoding, not like Ternary writing less. On MathQA, Ternary averaged 409 tokens, near the 512 ceiling, and still finished in 11.88 minutes, faster than Ridge at 16.89 and OrcaSAQ at 15.57, with more items correct. Short multiple choice is a different story. ARC-Challenge and HellaSwag all finished within 1.3 minutes. CMMLU was fastest on OrcaSAQ, at 1.07 minutes. MMLU-Pro was fastest on Ridge, at 1.43 minutes. Ternary spent 7.54 minutes there, slower than Ridge, and 15 items less accurate.
So "Ternary is faster" holds on the sum of 700 items. It is faster on the long GSM8K and MathQA blocks. It is not faster on every short suite, and it is not both faster and more accurate on the hardest multiple-choice suite.

The suite split changes the pick
The arm that trails the total leads on MathQA and drops the most on MMLU-Pro.
If the work looks like the frozen GSM8K slice, Ternary uses about one third of the wall clock, matches Ridge on correct items, and trails OrcaSAQ by 2. If it looks like MathQA, Ternary is faster and more accurate, but 41 items hit the 512 cap and 34 of those failed. Leaving the cap at 512 means accepting those cut failures. If the work looks like MMLU-Pro, where the choices are tighter and the contract is a letter, Ridge spends 1.43 minutes to get 63, OrcaSAQ spends 8.35 minutes to get 68, and Ternary spends 7.54 minutes to get 48. Speed and accuracy do not both land on Ternary there.
OrcaSAQ's total is only 6 items behind Ridge, but MathQA is 12 behind and MMLU-Pro is 5 ahead. It is not a slower copy of Ridge. Close totals still move when you open the suites.
When to use which
If the contract is objective extraction, pick Ridge. If the contract is wall clock on the same items, pick Ternary.
If the contract is "letters and numbers on this frozen list," Ridge's 549 is the high point of these three arms. OrcaSAQ is only 6 behind, and the totals nearly stick together unless MathQA is a large share of the work. If the contract is "same 4060 Ti, same 700 items, thinking off, when does it finish," Ternary's 34.01 minutes is the high point. The cost is 27 items, concentrated on MMLU-Pro, and entangled with the 512 ceiling.
What not to do with these numbers is just as concrete. Do not use them to judge writing taste. Do not use them to decide whether a tool loop connects. Do not load all three runtimes into 16 GB at once. Do not subtract a vendor thinking-mode average from this temperature-0 extraction rate. Do not write a 100-item slice as full MMLU or full GSM8K. And do not treat a 2.68x tok/s ratio as a 2.68x speedup on every item.
Where these numbers do not travel
This is not writing taste, not a tool loop, and not a vendor thinking-mode leaderboard.
The scope is one Windows workstation, one RTX 4060 Ti 16GB, and three time-shared servings on 2026-09-28. Each suite has 100 items, not the full benchmark, and this article does not add a confidence interval. A 6-item total gap is thin. Twenty-seven items is thicker, but 15 of the MMLU-Pro gap is entangled with the output cap and cannot be booked as pure capability. Ridge and Ternary weight files were not re-hashed in this run. The judge is the passed flag already in the result files, not a second extractor. Context length is not in the protocol fields that can be stated here. Temperature 0, thinking off, and 512 tokens are this deployment contract, not the sampling advice on the three model cards.
Inside that contract, the order is stable: the item recount matches the file headers, the item hashes match, and the request-error count is 0. The accuracy winner is Ridge. The speed winner is Ternary. They are not the same deployment.