IQ3_S MTP
It produced the highest balanced score, nearly halved the adaptive long-task time versus IQ3_S, and led the Researcher and Coder role scores. The tradeoff is memory: only 1,673 MiB remained free in the representative run.
A base-versus-MTP comparison of every GGUF in ISTA-DASLab’s Qwen3.8-27B GSQ-RCO release—measured for speed, memory, long context, coding, research, knowledge work, tools, and C-Suite output.
The fastest file was not automatically the best broad-use model. Quality was weighted around the actual jobs this local stack needs to perform; HumanEval+ contributed only 10% of the balanced index.
It produced the highest balanced score, nearly halved the adaptive long-task time versus IQ3_S, and led the Researcher and Coder role scores. The tradeoff is memory: only 1,673 MiB remained free in the representative run.
The speed-first choice, with a 32.7% elapsed-time reduction over its base pair.
The highest base-model balanced score and the top C-Suite role score.
The 96-second explainer covers the experimental setup, $2.95 cost, MTP speed and memory tradeoffs, quality results, 64K context, and role-specific recommendations.
One universal winner is less useful than a routing policy. These scores reweight the same evidence for five distinct jobs in a private AI staff.
Switch the chart between quality, coding, elapsed time, and remaining memory. Lower is better for time; higher is better for the other views.
| Model | Balanced | HE+ pass@1 | Adaptive time | Peak VRAM | Free VRAM | Selected profile |
|---|---|---|---|---|---|---|
| IQ3_S MTP WINNERQwen3.8-27B GSQ-RCO | 83.19 | 89.63% | 53.49s | 14,176 MiB | 1,673 MiB | MTP 3 · ubatch 256 |
| IQ3_SQwen3.8-27B GSQ-RCO | 81.34 | 87.80% | 105.79s | 13,286 MiB | 2,563 MiB | MTP 0 · ubatch 1024 |
| IQ3_XXSQwen3.8-27B GSQ-RCO | 81.25 | 90.24% | 98.51s | 11,692 MiB | 4,159 MiB | MTP 0 · ubatch 1024 |
| IQ2_XS MTPQwen3.8-27B GSQ-RCO | 80.80 | 90.24% | 48.29s | 11,270 MiB | 4,581 MiB | MTP 3 · ubatch 512 |
| IQ2_S MTPQwen3.8-27B GSQ-RCO | 80.66 | 85.98% | 49.62s | 12,120 MiB | 3,731 MiB | MTP 2 · ubatch 1024 |
| IQ2_SQwen3.8-27B GSQ-RCO | 80.04 | 84.76% | 65.45s | 10,896 MiB | 4,955 MiB | MTP 0 · ubatch 1024 |
| IQ2_XSQwen3.8-27B GSQ-RCO | 79.12 | 85.98% | 71.73s | 10,222 MiB | 5,629 MiB | MTP 0 · ubatch 1024 |
| IQ3_XXS MTPQwen3.8-27B GSQ-RCO | 78.68 | 87.80% | 49.76s | 13,066 MiB | 2,785 MiB | MTP 3 · ubatch 1024 |
MTP delivered meaningful acceleration in every matched pair, but it always consumed more VRAM and did not guarantee better broad-use quality.
The applied suite used 24 synthetic tasks per model across six practical categories. This prevented 164 coding problems from dominating a model intended for broad research, operations, and knowledge work.
| Model | Business | Research | Knowledge | Agent tools | Sysadmin | Executive content |
|---|---|---|---|---|---|---|
| IQ2_XS | 47.92 | 95.83 | 96.88 | 80.00 | 78.75 | 64.58 |
| IQ2_XS MTP | 56.25 | 95.83 | 96.88 | 80.00 | 78.75 | 60.42 |
| IQ2_S | 56.25 | 95.83 | 100.00 | 71.67 | 78.75 | 64.58 |
| IQ2_S MTP | 38.54 | 100.00 | 100.00 | 95.83 | 78.75 | 60.42 |
| IQ3_XXS | 57.29 | 92.71 | 96.88 | 75.83 | 83.75 | 72.92 |
| IQ3_XXS MTP | 51.04 | 92.71 | 96.88 | 75.83 | 78.75 | 64.58 |
| IQ3_S | 59.38 | 100.00 | 96.88 | 80.00 | 73.75 | 56.25 |
| IQ3_S MTP | 56.25 | 100.00 | 96.88 | 90.83 | 78.75 | 60.42 |
| Model | Evaluated | HumanEval passed | HumanEval+ passed | HE+ pass@1 |
|---|---|---|---|---|
| IQ2_XS | 164 | 150 | 141 | 85.98% |
| IQ2_XS MTP | 164 | 155 | 148 | 90.24% |
| IQ2_S | 164 | 145 | 139 | 84.76% |
| IQ2_S MTP | 164 | 145 | 141 | 85.98% |
| IQ3_XXS | 164 | 156 | 148 | 90.24% |
| IQ3_XXS MTP | 164 | 150 | 144 | 87.80% |
| IQ3_S | 164 | 151 | 144 | 87.80% |
| IQ3_S MTP | 164 | 152 | 147 | 89.63% |
| Model | C-Suite | Researcher | Coordinator | Librarian | Coder |
|---|---|---|---|---|---|
| IQ2_XS | 68.89 | 85.80 | 85.31 | 85.95 | 85.22 |
| IQ2_XS MTP | 72.84 | 86.22 | 84.90 | 85.74 | 87.35 |
| IQ2_S | 72.69 | 86.17 | 83.53 | 85.79 | 83.67 |
| IQ2_S MTP | 66.87 | 89.27 | 91.05 | 90.16 | 88.11 |
| IQ3_XXS | 72.95 | 85.70 | 85.15 | 86.05 | 87.57 |
| IQ3_XXS MTP | 69.41 | 84.24 | 83.53 | 84.49 | 85.35 |
| IQ3_S | 75.03 | 87.78 | 84.73 | 86.31 | 85.34 |
| IQ3_S MTP | 74.76 | 89.51 | 88.95 | 88.95 | 88.88 |
Long context was treated as a real deployment requirement, not a launch-parameter checkbox.
The fixed workload contained roughly 56,976 content tokens and ran inside a 73,728-token allocation.
Four Vast.ai shards ran matched base/MTP pairs on the same GPU class and software baseline. Scoring was completed in an isolated local Docker sandbox.
c8296709920f9c1ae168bfd5fe66f9f73637bd60