LOCALBENCH
Controlled local-AI benchmark · September 2026

Eight quants.
One practical decision.

A base-versus-MTP comparison of every GGUF in ISTA-DASLab’s Qwen3.8-27B GSQ-RCO release—measured for speed, memory, long context, coding, research, knowledge work, tools, and C-Suite output.

Explore the results Watch on YouTube ↗ Model files ↗
8GGUF files tested
1,312HumanEval+ evaluations
192Applied outputs scored
64KLong-context workload passed
$2.95Final Vast.ai benchmark cost

The verdict

The fastest file was not automatically the best broad-use model. Quality was weighted around the actual jobs this local stack needs to perform; HumanEval+ contributed only 10% of the balanced index.

Best general-purpose choice

IQ3_S MTP

It produced the highest balanced score, nearly halved the adaptive long-task time versus IQ3_S, and led the Researcher and Coder role scores. The tradeoff is memory: only 1,673 MiB remained free in the representative run.

83.19balanced index
Fastest adaptive runIQ2_XS MTP · 48.29s

The speed-first choice, with a 32.7% elapsed-time reduction over its base pair.

Best non-MTP choiceIQ3_S · 81.34

The highest base-model balanced score and the top C-Suite role score.

Watch the findings

The 96-second explainer covers the experimental setup, $2.95 cost, MTP speed and memory tradeoffs, quality results, 64K context, and role-specific recommendations.

ISTA-DASLab Qwen3.8-27B Quant BenchmarkOpen on YouTube ↗

Best model by job

One universal winner is less useful than a routing policy. These scores reweight the same evidence for five distinct jobs in a private AI staff.

01

C-Suite work

IQ3_S75.03
02

Researcher / Scout

IQ3_S MTP89.51
03

Hermes Coordinator

IQ2_S MTP91.05
04

Librarian / Knowledge

IQ2_S MTP90.16
05

Forge / Coder

IQ3_S MTP88.88

All eight results

Switch the chart between quality, coding, elapsed time, and remaining memory. Lower is better for time; higher is better for the other views.

ModelBalancedHE+ pass@1Adaptive timePeak VRAMFree VRAMSelected profile
IQ3_S MTP WINNERQwen3.8-27B GSQ-RCO83.1989.63%53.49s14,176 MiB1,673 MiBMTP 3 · ubatch 256
IQ3_SQwen3.8-27B GSQ-RCO81.3487.80%105.79s13,286 MiB2,563 MiBMTP 0 · ubatch 1024
IQ3_XXSQwen3.8-27B GSQ-RCO81.2590.24%98.51s11,692 MiB4,159 MiBMTP 0 · ubatch 1024
IQ2_XS MTPQwen3.8-27B GSQ-RCO80.8090.24%48.29s11,270 MiB4,581 MiBMTP 3 · ubatch 512
IQ2_S MTPQwen3.8-27B GSQ-RCO80.6685.98%49.62s12,120 MiB3,731 MiBMTP 2 · ubatch 1024
IQ2_SQwen3.8-27B GSQ-RCO80.0484.76%65.45s10,896 MiB4,955 MiBMTP 0 · ubatch 1024
IQ2_XSQwen3.8-27B GSQ-RCO79.1285.98%71.73s10,222 MiB5,629 MiBMTP 0 · ubatch 1024
IQ3_XXS MTPQwen3.8-27B GSQ-RCO78.6887.80%49.76s13,066 MiB2,785 MiBMTP 3 · ubatch 1024

What MTP changed

MTP delivered meaningful acceleration in every matched pair, but it always consumed more VRAM and did not guarantee better broad-use quality.

IQ2_XS

+48.5% faster
Adaptive time
71.73s → 48.29s
Elapsed reduction
32.7%
Balanced index
+1.68
Additional VRAM used
1,048 MiB

IQ2_S

+31.9% faster
Adaptive time
65.45s → 49.62s
Elapsed reduction
24.2%
Balanced index
+0.62
Additional VRAM used
1,224 MiB

IQ3_XXS

+98.0% faster
Adaptive time
98.51s → 49.76s
Elapsed reduction
49.5%
Balanced index
−2.57
Additional VRAM used
1,374 MiB

IQ3_S

+97.8% faster
Adaptive time
105.79s → 53.49s
Elapsed reduction
49.4%
Balanced index
+1.85
Additional VRAM used
890 MiB
How to read “faster”: speedup is calculated as base time divided by MTP time, minus one. The separate elapsed-reduction figure answers the more intuitive question: how much less wall-clock time did the run take?

Quality beyond coding

The applied suite used 24 synthetic tasks per model across six practical categories. This prevented 164 coding problems from dominating a model intended for broad research, operations, and knowledge work.

Applied capability scores — all models
ModelBusinessResearchKnowledgeAgent toolsSysadminExecutive content
IQ2_XS47.9295.8396.8880.0078.7564.58
IQ2_XS MTP56.2595.8396.8880.0078.7560.42
IQ2_S56.2595.83100.0071.6778.7564.58
IQ2_S MTP38.54100.00100.0095.8378.7560.42
IQ3_XXS57.2992.7196.8875.8383.7572.92
IQ3_XXS MTP51.0492.7196.8875.8378.7564.58
IQ3_S59.38100.0096.8880.0073.7556.25
IQ3_S MTP56.25100.0096.8890.8378.7560.42
HumanEval+ counts — 164 problems per model
ModelEvaluatedHumanEval passedHumanEval+ passedHE+ pass@1
IQ2_XS16415014185.98%
IQ2_XS MTP16415514890.24%
IQ2_S16414513984.76%
IQ2_S MTP16414514185.98%
IQ3_XXS16415614890.24%
IQ3_XXS MTP16415014487.80%
IQ3_S16415114487.80%
IQ3_S MTP16415214789.63%
Role-weighted scores — all models
ModelC-SuiteResearcherCoordinatorLibrarianCoder
IQ2_XS68.8985.8085.3185.9585.22
IQ2_XS MTP72.8486.2284.9085.7487.35
IQ2_S72.6986.1783.5385.7983.67
IQ2_S MTP66.8789.2791.0590.1688.11
IQ3_XXS72.9585.7085.1586.0587.57
IQ3_XXS MTP69.4184.2483.5384.4985.35
IQ3_S75.0387.7884.7386.3185.34
IQ3_S MTP74.7689.5188.9588.9588.88

64K context

Long context was treated as a real deployment requirement, not a launch-parameter checkbox.

8/8completed the long-context workload

The fixed workload contained roughly 56,976 content tokens and ran inside a 73,728-token allocation.

What the test establishes

  • Every quant completed the same long Hermes-style task at the target class needed for a 64K deployment.
  • Every file also successfully allocated context through 131,072 tokens on the 16GB test GPU.
  • Flash Attention and Q4_0 K/V cache were held as the approved long-context baseline.
  • The largest MTP model still fit, but IQ3_S MTP left only 1,673 MiB of representative free VRAM.
Important: a successful 131K allocation probe is not the same as completing a 131K-token reasoning workload. The strongest end-to-end evidence here is the approximately 57K-token content task; 131K should be described as allocation capacity.

Methodology

Four Vast.ai shards ran matched base/MTP pairs on the same GPU class and software baseline. Scoring was completed in an isolated local Docker sandbox.

Controlled platform

NVIDIA RTX 5060 Ti 16GB

  • GPU VRAM reported: 16,311 MiB
  • CPU: AMD EPYC 7742, 64 threads
  • Driver: 595.91.07
  • llama.cpp: 0.5.0-dev
  • Commit: c8296709920f9c1ae168bfd5fe66f9f73637bd60
  • Four shards; one matched quant pair per shard
  • Final benchmark rental cost: $2.95
Balanced-index weights

Broad use beats one benchmark

Business analysis20% Research20% Knowledge work20% Agent / tool behavior15% Executive content10% HumanEval+10% Practical sysadmin5%
How the suite was scored
Each model produced three repeats of the fixed long-context workload plus tuned adaptive runs. The applied suite contributed 24 synthetic outputs per model across six categories. HumanEval+ contributed 164 problems per model. Role scores reweighted these same categories for C-Suite work, research, coordination, knowledge management, and coding. Speed and memory were reported separately so a fast but weaker response could not win simply by finishing first.
Limitations and interpretation
This is a controlled comparison of these eight files on one GPU class, not a universal model leaderboard. Applied-task samples are intentionally practical but finite; small score differences should be treated as directional. Rental prices vary over time. MTP gains depend on draft-token acceptance, workload, and tuning. The 131K result is an allocation probe, not a completed 131K active-token workload.
Decision rule
Choose IQ3_S MTP when one model must cover many jobs. Choose IQ3_S for C-Suite analysis without MTP overhead, IQ2_S MTP for coordinator and librarian roles, and IQ2_XS MTP when speed and headroom matter most. In a multi-model private AI staff, route by role rather than forcing one quant to do everything.