HMICIndependent local model test
HardwareRTX 5090 / 32 GB
Published18 September 2026
The Assembled Worker / Local model benchmark 01

Qwen deployment comparison

Qwen3.6-35B reached 3.82× the median decode rate.
Why start with Qwen3.8-27B?

Because speed and cold reliability selected different winners. The sparse 35B model was the stronger platform to build around. The dense 27B model was easier to trust before calibration.

01

The practical ruling

Cold starting point

Start with 27B.

It interpreted ordinary tasks more consistently and held exact response shape without a task-specific operating card.

Calibrated platform

Build around 35B.

It was dramatically faster, fit in less memory than its name suggests, and responded strongly when given a compact decision procedure.

3.82xMeasured decode-rate ratio
20/2427B grounded cold passes
12→1635B passes after one compact card
02

What was measured

ArmMedian per-call decodeGrounded holdoutExact surfaceMedian short completion
Qwen3.8-27B Q6_K
native
53.8 tok/s20/2424/240.920 s
Qwen3.6-35B-A3B Q4_K_M
native
205.3 tok/s12/2419/240.330 s
Qwen3.6-35B-A3B Q4_K_M
compact operating card
same deployment class16/2422/240.270 s

Speed figures summarize 48 measured calls on the same RTX 5090 deployment. The fresh holdout contained twelve unseen cases at two seeds: state extraction, exact delivery, causal transfer, and compressed operator commands.

03

What the aggregate score concealed

35B's weakness was often interface behavior.

The native 35B arm lost repeatable points on output shape and compressed operator language. One small card restored the entire causal-transfer family from 2/6 to 6/6. That is useful calibration response, even though the card did not make 35B universally better.

Some failures needed a different tool.

Every arm missed the unseen integer-only chain. A syntax-only CPU renderer changed no semantic score. That failure calls for a deterministic executor or verifier, not more formatting instructions.

04

Boundary of the claim

This is a deployment comparison, not a clean quantization theorem.

The two measured arms used different source packages and different quant levels: 27B Q6_K versus 35B-A3B Q4_K_M. The result supports a practical choice between those installed deployments. It does not isolate whether architecture, quantization, source packaging, or their interaction caused every observed difference.

05

The next test removes that ambiguity

Matched quant matrix

  • Q4_K_M versus Q6_K from one pinned source
  • RTX 5090, one RTX 3090, and two RTX 3090s
  • Identical context, runtime settings, prompts, and scoring
  • Speed, time to first token, memory, semantic accuracy, and exact format

The question

Where does Q6 buy measurable task capability over Q4_K_M, and where does that gain justify its memory and speed cost?

This is the next video and public result.

06

Watch the full analysis

07

Give us a failure worth testing

If you run either model, send the exact configuration and the first real task where it breaks. Useful requests become candidate cells in the next benchmark hand.

MODEL / SOURCE / QUANT / GPU / SYSTEM RAM / CONTEXT / RUNTIME / PROMPT / EXPECTED BEHAVIOR / OBSERVED FAILURE

Request a benchmark
08

Local evidence lineage

  • Fresh holdout: 72/72 calls completed on RTX 5090.
  • 27B artifact SHA-256: c9c206812fbe4ac7b76a729e25928b63f2ae89d37f69da7a71c20aec763cd436
  • 35B artifact SHA-256: 4ac6a06bce551257267f49ad2226f8671a22519ccc1a4dde9d5b433d1f2a410d
  • Scoring separates grounded semantic correctness from exact surface compliance.
09

Reproduce or audit it

Two experiments, two denominators

The download keeps the 48-call timing experiment separate from the 72-call frozen holdout. Every published statistic identifies its experiment, denominator, unit, aggregation method and source file.

Sanitized evidence pack

Includes public fixtures, exact model hashes, per-call outputs and measurements, calibration card, limitations, failures and deterministic manifests.

Download ZIP View manifest